AML-MLsys Engineer

Posted 6 months ago
$187.2K / year

Are you applying to the internship?

Job Description

AML-MLsys Engineer | ByteDance

The Tone:
This is a full-time engineering role at ByteDance, focused on advanced machine learning systems within the US, while collaborating with a global team. ByteDance builds a suite of popular products including TikTok, Lemon8, and CapCut, aiming to inspire creativity and enrich life. This role is crucial for developing and maintaining the massively distributed ML training and inference systems that power cutting-edge applications like Large Language Models (LLM), AIGC, and Artificial General Intelligence (AGI). It offers a unique opportunity to shape the foundational infrastructure for future AI innovations across the company’s global operations.

The TL;DR
• Role: Full Time
• Type: Full-time
• Location: US (global team collaboration)
• Pay: $122574–$187200 yearly
• Team: AML-MLsys team
• Mission: To provide high-performance, highly reliable, and scalable systems for cutting-edge applications such as LLM/AIGC/AGI.
• Tech Stack: GPU, NPU, RDMA, Storage, Python, PyTorch, CUDA, FSDP, DeepSpeed, JAX SPMD, Megatron-LM, Verl, TensorRT-LLM, ORCA, VLLM, SGLang

What You’ll Actually Do
• Development: Develop and optimize LLM training, inference, and Reinforcement Learning (RL) frameworks.
• Collaboration: Work closely with model researchers to scale LLM training and RL to the next level of performance.
• Optimization: Be responsible for GPU and CUDA performance optimization to create an industry-leading high-performance LLM training, inference, and RL engine.
• System Building: Build large-scale heterogeneous systems, integrating technologies like GPU, NPU, RDMA, and Storage.
• Operations: Ensure the stable and reliable operation of massively distributed ML training and Inference systems/services around the world.

The Must-Haves
• Background: Bachelor’s degree or above in computer science, electronics, automation, software, or a related field.
• Experience: Proficient in algorithms and data structures, familiar with Python, and understand the basic principles of deep learning algorithms and neural network architectures.
• Skills: Familiarity with deep learning training frameworks such as PyTorch.
• Bonus: Proficient in GPU high-performance computing optimization technology on CUDA, with an in-depth understanding of computer architecture. Familiar with parallel computing optimization, memory access optimization, low-bit computing, and related techniques. Experience with frameworks like FSDP, DeepSpeed, JAX SPMD, Megatron-LM, Verl, TensorRT-LLM, ORCA, VLLM, SGLang, and knowledge of LLM models, especially with experience in accelerating LLM model optimization.

Related Jobs

NYL Insurance
Posted 18 hours ago Hybrid - 3 days per week - New York, USA Software Development $124K - $177K / year