Graduate AI Infrastructure Engineer

Posted 7 months ago
$341.73K / year

Are you applying to the internship?

Job Description

Graduate AI Infrastructure Engineer | ByteDance

The Tone:
This is an early career position at ByteDance, located in various US locations. The company’s Seed Infrastructures team is dedicated to pioneering advanced AI foundation models by overseeing distributed training, reinforcement learning frameworks, high-performance inference, and heterogeneous hardware compilation technologies. This role is crucial for conducting research and development on large-scale AI infrastructure, supporting the efficient training and post-training of cutting-edge models and tackling complex challenges to unlock limitless growth.

The TL;DR
• Role: Early Career
• Location: US (various locations)
• Pay: $177688–$341734 yearly
• Team: Seed Infrastructures team, which oversees distributed training, reinforcement learning frameworks, high-performance inference, and heterogeneous hardware compilation technologies for AI foundation models.
• Mission: Conduct research and development on large-scale AI infrastructure to support efficient training and post-training of foundation models, multimodal LLMs, and image/video generation models.
• Tech Stack: Python and/or C++, PyTorch and distributed training tools

What You’ll Actually Do
• Research: Conduct research and development on large-scale AI infrastructure to support efficient training and post-training of foundation models, multimodal LLMs, and image/video generation models.
• Optimize: Design and optimize distributed training strategies, including various parallelism techniques, computation–communication overlap, and large-scale GPU cluster scaling.
• Prototype: Prototype and improve end-to-end reinforcement learning (RL) training systems, covering rollout generation, policy optimization, evaluation, and iterative deployment workflows.
• Build: Build scalable and fault-tolerant infrastructure that operates reliably under dynamic workloads and heterogeneous compute environments.
• Analyze: Analyze performance bottlenecks across the training stack and develop principled optimization approaches to improve throughput, efficiency, and stability.

The Must-Haves
• Background: Individuals completing or having recently completed a PhD degree in Software Development, Computer Science, Computer Engineering, or a related technical discipline, with a strong background in distributed systems, large-scale machine learning systems, or deep learning infrastructure.
• Experience: Research or hands-on experience in training or optimizing large-scale models (e.g., LLMs, multimodal models, RL systems), with an understanding of parallelism strategies and distributed training concepts. Familiarity with reinforcement learning workflows such as rollout generation, policy optimization, and evaluation loops.
• Skills: Proficiency in programming (e.g., Python and/or C++) and experience with modern ML frameworks (e.g., PyTorch and distributed training tools).

Related Jobs

AI-enabled Builder Early CareerFull Time
Stripe
Posted 18 hours ago In-person San Francisco, CA Software Development $120K - $180K / year