Graduate Software Engineer, AI Foundation Models Infrastructure

Posted 7 months ago
$341.73K / year

Are you applying to the internship?

Job Description

Graduate Software Engineer, AI Foundation Models Infrastructure | ByteDance

The Tone:
This is an early career role at ByteDance, focused on AI foundation models infrastructure. ByteDance is dedicated to pioneering advanced AI foundation models, leading in cutting-edge research, and driving technological and societal advancements through innovative products that inspire creativity and enrich life. This role is crucial for improving the reliability, performance, and scalability of the core systems that enable the development and deployment of these large-scale AI models.

The TL;DR
• Role: Early Career
• Location: Undisclosed (US, based on compensation & benefits provided)
• Pay: $177688–$341734 yearly
• Team: Seed Infrastructures team, overseeing distributed training, reinforcement learning framework, high-performance inference, and heterogeneous hardware compilation technologies.
• Mission: Enhance the reliability, performance, and efficiency of large-scale AI foundation model training and inference systems.
• Tech Stack: C++, Python, PyTorch, CUDA, NCCL, GPU, torch.profiler, Nsight, FSDP, Megatron-LM

What You’ll Actually Do
• System Improvement: Improve the reliability and performance of large-scale training systems across pre-training, fine-tuning, evaluation, and inference.
• Tool Development: Build observability, profiling, and debugging tools specifically for distributed ML workloads.
• Performance Optimization: Identify and optimize performance bottlenecks spanning GPU, networking, and storage layers.
• Framework Contribution: Contribute to the development and enhancement of distributed training frameworks in multi-GPU and multi-node environments.
• Cross-functional Collaboration: Collaborate with model and infrastructure teams to collectively improve system scalability and overall efficiency.

The Must-Haves
• Background: Early career individual completing or recently completed a PhD degree in Software Development, Computer Science, Computer Engineering, or a related technical discipline.
• Experience: Strong programming capabilities in C++ and Python, a solid understanding of PyTorch training workflows and distributed runtime behavior, and familiarity with CUDA execution, NCCL communication, and GPU systems fundamentals.
• Skills: C++, Python, PyTorch, CUDA, NCCL, GPU systems fundamentals.
• Bonus: Experience with performance profiling and debugging tools (e.g., torch.profiler, Nsight), familiarity with distributed training or parallelization strategies (e.g., FSDP, Megatron-LM), and the ability to analyze and optimize performance within complex ML training systems.

Related Jobs

AI-enabled Builder Early CareerFull Time
Stripe
Posted 21 hours ago In-person San Francisco, CA Software Development $120K - $180K / year