Graduate AI Infrastructure Engineer

Posted 6 months ago
$341.73K / year

Are you applying to the internship?

Job Description

Graduate AI Infrastructure Engineer | ByteDance

The Tone:
This is an early career position at ByteDance, located in the US. ByteDance creates innovative products built to help people authentically express themselves, discover, and connect. This role is crucial for advancing AI foundation model development, offering an unparalleled opportunity to pursue bold ideas, tackle complex challenges, and unlock limitless growth in the field of AI infrastructure.

The TL;DR
• Role: Early Career
• Type: Full-time
• Location: US-based (specific city not listed)
• Pay: $177688–$341734 annually
• Team: Seed Infrastructures team is responsible for foundational technologies that power AI foundation models.
• Mission: This person designs, builds, and improves systems that enable large-scale agent development and evaluation.
• Tech Stack: Python, Ray, PyTorch Distributed, Kubernetes

What You’ll Actually Do
• Design: Design and build scalable agent evaluation frameworks for multi-step, tool-using, and environment-interacting agents.
• Develop: Develop reproducible rollout systems, including trajectory logging, execution formats, and experiment replay capabilities.
• Build: Build distributed execution pipelines to support large-scale agent evaluation, RLHF, and other post-training workloads.
• Create: Develop sandboxed and cloud-based environments to ensure safe and reliable agent experimentation.
• Enhance: Improve the scalability, reliability, and performance of agent training and evaluation systems.

The Must-Haves
• Background: Doctorate degree in Software Development, Computer Science, Computer Engineering, or a related technical discipline. Early Career.
• Experience: Strong programming skills in Python and demonstrated experience building large-scale systems or open-source frameworks. Familiarity with reinforcement learning workflows (e.g., rollout, trajectory collection, policy optimization). Experience with distributed systems or parallel execution frameworks (e.g., Ray, PyTorch Distributed, Kubernetes).
• Skills: Python programming, distributed systems, parallel execution frameworks, reinforcement learning workflows, scalability, system reliability.
• Bonus: Experience in agent systems, evaluation frameworks, or RLHF pipelines. Open-source contributions or publications in ML systems, agent infrastructure, or distributed RL.

Related Jobs