Are you applying to the internship?
Job Description
Senior Software Engineer – AI/ML Platform | [Your Company Name]
The Tone:
This is a full-time role at [Your Company Name], with remote options available. [Your Company Name] is a leader in [industry], dedicated to transforming how the world [company’s core mission]. This role is crucial for designing, building, and scaling the fundamental infrastructure and tools that empower our data scientists and machine learning engineers to create advanced AI models, directly impacting our product development and intelligent application future.
The TL;DR
• Role: Full Time
• Type: Full-time
• Location: Flexible / Remote Options Available
• Team: AI/ML Platform team
• Mission: Design, build, and scale the core infrastructure and tools that empower data scientists and machine learning engineers to develop, deploy, and manage AI models across our product suite.
• Tech Stack: Python, Go, Java, C++, AWS, GCP, Azure, Docker, Kubernetes, Spark, Flink, Kafka, object storage, relational databases, NoSQL databases, TensorFlow, PyTorch, Scikit-learn, Terraform, CloudFormation
What You’ll Actually Do
• Design and develop scalable, robust, and high-performance infrastructure for AI/ML model training, inference, and deployment.
• Build and maintain core platform services, including feature stores, model registries, experiment tracking, and monitoring systems.
• Optimize existing AI/ML pipelines for performance, cost-efficiency, and reliability, leveraging cloud-native technologies.
• Collaborate closely with data scientists and ML engineers to understand their needs and translate them into platform features and improvements.
• Drive architectural decisions and provide technical leadership on complex projects within the team and across engineering.
The Must-Haves
• Background: Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field, with a focus on building scalable backend systems or infrastructure.
• Experience: 5+ years of professional software development experience, including extensive work with cloud platforms (AWS, GCP, or Azure) and containerization technologies (Docker, Kubernetes).
• Skills: Proficiency in at least one modern programming language (Python, Go, Java, or C++); solid understanding of distributed systems principles, microservices architecture, and API design; strong problem-solving skills with an ability to debug complex systems and optimize performance; excellent communication and collaboration skills.
• Bonus: Experience building or contributing to AI/ML platforms or MLOps tools; familiarity with ML frameworks such as TensorFlow, PyTorch, or Scikit-learn; experience with infrastructure-as-code tools like Terraform or CloudFormation; knowledge of big data technologies and data warehousing concepts.