Are you applying to the internship?
Job Description
Senior AI/ML Platform Engineer | InnovateCorp
The Tone:
This is a full-time role at InnovateCorp, with options for remote work. InnovateCorp is a technology pioneer dedicated to building intelligent systems that solve complex real-world problems using artificial intelligence and machine learning. This pivotal role is instrumental in designing, building, and maintaining the core infrastructure and tools that power all AI and ML initiatives, shaping the future of AI within the company.
The TL;DR
• Role: Full Time
• Type: Full-time
• Location: Remote
• Mission: To design, build, and maintain scalable, reliable, and efficient AI/ML platforms and services that enable seamless model development, deployment, and monitoring.
• Tech Stack: AWS, Azure, GCP, Docker, Kubernetes, Terraform, CloudFormation, Pulumi, Sagemaker, Vertex AI, MLflow, Spark, Flink, Kafka, Python, Go, Java, Scala
What You’ll Actually Do
• Design & Develop: Architect, build, and maintain robust, scalable, and secure AI/ML platforms and services using cloud-native technologies and containerization.
• Tooling & Automation: Develop automated CI/CD pipelines for ML models, experiment tracking systems, feature stores, and model serving infrastructure.
• Performance & Optimization: Optimize the performance, reliability, and cost-efficiency of the ML infrastructure, including distributed training and inference systems.
• Collaboration: Work closely with data scientists, ML engineers, and other engineering teams to understand their needs and provide technical leadership and support.
• Best Practices: Advocate for and implement best practices in MLOps, software engineering, security, and data governance within the AI/ML ecosystem.
The Must-Haves
• Background: Bachelor’s or Master’s degree in Computer Science, Engineering, or a related quantitative field, with core domain knowledge in software engineering and AI/ML platforms.
• Experience: 5+ years of professional experience in software engineering, with at least 3 years focused on building and scaling AI/ML platforms or large-scale distributed systems.
• Skills: Expert-level proficiency in Python, hands-on experience with at least one major cloud provider (AWS, Azure, or GCP), strong understanding and practical experience with Docker and Kubernetes, and deep understanding of MLOps principles.
• Bonus: Experience contributing to open-source ML platforms or tools, knowledge of security best practices in cloud environments, and prior experience with specific ML frameworks like TensorFlow or PyTorch.