Are you applying to the internship?
Job Description
Software Engineer, AI/ML Platform | InnovateX Solutions
The Tone:
This is a full-time role at InnovateX Solutions, located in [City, State] with hybrid remote options. InnovateX Solutions creates advanced platforms that support businesses and individuals worldwide. This role is crucial for developing the core infrastructure and tools that enable our data scientists and ML engineers to build, deploy, and manage AI/ML models effectively, directly influencing our ability to deliver value to customers.
The TL;DR
• Role: Full Time
• Type: Full-time
• Location: [City, State] (Hybrid Remote Options)
• Team: AI/ML Platform team
• Mission: Design, build, and maintain robust, scalable infrastructure and tools for developing, deploying, and managing AI/ML models across the company.
• Tech Stack: Python, Java, Go, C++, AWS, Azure, GCP, EC2, ECS, Kubernetes, S3, ADLS, GCS, Lambda, Functions, EMR, Dataproc, Spark, Flink, Kafka, PostgreSQL, Cassandra, MongoDB, Docker, Terraform, CloudFormation, TensorFlow, PyTorch, Scikit-learn, Kubeflow, MLflow, Airflow, Sagemaker, Vertex AI
What You’ll Actually Do
• Design & Develop: Architect, implement, and maintain scalable, reliable, and high-performance AI/ML platform components, including data pipelines, model training frameworks, inference services, and monitoring tools.
• Manage Infrastructure: Deploy and manage ML workloads using cloud infrastructure (AWS, Azure, GCP), ensuring optimal resource utilization and cost efficiency.
• Build Tooling: Create and enhance internal tools and automation scripts to streamline the ML lifecycle, from experimentation and data preparation to model deployment and monitoring.
• Collaborate: Partner with data scientists, ML engineers, product managers, and other engineering teams to understand needs and deliver platform capabilities that accelerate their work.
• Optimize Performance: Identify and resolve performance bottlenecks to ensure the platform efficiently handles large-scale data and complex ML models.
The Must-Haves
• Background: Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience, with a focus on building scalable backend systems or platform infrastructure.
• Experience: 3+ years of professional software engineering experience, including hands-on work with at least one major cloud provider (AWS, Azure, GCP) and its relevant services (e.g., EC2/ECS/Kubernetes, S3/ADLS/GCS, Lambda/Functions, EMR/Dataproc), as well as large-scale data processing frameworks and databases.
• Skills: Strong proficiency in Python (highly preferred), Java, Go, or C++; experience designing and implementing distributed systems and microservices; a basic understanding of machine learning concepts and the MLOps lifecycle; excellent analytical, problem-solving, and debugging skills.
• Bonus: Experience building or contributing to AI/ML platforms or MLOps tools (e.g., Kubeflow, MLflow, Airflow, Sagemaker, Vertex AI); familiarity with containerization (Docker, Kubernetes) and CI/CD pipelines/infrastructure as code (Terraform, CloudFormation); knowledge of machine learning frameworks (e.g., TensorFlow, PyTorch, Scikit-learn).