Senior Software Engineer, AI Platform

Posted 7 months ago

Are you applying to the internship?

Job Description

Senior Software Engineer, AI Platform | Innovatech Solutions Inc.

The Tone:
This is a full-time role at Innovatech Solutions Inc., located in San Francisco, CA (Hybrid Remote). Innovatech Solutions Inc. is a technology company focused on solving complex challenges by creating transformative software and artificial intelligence products. This role is crucial for designing, developing, and deploying robust, scalable, and high-performance infrastructure and services that power the company’s next-generation AI offerings. Your work will directly empower businesses and individuals with intelligent tools, impacting millions and helping shape the future of technology by enabling efficient model training, deployment, and monitoring at scale.

The TL;DR
• Role: Full Time
• Type: Full-time
• Location: San Francisco, CA (Hybrid)

• Team: Engineering – Artificial Intelligence team, reports to Director of AI Engineering
• Mission: This person designs, develops, and deploys core AI/ML platform components, enabling data scientists and ML engineers to train, deploy, and monitor models efficiently at scale.
• Tech Stack: Python, Go, Java, C++, AWS, Azure, GCP, Kubernetes, Docker, Spark, Kafka, Flink, PostgreSQL, NoSQL databases, MLflow, Kubeflow, Airflow, CI/CD, TensorFlow, PyTorch, JAX

What You’ll Actually Do
• Platform Architecture: Lead the design, architecture, and implementation of critical components for our AI/ML platform, including data pipelines, model training frameworks, inference engines, and MLOps tools.
• System Optimization: Optimize existing systems and build new ones to handle increasing data volumes and computational demands, ensuring high availability, reliability, and performance.
• Cross-functional Partnership: Work closely with data scientists, ML engineers, and product managers to understand requirements, define technical specifications, and deliver integrated solutions.
• Cloud Infrastructure: Contribute to the management and automation of cloud infrastructure (AWS, Azure, GCP) supporting our AI platform, leveraging technologies like Kubernetes, Docker, and CI/CD pipelines.
• Technical Guidance: Mentor junior engineers, conduct code reviews, and contribute to setting best practices for software development, testing, and deployment.

The Must-Haves
• Background: Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field, with core domain knowledge in building and scaling distributed systems or AI/ML infrastructure.
• Experience: 5+ years of professional software development experience, with at least 2-3 years specifically focused on building and scaling distributed systems or AI/ML infrastructure.
• Skills:
* Programming Proficiency: Expert-level in Python (highly preferred), Go, Java, or C++.
* Distributed Systems: Strong understanding and practical experience with distributed systems principles, microservices architectures, and concurrency.
* Cloud Platforms: Hands-on experience with AWS, Azure, or GCP, including services relevant to AI/ML (e.g., S3/Blob Storage, EC2/VMs, Kubernetes, Sagemaker/Vertex AI).
* Data Technologies: Experience with big data technologies (e.g., Spark, Kafka, Flink) and databases (e.g., PostgreSQL, NoSQL databases).
* MLOps Knowledge: Familiarity with MLOps concepts, tools, and best practices for model lifecycle management (e.g., MLflow, Kubeflow, Airflow, CI/CD for ML).
• Bonus: Experience with Docker and Kubernetes; familiarity with deep learning frameworks like TensorFlow, PyTorch, or JAX; previous experience contributing to open-source AI/ML infrastructure projects; understanding of machine learning algorithms; experience in an agile development environment.

Related Jobs