Are you applying to the internship?
Job Description
Data Engineer Intern | MedLaunch
The Tone:
This is an internship at MedLaunch. MedLaunch is a startup building a transformative healthcare accreditation platform, revolutionizing how healthcare organizations manage compliance and improve quality. This role offers a unique opportunity for aspiring Data Engineers to gain significant hands-on experience, contributing directly to a production-grade healthcare platform. Your contributions will directly impact our product and customers.
The TL;DR
• Role: Internship
• Type: Full-time
• Location: Remote
• Team: Engineering Team
• Mission: Develop critical components of our platform’s data infrastructure, from pipelines to analytics, directly impacting product functionality and customer insights.
• Tech Stack: Python, SQL, AWS (Lambda, S3, EMR, Athena, CloudWatch), MongoDB Atlas, PostgreSQL, Redis, Spark, Delta Lake, SageMaker, MLflow, Git, HIPAA
What You’ll Actually Do
• Data Pipeline Development: Build robust ETL pipelines using Python and SQL for processing diverse healthcare data sources, including healthcare facility lookups and compliance tracking.
• Data Storage and Optimization: Work hands-on with MongoDB Atlas and S3, design optimized SQL schemas, implement S3-based data lake architecture, and build caching systems using Redis.
• Analytics and ML Support: Create reliable data feeds for executive dashboards and KPI tracking, build healthcare-specific analytics pipelines, and establish data infrastructure for Machine Learning models.
• Healthcare Compliance and Quality: Implement stringent HIPAA-compliant data processing and audit trail systems, build data governance standards, and create automated monitoring for data quality issues.
The Must-Haves
• Background: Aspiring Data Engineer with a strong foundational understanding of data engineering principles and readiness for significant contributions.
• Experience: 1-3+ years with Python and SQL for data processing; 1-2 years AWS experience with services like Lambda or S3; practical knowledge of SQL databases and/or NoSQL (MongoDB preferred); solid understanding of ETL/data pipeline concepts and batch processing.
• Skills: Experience with data transformation and cleaning techniques; understanding of data warehousing and data lake concepts; basic knowledge of data quality and validation techniques; familiarity with version control (Git).
• Bonus: Experience in healthcare or regulated industries; knowledge of Spark, Delta Lake, or distributed computing frameworks; understanding of ML pipeline integration concepts; HIPAA compliance knowledge.