AI & Data Science Intern

Posted 6 months ago

Are you applying to the internship?

Job Description

AI & Data Science Intern | Progress Rail, a Caterpillar company

The Tone:
This is an internship at Progress Rail, located in Albertville, AL or Independence, KS. Progress Rail, a Caterpillar company, operates as an integrated rolling stock and infrastructure provider, delivering a comprehensive range of products and services to both domestic and international railroad customers. The company offers end-to-end railway solutions, encompassing everything from locomotives, transit, freight cars, and engines, to tracks, signals, and advanced technology, ensuring customers’ diverse rail needs are met. This crucial role supports the design and development of an AI-enabled solution, focusing on profiling, cleansing, standardizing, and modeling enterprise data within a Snowflake cloud data platform. The intern’s contributions will directly improve data quality, governance, and analytical readiness, accelerating the organization’s journey toward an AI-ready, governed enterprise data ecosystem.

The TL;DR
• Role: Internship
• Location: Albertville, AL or Independence, KS
• Team: Collaborates closely with Data Engineers, Data Architects, and Analytics teams.
• Mission: To support the design and development of an AI-enabled solution for profiling, cleansing, standardizing, and modeling enterprise data within a Snowflake cloud data platform.
• Tech Stack: Snowflake, Python, SQL, pandas, NumPy, scikit-learn, Power BI

What You’ll Actually Do
• Data Quality Analysis: Analyze structured enterprise datasets to identify data quality issues, and develop automated data profiling and quality checks using Python and SQL.
• Data Transformation: Design and implement data cleansing and standardization logic, supporting ELT workflows to transform raw source data into curated datasets within Snowflake.
• Data Modeling & Integration: Assist with logical and physical data modeling in Snowflake and write optimized SQL queries for validation, analytics, and downstream consumption.
• Applied AI / Machine Learning: Explore and prototype machine learning techniques to detect anomalies, automate mappings, or classify data attributes, assisting in evaluating AI-driven approaches.
• Cross-functional Collaboration: Work closely with Data Engineering, Analytics, and Governance teams, create clear technical documentation, and present findings and recommendations to stakeholders.

The Must-Haves
• Background: Currently pursuing a Bachelor’s or Master’s degree in Data Science, Computer Science, Engineering, Statistics, or a related field, with a foundational understanding of data structures, statistics, and analytics concepts.
• Experience: Coursework or hands-on experience in Python and SQL, exposure to Snowflake, cloud data platforms, or data warehousing concepts, and a basic understanding of ETL/ELT pipelines and data modeling techniques.
• Skills: Strong analytical thinking and problem-solving skills, ability to learn quickly, and capability to work with complex, real-world datasets. Familiarity with pandas, NumPy, scikit-learn, or similar Python libraries.
• Bonus: Experience with Power BI, data visualization, or metadata tools is a plus, along with an interest in AI, machine learning, and data governance within enterprise environments.

Related Jobs