PhD Research Intern

Posted 5 months ago

Are you applying to the internship?

Job Description

PhD Research Intern | Higharc

The Tone:
This is an internship at Higharc, located remotely in the United States. Higharc is a VC-backed startup dedicated to transforming how new homes are designed and built. This role is a 12-week engagement within the Special Projects team, focusing on cutting-edge research. The intern’s work will contribute directly to Higharc’s machine learning products through publishable research and reproducible code.

The TL;DR
• Role: Internship
• Type: Temporary
• Location: Remote

• Team: Special Projects team.
• Mission: To produce a publishable research contribution alongside reproducible code for integration with Higharc ML products.
• Tech Stack: PyTorch, RoboFlow, Modal, WandB.

What You’ll Actually Do
• Data Management: Build data pipelines and extract data from Higharc’s existing datasets.
• Pipeline Development: Design, implement, and execute semi-supervised and weakly-supervised Vision-Language Model (VLM) and segmentation training pipelines.
• Evaluation Design: Develop evaluation suites and error taxonomies for targeted multimodal tasks.
• Experimentation: Run rigorous ablations and scaling experiments, track results, and maintain reproducibility and research hygiene throughout.
• Documentation & Presentation: Document findings and present results through technical reports, demos, and a submission-ready draft.

The Must-Haves
• Background: Active enrollment in a PhD program in Computer Science, Machine Learning, or a related field at a U.S. institution.
• Experience: Demonstrated research ability through publications, strong preprints, open-source research code, or equivalent evidence of research impact; and experience building deep learning training workflows.
• Skills: Strong Python programming skills, solid understanding of computer vision, transformers, representation learning, ML experimentation practices, and comfort working in conventional research and engineering stacks (RoboFlow, Modal, WandB, or similar).
• Bonus: Experience with vision-language models or multimodal foundation models; experience designing and deploying semi-supervised learning methods (pseudo-labeling, self-training, distillation, consistency regularization) and familiarity with common failure modes; experience with multi-GPU training and large-scale experimentation; or familiarity with AEC data and workflows such as CAD/BIM concepts, plan understanding, or domain-specific labeling.

Related Jobs