Data Infrastructure Engineer, Intern
We're building the data engine behind Physical AI: the pipelines that turn raw multi-sensor capture into the training-grade datasets robot foundation models learn from. As our Data Infrastructure Intern, you'll work alongside the team on the processing and quality-control pipeline that ingests terabytes of egocentric video, tactile and IMU streams, and pose trajectories, and turns them into datasets our customers train on.
We're a high-growth company: you'll work directly with the founders and with many world-tier robotics companies, in a hands-on role whose core challenges are infrastructure at scale — moving, storing, and processing massive volumes of multi-sensor data cheaply and reliably — alongside video processing, time synchronization across devices, trajectory validation, and the automated QC checks that decide whether an episode ships or gets thrown out.
What you'll do
- Contribute to the data processing pipeline that ingests raw multi-device captures (RGB video, IMU, magnetic encoder streams, and 6-DoF pose trajectories) and converts them into standardized training-ready datasets.
- Build automated quality-control checks with the team: dropped-frame and timestamp-gap detection, cross-device clock drift and jitter analysis, trajectory discontinuity and outlier detection, sensor dropout, and episode-level pass/fail gates.
- Apply computer vision to QC: run and evaluate hand detection and reconstruction models (MediaPipe, WiLoR, HaMeR, YOLO-family) to measure hand coverage and visibility in egocentric video, and use OpenCV for frame-level checks, undistortion, and calibration validation.
- Work with our cloud infrastructure and databases to keep data organized, queryable, and cheap to store at scale.
- Have the opportunity to build internal web tooling: dashboards and review UIs that let the team see what's wrong with a dataset or an episode in seconds rather than hours, backed by clean APIs.
- Help review quality check results and iterate on improvements to the pipeline.
What we're looking for
- Currently pursuing (or recently completed) a BS/MS/PhD in CS, EE, robotics, or a related field, available for a full-time in-person internship in San Francisco.
- Strong Python (NumPy and the scientific stack), plus working knowledge of SQL and at least one of C++ or TypeScript/JavaScript. You should be comfortable moving between a processing script, a query, and a frontend in the same week.
- Familiarity with AWS (S3, EC2, batch/serverless compute, IAM basics) and with designing or using a relational database schema.
- Hands-on experience with OpenCV and standard CV workflows, and exposure to running or fine-tuning detection/pose models, especially hand detection or human pose estimation.
- Basic understanding of web application architecture: client/server split, REST APIs, and how a dashboard actually gets its data.
- Comfort with time-series and multi-modal data: resampling, interpolation, and alignment across streams with different rates and clocks.
- A quantitative, skeptical mindset about data quality. You assume the data is broken until a check proves otherwise.
Nice to have
- CUDA programming and GPU acceleration.
- Familiarity with ML frameworks such as PyTorch.
- Hands-on work with robotics data: ROS/ROS 2, rosbag or MCAP, Zarr/Parquet, URDF, or robot learning datasets (LeRobot, Open X-Embodiment, UMI, DROID).
- Video processing at scale: ffmpeg, codec and container tradeoffs, frame-accurate seeking.
- Rigid-body transforms and pose math: quaternions, coordinate frame conventions, hand-eye and multi-camera calibration.
- Distributed or parallel compute (Ray, Dask, multiprocessing) and containerized workflows.
- Prior experience with imitation learning, VLA training data, or motion capture pipelines.
Compensation
- Competitive hourly rate + equity consideration for exceptional interns. Return-offer track for full-time.
Why join
You'll work in person in San Francisco alongside a team building foundational infrastructure for embodied intelligence, with real scope over the tooling and checks that determine data quality, plus the rare intern experience of shipping something that goes straight into production and into customers' training runs.