Data Infrastructure Engineer
We're building the data engine behind Physical AI: the pipelines that turn raw multi-sensor capture into the training-grade datasets robot foundation models learn from. As our Data Infrastructure Engineer, you'll own the processing and quality-control pipeline that ingests terabytes of egocentric video, tactile and IMU streams, and pose trajectories, and turns them into datasets our customers train on.
We're a high-growth company: you'll work directly with the founders and with many world-tier robotics companies, in a hands-on role whose core challenges are infrastructure at scale — moving, storing, and processing massive volumes of multi-sensor data cheaply and reliably — alongside video processing, time synchronization across devices, trajectory validation, and the automated QC checks that decide whether an episode ships or gets thrown out.
What you'll do
- Own and evolve the data processing pipeline that ingests raw multi-device captures (RGB video, IMU, magnetic encoder streams, and 6-DoF pose trajectories) and converts them into standardized training-ready datasets.
- Design and build automated quality-control checks: dropped-frame and timestamp-gap detection, cross-device clock drift and jitter analysis, trajectory discontinuity and outlier detection, sensor dropout, and episode-level pass/fail gates.
- Apply computer vision to QC: run and evaluate hand detection and reconstruction models (MediaPipe, WiLoR, HaMeR, YOLO-family) to measure hand coverage and visibility in egocentric video, and use OpenCV for frame-level checks, undistortion, and calibration validation.
- Design and operate the cloud infrastructure and databases that keep data organized, queryable, and cheap to store at scale.
- Build internal web tooling: dashboards and review UIs that let the team see what's wrong with a dataset or an episode in seconds rather than hours, backed by clean APIs.
- Define and track data-quality metrics, review quality check results, and root-cause quality regressions across the pipeline.
What we're looking for
- 3–6+ years of experience in data infrastructure, data engineering, or backend systems, with a track record of shipping work into production.
- Strong Python (NumPy and the scientific stack), plus working knowledge of SQL and at least one of C++ or TypeScript/JavaScript. You should be comfortable moving between a processing script, a query, and a frontend in the same week.
- Hands-on experience with AWS (S3, EC2, batch/serverless compute, IAM) and with designing relational database schemas for production workloads.
- Experience with OpenCV and standard CV workflows, and exposure to running or fine-tuning detection/pose models, especially hand detection or human pose estimation.
- Solid understanding of web application architecture: client/server split, REST APIs, and how a dashboard actually gets its data.
- Solid understanding of systems fundamentals: operating systems (processes, memory, I/O, scheduling), GPU kernels and how work actually executes on an accelerator, and distributed systems (partitioning, replication, consistency, fault tolerance).
- Comfort with time-series and multi-modal data: resampling, interpolation, and alignment across streams with different rates and clocks.
- A quantitative, skeptical mindset about data quality. You assume the data is broken until a check proves otherwise.
Nice to have
- CUDA programming and GPU acceleration.
- Familiarity with ML frameworks such as PyTorch.
- Hands-on work with robotics data: ROS/ROS 2, rosbag or MCAP, Zarr/Parquet, URDF, or robot learning datasets (LeRobot, Open X-Embodiment, UMI, DROID).
- Video processing at scale: ffmpeg, codec and container tradeoffs, frame-accurate seeking.
- Rigid-body transforms and pose math: quaternions, coordinate frame conventions, hand-eye and multi-camera calibration.
- Distributed or parallel compute (Ray, Dask, multiprocessing) and containerized workflows.
- Prior experience with imitation learning, VLA training data, or motion capture pipelines.
Compensation
- Competitive cash base + performance-based bonus + equity.
Why join
You'll work in person in San Francisco alongside a team building foundational infrastructure for embodied intelligence, with direct ownership over the pipelines, tooling, and checks that determine the data quality our customers' models train on.