← All openings
GI · Open role

Data Infrastructure Engineer

San Francisco, CAIn personFull-time

We're building the data engine behind Physical AI: the pipelines that turn raw multi-sensor capture into the training-grade datasets robot foundation models learn from. As our Data Infrastructure Engineer, you'll own the processing and quality-control pipeline that ingests terabytes of egocentric video, tactile and IMU streams, and pose trajectories, and turns them into datasets our customers train on.

We're a high-growth company: you'll work directly with the founders and with many world-tier robotics companies, in a hands-on role whose core challenges are infrastructure at scale — moving, storing, and processing massive volumes of multi-sensor data cheaply and reliably — alongside video processing, time synchronization across devices, trajectory validation, and the automated QC checks that decide whether an episode ships or gets thrown out.

What you'll do

  • Own and evolve the data processing pipeline that ingests raw multi-device captures (RGB video, IMU, magnetic encoder streams, and 6-DoF pose trajectories) and converts them into standardized training-ready datasets.
  • Design and build automated quality-control checks: dropped-frame and timestamp-gap detection, cross-device clock drift and jitter analysis, trajectory discontinuity and outlier detection, sensor dropout, and episode-level pass/fail gates.
  • Apply computer vision to QC: run and evaluate hand detection and reconstruction models (MediaPipe, WiLoR, HaMeR, YOLO-family) to measure hand coverage and visibility in egocentric video, and use OpenCV for frame-level checks, undistortion, and calibration validation.
  • Design and operate the cloud infrastructure and databases that keep data organized, queryable, and cheap to store at scale.
  • Build internal web tooling: dashboards and review UIs that let the team see what's wrong with a dataset or an episode in seconds rather than hours, backed by clean APIs.
  • Define and track data-quality metrics, review quality check results, and root-cause quality regressions across the pipeline.

What we're looking for

  • 3–6+ years of experience in data infrastructure, data engineering, or backend systems, with a track record of shipping work into production.
  • Strong Python (NumPy and the scientific stack), plus working knowledge of SQL and at least one of C++ or TypeScript/JavaScript. You should be comfortable moving between a processing script, a query, and a frontend in the same week.
  • Hands-on experience with AWS (S3, EC2, batch/serverless compute, IAM) and with designing relational database schemas for production workloads.
  • Experience with OpenCV and standard CV workflows, and exposure to running or fine-tuning detection/pose models, especially hand detection or human pose estimation.
  • Solid understanding of web application architecture: client/server split, REST APIs, and how a dashboard actually gets its data.
  • Solid understanding of systems fundamentals: operating systems (processes, memory, I/O, scheduling), GPU kernels and how work actually executes on an accelerator, and distributed systems (partitioning, replication, consistency, fault tolerance).
  • Comfort with time-series and multi-modal data: resampling, interpolation, and alignment across streams with different rates and clocks.
  • A quantitative, skeptical mindset about data quality. You assume the data is broken until a check proves otherwise.

Nice to have

  • CUDA programming and GPU acceleration.
  • Familiarity with ML frameworks such as PyTorch.
  • Hands-on work with robotics data: ROS/ROS 2, rosbag or MCAP, Zarr/Parquet, URDF, or robot learning datasets (LeRobot, Open X-Embodiment, UMI, DROID).
  • Video processing at scale: ffmpeg, codec and container tradeoffs, frame-accurate seeking.
  • Rigid-body transforms and pose math: quaternions, coordinate frame conventions, hand-eye and multi-camera calibration.
  • Distributed or parallel compute (Ray, Dask, multiprocessing) and containerized workflows.
  • Prior experience with imitation learning, VLA training data, or motion capture pipelines.

Compensation

  • Competitive cash base + performance-based bonus + equity.

Why join

You'll work in person in San Francisco alongside a team building foundational infrastructure for embodied intelligence, with direct ownership over the pipelines, tooling, and checks that determine the data quality our customers' models train on.

Apply via emailSend your resume and a note on what you've built to join@gilabs.xyz.