Skip to content

Jobs

Merlin Labs

Robot type
Drone · Defense
Location
BostonMassachusettsUSA
Job type
Artificial Intelligence
Posted
Sep 17, 2026
Salary
$125,000–$220,000 a year
Full-time

Data Loop Pipeline Engineer

Job description

Merlin Labs — Data Loop Pipeline Engineer (Boston, Massachusetts · San Francisco, California (Remote OK)).

Job responsibilities

  • Build and operate ingestion pipelines for flight test, simulation and operational data, including format normalization, time alignment across sources, and quality validation at the gate.
  • Implement dataset curation and versioning with full lineage: what went into a dataset, from where, processed how, and by which version of which tool.
  • Build automated mining and triage for the rare and interesting — anomalies, disagreements between learned and rule-based systems, near-boundary events — so the flywheel prioritizes signal over volume.
  • Own labeling workflows and tooling, including quality control and inter-annotator agreement where human labeling is involved.
  • Close the loop: instrument deployed-system behavior so that operational signal reliably reaches the next training cycle without manual shepherding.
  • Work with Flight Test & Operations on logging and instrumentation requirements, and help resolve access constraints where data is captured by third-party systems.
  • Monitor pipeline health and dataset coverage; alert on drift, gaps and silent breakage.

Job requirements

  • You think data infrastructure is a craft. You have built pipelines that ran unattended for months and you have felt the specific satisfaction of a dataset that is exactly what it claims to be.
  • You will own the machinery that turns flight and simulation output into trustworthy, versioned, traceable training data — and turns deployed behavior back into the next training set.
  • Degree in Computer Science, Artificial Intelligence, Data Science, Computer Engineering, Applied Math, or a related subject.
  • 3+ years in data engineering, with production ownership of pipelines feeding ML training.
  • Strong Python and SQL; solid grasp of data modeling, storage formats and processing frameworks.
  • Experience with dataset versioning and lineage tooling.
  • Comfort with high-volume time-series and multimodal sensor data — telemetry, video, audio, structured logs.
  • Care about correctness. In this role a quiet data bug becomes a model behavior nobody can explain.

Similar jobs