Skip to content

Jobs

Avride

Robot type
Autonomous Vehicle
Location
AustinTexasUSA
Job type
Business Operations
Posted
Aug 19, 2026
Full-time

Lead Evaluation Engineer – Scenario Coverage & Datasets (Autonomous Driving)

Job description

Avride builds autonomous driving technology for vehicles and delivery robots. Our QA organization measures how well that technology drives, largely in simulation — and an evaluation is only as good as the data behind it. So we build our own datasets, keep track of what they cover, and keep widening that coverage as the technology takes on more of the road.

Job responsibilities

  • Own the evaluation dataset. What is in it, what is missing, and what it lets us claim about autonomous driving quality.
  • Own the coverage model. Our scenario taxonomy and ODD parameter space have to stay accurate as the technology matures and the operating environment changes.
  • Close the distance between scenarios and scenes. For each thin area, choose the method that will actually produce the scenes you need — mining, simulation or the field — and know what it costs before you spend it.
  • Set the mining agenda for the QA team. Decide what each area owner should be looking for next, review what comes back, and keep a regular rhythm for collecting their feedback.
  • Be the internal customer for our mining tooling. Use it yourself, turn what you learn into requirements for the development and analytics teams, and keep the feedback loop between QA and engineering running so the…
  • Bring in what works elsewhere. Track how other AV programs, research groups and vendors solve coverage, edge-case curation and data selection, and turn what is worth having into concrete proposals.
  • Work through the teams you depend on. Coverage is only visible once scenes are labeled and only measurable when the right metrics exist, which makes labeling and analytics standing partners rather than occasional ones.
  • Report coverage and readiness. Put them in a form a release decision can be made on — clear about confidence and about blind spots.

Job requirements

  • Evaluation, validation, or simulation in AV, robotics, or perception
  • Detection engineering or threat detection , where the attack space is unbounded, there's no ground truth per event, and you're always trading false positives against false negatives
  • Search relevance or recommendations , building judged sets with coverage across query types
  • LLM or ML model evaluation , building eval sets and benchmarks and finding what they fail to catch
  • Speech recognition, medical AI, or fraud and risk modeling , curating test sets for the long tail, edge cases, and drift
  • 8+ years in software testing, test engineering, validation, or evaluation, including 2+ years owning the test or evaluation strategy for a full system rather than a feature area.
  • A track record of taking on more. You've owned an area end to end, led or mentored engineers, and improved how your team works rather than only executing within it.
  • Coverage as a first-class problem. You can talk about equivalence classes, parameter spaces, risk-based prioritization, and what would have to be true for a dataset to be enough .

Similar jobs