- Robot type
- Autonomous Vehicle
- Location
- AustinTexasUSA
- Job type
- Manufacturing
- Posted
- Sep 1, 2026
Full-time
Senior Test Engineer - Metrics Quality
Job description
Every decision Avride makes about its autonomous driving software merge or revert, ship or hold, this approach or that one — rests on a metric computed over a set of recorded scenes. Our QA organization is what stands between a metric and a decision made on a metric that was quietly wrong.
Job responsibilities
- Own the acceptance of new and changed metrics. Before a metric is used to decide whether a change is safe to ship, you decide whether it can be: what it should count, what it should not, and whether the implementation…
- Build the case space. For each metric, work out the full set of situations it has to handle — including all the ones where it must stay silent — and keep that set current as the technology and the operating environment…
- Hunt the silent failures. A metric that returns a surprising number gets noticed. A metric that returns nothing where it should have returned something does not, and that is the class of defect you are here to find.
- Read the implementation against the case space. Not for code quality — for the real situations it does not handle, and will therefore never report.
- Build the test data the job needs. The situations a metric has to handle are rarely all sitting in the data already.
- Make the case for a fix. Take findings to the engineer who owns the metric with concrete examples and a sense of scale: how often it is wrong, and what decisions that changes.
- Think like the people who read the numbers. A metric that is right on every individual scene can still fail to reveal a degradation across the whole set. Say so before anyone builds a release gate on it.
- Own the release cycle for metrics. Metrics ship on their own cadence. You qualify each release, run and maintain the regression that catches silent changes in what a metric means rather than only in what it returns,…
Job requirements
- 6+ years in software testing or test engineering, with real depth in data-heavy or analytical systems.
- Strong Python and SQL. You write analysis scripts and independent reference computations as a matter of routine, not as an exception.
- Statistical literacy. Distributions, variance, sample size, aggregation traps, and the confidence to say "this difference is noise".
- Product thinking. You ask what a number is for before you test it. Testing against the letter of a specification is the starting point of this job, not the substance of it.
- Conviction that survives a "that's how it works." You will regularly be the one telling a colleague that their metric does not behave the way they expect.
- Fluent use of LLMs as a working tool — analysis, test-data generation, reading unfamiliar code, and cross-checking your own reasoning.
- Clear written English. Your findings are read by people who will act on them.
