Skip to content

Jobs

Nuro

Robot type
Autonomous Vehicle
Location
Mountain ViewCaliforniaUSA
Job type
Artificial Intelligence
Posted
Aug 11, 2026
Full-time

Software Engineer, Applied AI Infrastructure

Job description

Frontier models are fungible. Any team can rent the same intelligence we can, and the model we build on today will be replaced within a month. What is not fungible is the infrastructure that decides whether an autonomous system's output can be trusted — evaluation, verification, and the discipline to gate on evidence instead of impressions. Nuro has spent a decade building exactly that discipline for a robot that drives on public roads, and this team turns it inward: we build the platform that lets AI agents operate autonomously inside Nuro's own engineering organization, under the same standard of proof we apply to the vehicle.

Our mandate is to amplify the output of every engineer and…

Job responsibilities

  • Build the closed-loop measurement layer that tells us, per workflow, whether agent output is accepted, reverted, or overridden — and use it to decide where autonomy expands and where it gets pulled back.
  • Take the autoresearch loop from assisted to unattended for a bounded class of experiments, including the eval and confidence machinery required to run it without a human in the loop.
  • Design the isolation and permissioning model that lets agents act on production repositories and infrastructure with an auditable record of what they did and why.

Job requirements

  • 3+ years of software engineering experience (or 2+ with a Master's) in computer science, engineering, or equivalent practical experience. Staff-level candidates should bring correspondingly deeper scope and ownership.
  • Deep, current taste in LLM research. You understand how a model is trained from scratch — data, tokenization, architecture, pretraining dynamics, the full post-training stack of supervised fine-tuning, preference…
  • You know what happens under the hood at inference. Attention and KV-cache behavior, batching and scheduling, quantization, speculative decoding, prefix caching, context handling, and how each trades off latency,…
  • You have built and operated LLM-based agent systems in production — tool use, orchestration, sandboxing, retrieval, memory — and you know where they break.
  • Strong backend and distributed systems background at scale: cloud infrastructure, service design, storage, queuing, and the judgment to build things that stay up.
  • Strong programming skills in Python.
  • You are opinionated about evaluation. You have argued with someone about whether a benchmark measured anything real, and you were right.
  • You work end-to-end and do not need the problem pre-decomposed. This role has more surface than a specification.

Similar jobs