Skip to content

Jobs

Waymo

Robot type
Autonomous Vehicle
Location
Mountain ViewCaliforniaUSA
Job type
Artificial Intelligence
Posted
Apr 28, 2026
Salary
$213,000–$263,000 a year
Full-time

Senior Machine Learning Engineer, Runtime and Serving

Job description

The ML Optimization team at Waymo provides a set of tools to support and automate the lifecycle of the machine learning workflow, including feature and experiment management, model development, optimization and monitoring. These efforts have resulted in making machine learning more accessible to teams at Waymo, including Perception, Planner, Research and Simulation.

We are looking for engineers with ML software & systems expertise to help b uild the next generation Waymo onboard ML inference engine for Waymo fundamental model. You'll work across the entire ML stack from the system perspective, from efficient deep learning models, model compression, ML software (e.g.

Job responsibilities

  • Architect and develop an efficient, high-performance ML runtime and serving system tailored for both onboard autonomous vehicle compute and large-scale, offboard data center environments.
  • Lead the integration and feature development for ML inference runtimes across both domains, balancing the strict real-time latency and memory constraints of onboard systems with the high-throughput, highly concurrent…
  • Drive the strategic migration of ML workloads toward a JAX-native runtime architecture, which includes extending and modifying underlying ML compilers and runtimes (e.g., OpenXLA/PjRT, TensorRT).
  • Collaborate with world-class Waymo ML practitioners across perception, planner, and research to analyze system-level ML workloads and apply hardware-aware compute optimizations.
  • Design and build robust tooling for profiling, benchmarking, and identifying system-level bottlenecks across the end-to-end ML software stack.

Job requirements

  • B.S. or M.S. in CS, EE, Deep Learning or a related field
  • 5+ years of professional software engineering experience focused on building, scaling, or maintaining ML systems and infrastructure.
  • 5+ years production programming in C++.
  • 3+ years of production experience in Python and major deep learning frameworks (e.g., PyTorch, JAX).
  • Experience optimizing ML software for hardware accelerators (e.g., GPUs, TPUs, custom silicon).
  • Experience building low-latency, highly concurrent distributed backend systems.

Similar jobs