- Robot type
- Autonomous Vehicle
- Location
- Mountain ViewCaliforniaUSA
- Job type
- Artificial Intelligence
- Posted
- Apr 28, 2026
- Salary
- $213,000–$263,000 a year
Senior Machine Learning Engineer, Runtime and Serving
Job description
The ML Optimization team at Waymo provides a set of tools to support and automate the lifecycle of the machine learning workflow, including feature and experiment management, model development, optimization and monitoring. These efforts have resulted in making machine learning more accessible to teams at Waymo, including Perception, Planner, Research and Simulation.
We are looking for engineers with ML software & systems expertise to help b uild the next generation Waymo onboard ML inference engine for Waymo fundamental model. You'll work across the entire ML stack from the system perspective, from efficient deep learning models, model compression, ML software (e.g.
Job responsibilities
- Architect and develop an efficient, high-performance ML runtime and serving system tailored for both onboard autonomous vehicle compute and large-scale, offboard data center environments.
- Lead the integration and feature development for ML inference runtimes across both domains, balancing the strict real-time latency and memory constraints of onboard systems with the high-throughput, highly concurrent…
- Drive the strategic migration of ML workloads toward a JAX-native runtime architecture, which includes extending and modifying underlying ML compilers and runtimes (e.g., OpenXLA/PjRT, TensorRT).
- Collaborate with world-class Waymo ML practitioners across perception, planner, and research to analyze system-level ML workloads and apply hardware-aware compute optimizations.
- Design and build robust tooling for profiling, benchmarking, and identifying system-level bottlenecks across the end-to-end ML software stack.
Job requirements
- B.S. or M.S. in CS, EE, Deep Learning or a related field
- 5+ years of professional software engineering experience focused on building, scaling, or maintaining ML systems and infrastructure.
- 5+ years production programming in C++.
- 3+ years of production experience in Python and major deep learning frameworks (e.g., PyTorch, JAX).
- Experience optimizing ML software for hardware accelerators (e.g., GPUs, TPUs, custom silicon).
- Experience building low-latency, highly concurrent distributed backend systems.
Similar jobs
Waymo · Artificial Intelligence
Staff Tech Lead, ML Data Infrastructure and Inference Platform
Mountain View, California
Waymo · Artificial Intelligence
2027 Summer Intern, MS/PhD, Road Understanding, ML Engineer
Mountain View, California
Waymo · Artificial Intelligence
Research Scientist, World Model Post-Training
Mountain View, California · San Francisco, California · New York, New York
Waymo · Artificial Intelligence
2027 Summer Intern, MS/PhD, Machine Learning Engineer - Simulator Realism Evaluation
San Francisco, California
