- Robot type
- Autonomous Vehicle
- Location
- Mountain ViewCaliforniaUSA
- Job type
- Artificial Intelligence
- Posted
- Jul 22, 2026
- Salary
- $251,000–$310,000 a year
Staff Machine Learning Engineer - Vision-Language Foundation Models
Job description
The Team & Mission: In the Oracle Perception team, our mission is to build the ultimate cognitive engine for autonomous driving. We are pioneering the use of large multimodal foundation models (e.g., Gemini) to build a powerful offboard reasoning and data flywheel system. We are moving beyond traditional perception to true scene understanding and driving actions—building offboard models that can comprehend complex driving problems, predict object/scene dynamics, and deduce driving paths with logical rationale.
Our core focus is advancing the VLM foundation itself. By pushing the boundaries of multimodal pre-training and state-of-the-art post-training (SFT, RL) , we are creating models…
Job responsibilities
- Drive Pre-training & Domain Adaptation: Lead the technical strategy for curating and constructing massive-scale, high-quality multimodal pre-training datasets.
- Lead Post-Training & Reasoning Enhancement: Design and implement state-of-the-art fine-tuning (SFT) and Reinforcement Learning (RLHF/RLAIF, DPO/GRPO/PPO) pipelines.
- Pioneer the VLM Data Flywheel: Architect the highly scalable inference and evaluation pipelines that leverage these trained Gemini-class models to autonomously source, sample, and autolabel critical edge cases,…
- Define Training Recipes & Scaling Laws: Conduct rigorous ablation studies to optimize model architectures, token budgets, and loss functions.
- Drive Cross-Functional AI Strategy: Act as the principal technical visionary across ML Infra, Perception, Behavior, and AI Foundation teams.
- Provide Staff-Level Technical Leadership: Own the long-term technical roadmap for foundation model development. Mentor senior engineers, lead rigorous design reviews, and establish standard-setting engineering…
Job requirements
- Master’s degree in Computer Science, AI, ML, or a related technical field.
- 8+ years of hands-on experience designing, training, and scaling deep learning models, with at least 3+ years focused deeply on training Large Language Models (LLMs) or Vision-Language Models (VLMs) .
- Proven expertise in the full lifecycle of Foundation Models: from pre-training data curation (interleaved formats, tokenization) and distributed training to advanced post-training techniques.
- Expert-level understanding of training infrastructure and distributed paradigms (e.g., FSDP, Megatron, JAX/Pax) required for training massive models reliably.
- Expert-level software engineering fundamentals using Python, PyTorch, or JAX, with a track record of building reliable, highly scalable ML systems.
- Proven ability to operate with high ambiguity, define technical roadmaps, and drive complex, multi-quarter technical initiatives across multiple teams in a fast-paced environment.
Similar jobs
Waymo · Artificial Intelligence
Staff Tech Lead, ML Data Infrastructure and Inference Platform
Mountain View, California
Waymo · Artificial Intelligence
2027 Summer Intern, MS/PhD, Road Understanding, ML Engineer
Mountain View, California
Waymo · Artificial Intelligence
Research Scientist, World Model Post-Training
Mountain View, California · San Francisco, California · New York, New York
Waymo · Artificial Intelligence
2027 Summer Intern, MS/PhD, Machine Learning Engineer - Simulator Realism Evaluation
San Francisco, California
