Skip to content

Jobs

PlusAI

Robot type
Autonomous Vehicle
Location
Santa ClaraCaliforniaUSA
Job type
Artificial Intelligence
Posted
Sep 14, 2026
Salary
$170,000–$260,000 a year
Full-time

Senior/Staff Research Engineer — Vision-Language-Action Models (Autonomous Driving)

Job description

You will join our core AI team at the frontier of autonomous decision-making, building the Vision-Language-Action (VLA) models that form SuperDrive's reasoning layer. You'll train VLA models that generate high-level driving decisions and trajectory guidance for on-board strategic decision-making, and design the knowledge distillation and compression techniques that transition large models onto on-board compute.

Job responsibilities

  • Design, train, and evaluate Vision-Language-Action models that generate high-level driving decisions and trajectory guidance in support of Plus's reasoning layer.
  • Own a VLA workstream end to end — data, architecture, large-scale training, and on-vehicle validation.
  • Build training and evaluation pipelines and rigorous metrics for VLA performance in driving contexts.
  • Develop distillation and compression recipes to deploy large reasoning models on on-board compute.
  • Apply SFT and RL post-training to improve reasoning, robustness, and long-tail behavior.
  • Collaborate with perception, planning, and platform teams to bring models from research to production

Job requirements

  • M.S. minimum, Ph.D. preferred in CS, EE, Mathematics, Statistics, or a related field.
  • 3+ years implementing and training models in a deep learning framework (PyTorch, TensorFlow, or JAX).
  • Direct, hands-on experience training vision-language / vision-language-action models.
  • Hands-on experience with model training, evaluation, and deployment in production.
  • Thorough understanding of state-of-the-art vision-language / VLA models, diffusion, flow matching, and transformers.
  • Experience with large-scale / distributed model training.
  • Model distillation, quantization, and inference optimization (ONNX/TensorRT, mixed precision, custom kernels).
  • SFT and RL post-training of large multimodal models.

Similar jobs