Skip to content

Jobs

Toyota Research Institute

Robot type
Robot AI and Software
Location
Los AltosCaliforniaUSA
Job type
Artificial Intelligence
Posted
Jun 30, 2026
Full-time

Senior Research Engineer, Computer Vision (LFV/WFM)

Job description

At Toyota Research Institute (TRI), we’re on a mission to improve the quality of human life. We’re developing new tools and capabilities to amplify the human experience. To lead this transformative shift in mobility, we’ve built a world-class team advancing the state of the art in AI, robotics, driving, and material sciences.

The Learning From Videos (LFV) team develops world foundation models that leverage large-scale multi-modal data (RGB, depth, flow, semantics, actions, tactile, audio, etc.) from multiple domains to power downstream embodied AI tasks.

Our team is looking for a Research Engineer to help develop and deploy our world foundation models (WFMs) toward their key milestones in…

Job responsibilities

  • Collaborate directly with research scientists to implement, iterate on, and evaluate new architectures, objectives, datasets, and training strategies.
  • Build and maintain scalable pipelines for ingesting, converting, validating, and serving heterogeneous datasets (multi-view, multi-modal, multi-embodiment, etc.), across robotics and autonomous driving, into unified…
  • Support and optimize large-scale distributed training of world foundation models on multi-GPU and multi-node clusters.
  • Develop tools for dataset inspection, experiment tracking, model evaluation, GPU resource management, and visualization. Automate repetitive workflows to improve team velocity.
  • Work with other TRI teams and Toyota affiliates to set up shared pipelines, onboard their data, and support joint training and evaluation efforts.
  • Produce maintainable, well-documented code. Contribute to internal tooling and open-source releases to the scientific community.

Job requirements

  • Master’s or PhD in Computer Science, Electrical Engineering, Machine Learning, or a related field, with a minimum of 3 years of relevant experience and strong software engineering skills.
  • Deep proficiency in Python, PyTorch, and the Unix/Linux toolchain. Comfort working in terminal-heavy, SSH-based workflows on shared GPU clusters.
  • Hands-on experience with large-scale deep learning training, including distributed training (DDP, FSDP, DeepSpeed, or similar), GPU profiling, and debugging training failures at scale.
  • Experience building data pipelines for heterogeneous or multi-modal datasets (images, video, depth, point clouds, actions, etc).
  • Experience with video diffusion models, 3D/4D reconstruction, and multi-view geometry.
  • You are proactive, self-directed, and comfortable operating with ambiguity in a research-driven environment that spans multiple divisions.
  • You are a reliable teammate who communicates clearly and takes ownership of problems end-to-end.
  • Experience with cloud training infrastructure (AWS SageMaker, EC2) and containerized workflows (Docker, Kubernetes).

Similar jobs