Skip to content

Jobs

Applied Intuition

Robot type
Autonomous Vehicle · Robot AI and Software
Location
SunnyvaleCaliforniaUSA
Job type
Artificial Intelligence
Posted
Sep 16, 2024
Salary
$126,000–$423,000 a year
Full-time

Research Engineer - AI/RL Infrastructure

Job description

Applied Intuition, Inc. is powering the future of physical AI. Founded in 2017 and now valued at $15 billion, the Silicon Valley company is creating the digital infrastructure needed to bring intelligence to every moving machine on the planet. Applied Intuition services the automotive, defense, trucking, construction, mining and agriculture industries in three core areas: tools and infrastructure, operating systems, and autonomy. Eighteen of the top 20 global automakers, as well as the United States military and its allies, trust the company’s solutions to deliver physical intelligence. Applied Intuition is headquartered in Sunnyvale, California, with offices in Washington, D.C.;

Job responsibilities

  • Design and build training and evaluation infrastructure to support our current AI research directions, orchestrating massive GPU clusters to process PBs of multimodal sensor data
  • Build robust benchmarking, continuous evaluation, and regression tracking systems to measure model performance across diverse, long-tail real-world driving distributions
  • Develop large-scale data sampling, dataset generation, and advanced data curation pipelines, leveraging state-of-the-art AI models to power a closed-loop data flywheel
  • Enable high-throughput distributed training across heterogeneous cloud environments, focusing on reliability, efficiency, and cost-aware scaling
  • Collaborate closely with AI research, autonomy, and platform teams to translate cutting-edge research into production-ready systems

Job requirements

  • Experience building and operating production-grade software systems across the full machine learning lifecycle, including training, evaluation, data, and deployment
  • Opinions about building a company-wide platform for ML training, evaluation, and deployment
  • Experience with performance engineering and compute acceleration for large-scale ML training, including profiling, bottleneck analysis, and optimization
  • Strong systems-level debugging skills to diagnose and resolve issues in large-scale distributed training, spanning model code, data pipelines, runtimes, and cluster infrastructure
  • Deep familiarity with the open-source ML and systems ecosystem, with judgment on when to adopt open source versus build in-house
  • Technical experience in: Pytorch, CUDA, Ray, Flyte, K8s

Similar jobs