Skip to content

Jobs

Applied Intuition

Robot type
Autonomous Vehicle · Robot AI and Software
Location
SunnyvaleCaliforniaUSA
Job type
Artificial Intelligence
Posted
Aug 12, 2026
Salary
$215,000–$285,000 a year
Full-time

AI Performance Engineer

Job description

Applied Intuition, Inc. is powering the future of physical AI. Founded in 2017 and now valued at $15 billion, the Silicon Valley company is creating the digital infrastructure needed to bring intelligence to every moving machine on the planet. Applied Intuition services the automotive, defense, trucking, construction, mining and agriculture industries in three core areas: tools and infrastructure, operating systems, and autonomy. Eighteen of the top 20 global automakers, as well as the United States military and its allies, trust the company’s solutions to deliver physical intelligence. Applied Intuition is headquartered in Sunnyvale, California, with offices in Washington, D.C.;

Job responsibilities

  • Profile and optimize distributed training end to end - data loading and preprocessing, augmentation, kernel execution, gradient communication, and checkpointing
  • Optimize large-scale offline and batch inference over petabyte-scale sensor logs: batching and scheduling strategies, quantization and low-precision execution, graph optimization, and accelerator saturation across…
  • Establish roofline and performance models for our workloads, quantify the gap between achieved and theoretical performance, and stack-rank optimization opportunities by impact and effort
  • Improve multi-node scaling efficiency: sharding and parallelism strategies, collective communication, interconnect utilization, and memory-bandwidth and kernel-fusion bottlenecks
  • Drive cluster goodput - reduce GPU idle time from input pipeline stalls, storage and network I/O, scheduling gaps, stragglers, and failure recovery on long-running jobs
  • Build the benchmarking, observability, and regression-detection tooling that keeps performance from silently degrading as models and code evolve
  • Collaborate with engineers across functions to solve complex data and compute problems at scale
  • Contribute to a team culture that values effective collaboration, technical excellence, and innovation

Job requirements

  • Hands-on ML performance engineering experience: profiling, roofline analysis, throughput optimization, and root-cause investigation in production systems
  • Experience with distributed multi-node training at scale (FSDP, DeepSpeed, Megatron, NCCL, or equivalent), including diagnosing scaling inefficiency as node count grows
  • Deep familiarity with GPU or accelerator performance concepts - memory bandwidth, kernel launch overhead, occupancy, quantization, collective communication
  • Experience with high-throughput or batch inference systems (NVIDIA Triton Inference Server, TensorRT, ONNX Runtime, Ray, or similar)
  • Fluency in Python and proficiency in C++ or another systems language
  • Excellent debugging, analytical, and problem-solving skills
  • A deep understanding of machine learning foundations, and the ability to develop technical solutions for problems with no established playbook

Similar jobs