- Robot type
- Autonomous Vehicle
- Location
- Santa ClaraCaliforniaUSA
- Job type
- Artificial Intelligence
- Posted
- Jul 31, 2026
- Salary
- $135,000–$200,000 a year
Software Engineer (SE / Sr SE), Data & ML Platform
Job description
All key offline workloads — large-scale data processing, simulation, auto-labeling, scenario mining, and model training — run on the compute platform this role owns. In this role, you will improve the reliability and efficiency of our Kubernetes infrastructure, make workload onboarding simpler and more self-service, and build reusable batch and workflow capabilities for petabyte-scale processing. We are looking for strong Kubernetes and platform-engineering fundamentals, depth in at least one adjacent area—distributed data processing, ML/GPU infrastructure, or multi-tenant compute systems—and the curiosity and ownership to grow across the others.
We are open to candidates at either the…
Job responsibilities
- Operate and evolve our production Kubernetes clusters end to end: bare-metal provisioning automation, highly available control planes, node lifecycle, GPU container runtime, networking, and storage
- Build safe, repeatable GitOps-based delivery for platform services and user applications using tools such as Argo CD, Helm, and Kustomize
- Develop shared multi-tenant platform capabilities for scheduling, resource isolation, storage, networking, access control, secrets, and observability while improving CPU/GPU utilization and cost efficiency
- Build and improve reusable distributed batch and workflow platforms for Spark data processing and GPU-based replay and simulation
- Ensure that your work is performed in accordance with the company’s Quality Management System (QMS) requirements and contribute to continuous improvement efforts
Job requirements
- BS, MS, or PhD in Computer Science or a related technical field, or equivalent practical experience
- Hands-on experience operating production Kubernetes clusters — node lifecycle, upgrades, troubleshooting — plus GitOps and infrastructure-as-code experience
- Experience with GPU or ML workload scheduling, queueing and priorities, fractional GPU sharing, autoscaling, or multi-tenant resource management
- Self-driven with a strong sense of ownership: a quick learner who is eager to take responsibility and drive projects forward end to end
- Experience with Ray or Kubeflow
- Experience with lakehouse technologies such as Delta Lake or Apache Iceberg
- Experience operating large-scale distributed data-processing and workflow systems, with hands-on depth in a system such as Apache Spark and working knowledge of Argo Workflows or an equivalent orchestrator
Similar jobs
PlusAI · Artificial Intelligence
Senior/Staff Research Engineer — Vision-Language-Action Models (Autonomous Driving)
Santa Clara, California
PlusAI · Artificial Intelligence
Senior/Staff Software Engineer (Machine Learning Runtime), Motion Planning
Santa Clara, California
PlusAI · Artificial Intelligence
Senior/Staff Machine Learning Engineer, Motion Planning
Santa Clara, California
PlusAI · Artificial Intelligence
Senior/Staff Machine Learning Engineer (Reinforcement Learning), Motion Planning
Santa Clara, California
