- Robot type
- Autonomous Vehicle
- Location
- AustinTexasUSA
- Job type
- Artificial Intelligence
- Posted
- Apr 1, 2026
Full-time
Software Engineer – ML Platform
Job description
The ML Platform team at Avride builds the infrastructure that powers large-scale ML training and data processing for autonomous driving. We sit between Cloud Platform and ML engineers, turning low-level compute, storage, and networking primitives into an ML platform that teams actually use — scalable orchestration, distributed compute, and production-grade tooling for the full model lifecycle.
Job responsibilities
- Build and scale our ML compute platform on Kubernetes, using Argo Workflows for training, evaluation, and data processing orchestration
- Design and implement core platform capabilities, including a Ray-based internal SDK for distributed execution, and multi-tenant resource governance — scheduling, priorities, quotas, and policy enforcement across GPU,…
- Improve end-to-end training throughput and platform efficiency by optimizing data access patterns, caching, and removing bottlenecks in storage, network, and resource contention
- Work directly with ML teams to debug complex workload issues, drive root-cause analysis, and turn recurring problems into platform-level fixes
- Evaluate, integrate and extend open-source tooling (Argo Workflows, Ray, Kubernetes ecosystem) to meet evolving platform needs
- Strong proficiency in Python or Go; C++ is a plus
- Track record of designing and building scalable, maintainable systems and services
- Experience operating production services end-to-end: APIs, reliability practices, observability
Similar jobs
Avride · Artificial Intelligence
Lead AI Infrastructure Engineer
Austin, Texas
Avride · Artificial Intelligence
Lead Data Scientist – Autonomous Driving
Austin, Texas
Avride · Artificial Intelligence
Machine Learning Engineer – Motion Planning & Prediction
Austin, Texas
Avride · Artificial Intelligence
Machine Learning Engineer
Austin, Texas
