- Robot type
- Autonomous Vehicle
- Location
- Santa ClaraCaliforniaUSA
- Job type
- Artificial Intelligence
- Posted
- Mar 17, 2026
- Salary
- $180,000–$240,000 a year
Senior AI Infrastructure Engineer
Job description
We have delivered complete, proprietary AV technology - an integration of software and hardware - to enable earlier successes for our clients in constrained Level 4 autonomy. By choosing the middle mile – with defined point-to-point delivery, we have simplified some of the more complex AV challenges, enabling us to achieve full autonomy ahead of competitors. Given extensive knowledge of Gatik’s well-defined, fixed route ODDs and hybrid architecture, we are able to hyper-optimize our models with exponentially less data, establish gate-keeping mechanisms to maintain explainability, and ensure continued safety of the system for unmanned operations.
Visit us at Gatik for more company…
Job responsibilities
- Distributed Training & ML Systems Support Scale Research Workloads: Enable researchers to scale complex models (VLA, World Models) across multi-node setups using PyTorch Distributed, and Ray Train.
- Performance Optimization: Architect and optimize multi-GPU setups, ensuring efficient model parallelism and data parallelism techniques across H100/A100 clusters.
- Networking & Hardware Tuning: Optimize low-level communication (e.g., NCCL tuning, InfiniBand, or RoCE v2) to minimize latency for 3D Gaussian Splatting (3DGS) and large-scale training.
- Intelligent Resource Scheduling: Optimize hardware utilization and cost-efficiency through Kubernetes-native GPU scheduling (NVIDIA GPU Operator, KubeFlow).
- Inference Performance Engineering: Deploy and scale optimized model artifacts using TensorRT, ONNX Runtime, and Triton Inference Server, fine-tuning pipelines for both real-time and batch processing
- Agentic Infrastructure & Automation Self-Healing AI Infrastructure: Architect and deploy Autonomous AI Agents (LangGraph, CrewAI, or AutoGen) to monitor GPU cluster health, enabling automated real-time triage of…
- Agentic DevOps & CI/CD: Develop agent-driven automation, such as Agentic PR Reviewers for infrastructure code and AI agents that proactively suggest model-specific Kubernetes resource optimizations.
- Agentic Data Curation: Support researchers in building "Data Machines" where AI agents autonomously curate, label, and verify high-priority edge cases from raw data.
Job requirements
- Experience: 5+ years in ML infrastructure, MLOps, or DevOps supporting high-scale compute environments.
- ML Expertise: Deep understanding of multi-GPU training strategies (FSDP, DeepSpeed, Ray Train) and high-performance networking (NCCL, InfiniBand).
- Infrastructure Automation: Mastery of Kubernetes, Terraform, and Helm, with a focus on GPU-native orchestration.
- AI Agent Frameworks: Proven experience building or supporting Agentic Workflows for infrastructure or data automation (e.g., using LLMs to drive DevOps tasks).
- Platform Mastery: Expertise in MLFlow, Argo Workflows, and Kubernetes.
- Containerization: Strong experience with Docker, Kubernetes, and Helm.
- Data & CI/CD: Proficiency in Apache Airflow, Kafka, Spark, and GitOps automation.
- Core Skills: Proficiency in Python and Bash; experience with Go or Rust is a plus
Similar jobs
Gatik · Artificial Intelligence
Machine Learning Engineer
Santa Clara, California
Gatik · Artificial Intelligence
Senior Systems & Safety Engineer – AI/ML
Santa Clara, California
Gatik · Artificial Intelligence
Senior Research Engineer, Perception
Mountain View, California · Santa Clara, California
Gatik · Artificial Intelligence
Senior Software Engineer, Perception
Santa Clara, California
