- Robot type
- Autonomous Vehicle
- Location
- Santa ClaraCaliforniaUSA
- Job type
- Software
- Posted
- Mar 17, 2026
- Salary
- $180,000–$240,000 a year
Senior Cloud Infrastructure Engineer
Job description
We have delivered complete, proprietary AV technology - an integration of software and hardware - to enable earlier successes for our clients in constrained Level 4 autonomy. By choosing the middle mile – with defined point-to-point delivery, we have simplified some of the more complex AV challenges, enabling us to achieve full autonomy ahead of competitors. Given extensive knowledge of Gatik’s well-defined, fixed route ODDs and hybrid architecture, we are able to hyper-optimize our models with exponentially less data, establish gate-keeping mechanisms to maintain explainability, and ensure continued safety of the system for unmanned operations.
Visit us at Gatik for more company…
Job responsibilities
- Cloud-Native Orchestration & Kubernetes Advanced K8s Management: Architect and maintain mission-critical Kubernetes clusters optimized for heavy GPU/TPU workloads.
- GPU Scheduling: Implement and optimize Kubernetes-native GPU scheduling (NVIDIA GPU Operator) to ensure maximum hardware utilization.
- Infrastructure as Code: Drive the "Everything as Code" philosophy using Terraform, Helm, and cloud-native tools.
- Self-Healing Infrastructure: Deploy Autonomous AI Agents (LangGraph, CrewAI) to monitor cluster health and enable automated triage of hardware failures and NCCL timeouts.
- Data Engineering & CI/CD Pipelines Autonomy Data Pipelines: Build large-scale pipelines using Apache Airflow, Kafka, and Spark to process raw sensor data into training-ready formats.
- GitOps: Implement robust GitOps workflows using ArgoCD, Gitlab CI/CD to automate the deployment of both infrastructure and model artifacts.
- Observability: Maintain deep visibility into infrastructure health and model serving performance using Prometheus, Grafana, and OpenTelemetry.
- Agentic DevOps & CI/CD: Develop agent-driven workflows to optimize the developer experience, such as automated PR reviewers for Terraform and AI agents that proactively suggest Kubernetes resource-limit adjustments…
Job requirements
- Experience: 5+ years in Cloud Infrastructure, DevOps, or MLOps supporting high-scale compute environments.
- Kubernetes Mastery: Deep expertise in K8s, Helm, and container orchestration.
- Orchestration & Tooling: Strong background in Apache Airflow, Argo Workflows, MLFlow, and Terraform.
- Distributed Systems: Practical experience supporting frameworks like Ray and PyTorch Distributed.
- Core Skills: Proficiency in Python, Bash scripting, and a solid understanding of IAM/RBAC.
- Distributed Training Expertise: Deep understanding of FSDP, and DeepSpeed.
- AI Agent Orchestration: Experience building Agentic Workflows (LangGraph, AutoGen) for infrastructure automation or data curation.
- Advanced Protocols: Familiarity with Model Context Protocol (MCP) to connect AI agents with infrastructure tools.
Similar jobs
Gatik · Software
Senior V&V Engineer – Software-in-the-Loop (SIL) & Simulation
Santa Clara, California
Gatik · Software
Embedded Software Engineer, Drive-By-Wire
Santa Clara, California
Gatik · Software
Senior/Staff GIS Engineer, HD Mapping
Santa Clara, California
Gatik · Software
Senior/Staff Full-Stack Engineer, Simulation Platform
Santa Clara, California
