Skip to content

Jobs

Gatik

Robot type
Autonomous Vehicle
Location
Santa ClaraCaliforniaUSA
Job type
Software
Posted
Mar 17, 2026
Salary
$180,000–$240,000 a year
Full-time

Senior Cloud Infrastructure Engineer

Job description

We have delivered complete, proprietary AV technology - an integration of software and hardware - to enable earlier successes for our clients in constrained Level 4 autonomy. By choosing the middle mile – with defined point-to-point delivery, we have simplified some of the more complex AV challenges, enabling us to achieve full autonomy ahead of competitors. Given extensive knowledge of Gatik’s well-defined, fixed route ODDs and hybrid architecture, we are able to hyper-optimize our models with exponentially less data, establish gate-keeping mechanisms to maintain explainability, and ensure continued safety of the system for unmanned operations.

Visit us at Gatik for more company…

Job responsibilities

  • Cloud-Native Orchestration & Kubernetes Advanced K8s Management: Architect and maintain mission-critical Kubernetes clusters optimized for heavy GPU/TPU workloads.
  • GPU Scheduling: Implement and optimize Kubernetes-native GPU scheduling (NVIDIA GPU Operator) to ensure maximum hardware utilization.
  • Infrastructure as Code: Drive the "Everything as Code" philosophy using Terraform, Helm, and cloud-native tools.
  • Self-Healing Infrastructure: Deploy Autonomous AI Agents (LangGraph, CrewAI) to monitor cluster health and enable automated triage of hardware failures and NCCL timeouts.
  • Data Engineering & CI/CD Pipelines Autonomy Data Pipelines: Build large-scale pipelines using Apache Airflow, Kafka, and Spark to process raw sensor data into training-ready formats.
  • GitOps: Implement robust GitOps workflows using ArgoCD, Gitlab CI/CD to automate the deployment of both infrastructure and model artifacts.
  • Observability: Maintain deep visibility into infrastructure health and model serving performance using Prometheus, Grafana, and OpenTelemetry.
  • Agentic DevOps & CI/CD: Develop agent-driven workflows to optimize the developer experience, such as automated PR reviewers for Terraform and AI agents that proactively suggest Kubernetes resource-limit adjustments…

Job requirements

  • Experience: 5+ years in Cloud Infrastructure, DevOps, or MLOps supporting high-scale compute environments.
  • Kubernetes Mastery: Deep expertise in K8s, Helm, and container orchestration.
  • Orchestration & Tooling: Strong background in Apache Airflow, Argo Workflows, MLFlow, and Terraform.
  • Distributed Systems: Practical experience supporting frameworks like Ray and PyTorch Distributed.
  • Core Skills: Proficiency in Python, Bash scripting, and a solid understanding of IAM/RBAC.
  • Distributed Training Expertise: Deep understanding of FSDP, and DeepSpeed.
  • AI Agent Orchestration: Experience building Agentic Workflows (LangGraph, AutoGen) for infrastructure automation or data curation.
  • Advanced Protocols: Familiarity with Model Context Protocol (MCP) to connect AI agents with infrastructure tools.

Similar jobs