Skip to content

Jobs

Zipline

Robot type
Drone
Location
South San FranciscoCaliforniaUSA
Job type
Software
Posted
Aug 12, 2026
Full-time

Sr. SWE Datacenter Automation

Job description

Our customers include the world’s largest and most prominent healthcare systems, governments, retailers, restaurants and global businesses who rely on us to save lives, reduce emissions, increase economic opportunity, and provide delivery from point A to point B as fast as possible. The drone is only 15% of what we’ve built to enable seamless, reliable, global operations.

Our system strengthens supply chains, reduces congestion, and gives people time back. With more than 140 million commercial autonomous miles safely flown, Zipline is redefining access to healthcare, consumer products, and food across the globe.

We operate at a global scale and are looking for practical problem solvers who…

Job responsibilities

  • Own end‑to‑end lifecycle for datacenter compute and storage: bare‑metal provisioning, hypervisor management, SAN/NVMe storage clusters, network configuration, and Kubernetes cluster lifecycle.
  • Design, build, and operate automation that reduces manual setup time and increases deployment velocity: PXE/firmware workflows, dynamic inventory, image generation, fleet-wide configuration drift detection, and…
  • Deliver measurable reliability and scale improvements: set SLIs/SLOs for provisioning time, node commissioning success rate, cluster upgrade success rate, and mean time to recover (MTTR); own meeting those targets.
  • Lead cross‑functional runbook and incident ownership for infra incidents affecting flight operations or telemetry: on‑call rotation, incident commander for datacenter platform incidents, postmortems and action items.
  • Instrument and maintain monitoring, alerting, and dashboards for hardware health, hypervisor performance, storage latency, Kubernetes control plane health, and cluster autoscaling behavior.
  • Implement cost, capacity, and lifecycle management: capacity planning for compute/storage, automated reclamation, firmware/BIOS/hypervisor patch pipelines, and cold‑standby / failover procedures for critical systems.
  • Execute hands‑on tasks when required: racking and cabling in datacenters, troubleshooting hardware failures, capture forensic logs, and coordinate physical repairs with vendors and field ops.

Similar jobs