- Robot type
- Defense
- Location
- El SegundoCaliforniaUSA
- Job type
- Software
- Posted
- Jul 30, 2026
- Salary
- $170,000–$195,000 a year
Full-time
Site Reliability Engineer
Job description
Picogrid is a leading venture-backed defense technology company founded to bridge the decades-long gap between modern technology and the critical demands of national security. Today, we're building the essential infrastructure to unify sensors, autonomy, and operators with our technology deployed in active operations around the world. Our mission is to deliver an operational advantage to secure the United States and its allies.
Job responsibilities
- Own, define and drive our reliability SLIs and SLOs for cloud deployments
- Own, define and drive our reliability SLIs and SLOs for our edge devices deployed in remote and sometimes contested areas
- Own the observability stack: Grafana, Prometheus, Loki, and OpenTelemetry, with dashboards versioned in git and alerting rules checked in alongside the code they watch
- Participate in on-call and incident response: log-first troubleshooting, blameless postmortems, and follow-up hardening
- Encode reliability into infrastructure as code
Job requirements
- 3+ years of experience as an SRE or related roles
- Deep Kubernetes operations experience: node lifecycle, workload scheduling, StatefulSets, graceful drains, and live cluster debugging
- Experience designing comprehensive observability dashboards and high signal-to-noise ratio alerting rules
- You are a competent and experienced incident responder practicing methodical evidence-first triage, blameless postmortems, and turning incidents into durable guardrails
- Production Terraform or OpenTofu experience
- Fluent in AWS including IAM, networking, multi-account environments, and account and workload hardening
- Experience managing high availability database deployments
- IoT or edge fleet operation experience
