Skip to content

Jobs

Anduril

Robot type
Drone · Defense
Location
Costa MesaCaliforniaUSA
Job type
Software
Posted
Mar 11, 2026
Salary
$166,000–$220,000 a year
Full-time

Site Reliability Engineer, Space

Job description

We are looking for a Site Reliability Engineer to join our Space team in Costa Mesa, CA someone who brings deep technical credibility and the leadership presence to match. This is not a management role, but it is a senior one: you'll be a technical anchor on a tight-knit team, the kind of engineer others look to when something breaks, when architecture decisions need to be made, and when the room needs someone who can cut through noise and drive clarity.

You will operate, maintain, harden, and evolve a classified networking capability that delivers operational Space Domain Awareness data for the United States Space Force.

You'll work hand in hand with Mission Software Engineers, Product…

Job responsibilities

  • Serve as a senior technical subject matter expert on a growing Product Operations team supporting a live, operational DoD system, helping set the standard for how the team operates, communicates, and solves problems
  • Stand up, harden, and maintain classified Linux-based infrastructure across globally distributed sites in air-gapped environments
  • Apply and enforce STIG compliance and system hardening standards across all deployed environments this is not checkbox work, you will own the security posture
  • Build, maintain, and evolve infrastructure as code using Ansible/Puppet automating provisioning, configuration management, and deployment pipelines
  • Write and maintain operational tooling, automation scripts, and workflow optimizations in Bash and Python
  • Architect and maintain comprehensive observability across the stack using Splunk/ELK/Open Search logging, alerting, and tracing for distributed systems
  • Manage and operate containerized workloads in Kubernetes at scale in production, classified environments
  • Provide real-time incident response, rapid root-cause analysis, and resolution of complex issues spanning application code and backing infrastructure

Job requirements

  • Active U.S. Secret security clearance (non-negotiable must currently hold and be able to maintain- preferred TS/SCI clearance)
  • 5+ years of hands-on experience with Kubernetes in production environments deploying, scaling, debugging, and managing containerized workloads. This is core to the role, not adjacent.
  • Expert-level Linux proficiency you live in the terminal, you understand the OS deeply (networking, filesystems, systemd, SELinux, kernel tuning) but this is an engineering and operations role, not a sysadmin role.
  • Proven experience with STIG compliance and system hardening you've applied STIGs, remediated findings, and understand the "why" behind the controls, not just the checklist
  • Strong experience standing up infrastructure in air-gapped/classified environments you understand the unique constraints and have operated in them
  • Proficiency with Ansible Puppet and Terraform for building, managing, and versioning infrastructure as code at scale
  • Strong scripting ability in Bash and Python writing production-grade automation, tooling, and operational scripts
  • Hands-on experience with observability tooling Prometheus, Grafana, and ELK Stack (Elasticsearch, Logstash, Kibana) for monitoring, logging, alerting, and tracing distributed systems

Similar jobs