- Robot type
- Drone · Defense
- Location
- Costa MesaCaliforniaUSA
- Job type
- Software
- Posted
- Sep 3, 2026
- Salary
- $166,000–$220,000 a year
Senior Infrastructure Reliability Engineer
Job description
Infrastructure Reliability Engineering (IRE) is a small but growing team responsible for the infrastructure and operations behind the core developer tools and on-prem compute platforms used across the entire engineering organization. We own the services every engineer depends on daily — source control, CI/CD, and artifact management — as well as the on-prem infrastructure behind simulation and GPU workloads. As the company’s on-prem footprint grows, this team is expanding its scope to provide SRE capabilities for on-prem systems, so there’s an opportunity to help shape that practice from the ground up.
You’ll own the full lifecycle — patching, upgrades, backups, scaling, and incident…
Job responsibilities
- Serve as a primary owner for critical services, including on-call and knowledge-sharing across the team
- Own the lifecycle of core self-hosted developer tools (e.g., RunAI, GitHub Enterprise Server, CircleCI, JFrog Artifactory/Xray)
- Design and implement automated systems for patching, backups (with validation), and upgrades
- Scale infrastructure to support a fast-growing engineering org
- Use Infrastructure-as-Code (Terraform) to manage environments
- Operate and troubleshoot systems using Docker, Kubernetes, and cloud platforms (AWS, GCP, Azure)
- Define and maintain SLOs for service availability, reliability, and performance
- Build and maintain monitoring, alerting, and observability for developer tool services
Job requirements
- Experience operating infrastructure outside of managed cloud services — bare-metal kubernetes and on-prem virtualization (VMware ESXi/vSphere)
- Experience operating production systems using Docker and Kubernetes
- Strong foundational knowledge of Linux (RHEL , Ubuntu)
- Proficiency with at least one cloud platform (AWS, GCP, or Azure)
- Experience managing infrastructure with Infrastructure-as-Code tools (e.g., Terraform/OpenTofu)
- Experience with configuration management tooling (e.g., Ansible, Puppet, Chef)
- Strong problem-solving skills with a focus on automation
- Scripting or software development experience (e.g., Python, Go, Bash)
Similar jobs
Anduril · Software
Staff Software Engineer, Discovery
Boston, Massachusetts · Costa Mesa, California
Anduril · Software
Senior Forward Deployed Software Engineer, Strategic Defense
Costa Mesa, California · Seattle, Washington
Anduril · Software
Technical Recruiter, Software - Air Dominance and Strike (Contract)
Boston, Massachusetts · Costa Mesa, California
Anduril · Software
Technical Recruiter, Software - Air Dominance and Strike (Contract)
Seattle, Washington · Costa Mesa, California
