Skip to content

Jobs

Anduril

Robot type
Drone · Defense
Location
ArlingtonVirginiaUSA
Job type
Software
Posted
Sep 14, 2026
Salary
$146,000–$220,000 a year
Full-time

Site Reliability Engineer, Discovery

Job description

As a Site Reliability Engineer in Anduril Cyber, you will solve a wide variety of problems involving networking, systems integration, distributed systems, and more, while making pragmatic engineering tradeoffs along the way. Your efforts will ensure that Anduril’s software is reliable, scalable, and deployable, in order to achieve critical national security outcomes. You will work closely with software developers, customers, and external vendors to get working offensive Cyber products into the hands of customers.

You will also be the steward of Anduril's mission and technological advantage in the room with customers — attending technical meetings, explaining why the system behaves the way…

Job responsibilities

  • Own the health of our deployed systems and keep them running with minimal downtime.
  • Automate and improve our software deployment processes into air-gapped, TS/SCI environments.
  • Design, build, and maintain the CI/CD and automated test infrastructure for Cyber’s complex hardware and software systems.
  • Develop metrics dashboards, TUIs, scripts, and other tools that automate common deployment steps or help to debug our software stack.
  • Drive engineering requirements based on onsite observations.
  • Perform root cause analysis and diagnose issues in mission-critical systems across our software stack, the Lattice OS stack, and external vendor services.
  • Build strong relationships with internal and external customers to identify technical solutions to their problems.
  • Drive continuous improvement by instrumenting systems, analyzing failures, and leading post-mortem events that span software, firmware, and hardware.

Job requirements

  • Currently possesses and is able to maintain an active U.S. TS/SCI security clearance.
  • Based in the DC metro area to support 3-5 days per week working on site at customer facilities.
  • 4+ years of experience in a Sys Admin, Site Reliability, DevOps, or Software Engineering role.
  • Deep, practical experience with Linux and Kubernetes (or a similar container orchestrator).
  • Working knowledge of network fundamentals and the ability to debug connectivity in a locked-down environment.
  • Experience delivering and maintaining systems on air-gapped and security-hardened networks.
  • Strong proficiency in Python or Bash for automation and debugging, and the ability to read and debug service code in a compiled language such as Go.
  • Excellent written and verbal communication skills for collaborating with a cross-functional engineering team and external customers.

Similar jobs