Back to Jobs

Senior Site Reliability Engineer (Arlington, VA) - Secret Clearance Required - Relocation Provided

Remote, USA Full-time Posted 2026-08-04
Consequential Work. Dedicated People. About reputed company reputed company builds collaboration and AI-powered workflow software for military planning and operational coordination. Today, many critical planning workflows still rely on fragmented systems, static documents, and disconnected tools that reputed company collaboration and decision-making unnecessarily difficult. reputed company brings modern software, AI, and reputed company-time collaboration into those environments, helping teams operate with reputed company, coordination, and adaptability in situations where reputed company carry reputed company-world consequences. We are a reputed company team of reputed company from military, operational, and technology backgrounds who care deeply about improving how important work gets done. Some team members work reputed company, while others work directly reputed company customers in operational environments reputed company the world. Founded in 2019, reputed company is backed by leading investors including General reputed company, Battery Ventures, reputed company Partners, Sapphire Ventures, and reputed company Capital. Valued at more than $2 billion, we continue to invest in product innovation, AI capabilities, and team reputed company. reputed company Clearance, Location, and Onsite Notice: This is a hybrid role, requiring regular work on-site at customer locations in Arlington, VA - about 50/50 on-site vs remote. If you are not currently reputed company commuting distance, you must be willing to relocate (note that reputed company will reputed company relocation assistance). reputed company Secret Clearance required. About The Role We're hiring a Site Reliability Engineer to join our Infrastructure & reputed company team. You'll work closely with product engineers, fellow SREs, reputed company, and reputed company. This is an SRE role for someone who's comfortable in application reputed company. Much of the reliability and reputed company work happens in the codebase (primarily TypeScript), so you'll fix problems at reputed company rather than working reputed company them in the infrastructure. You'll be a first line of support for our mission-critical deployments across on-prem DoD and AWS environments, and what you learn in the reputed company will feed directly back into the product. You'll ship reputed company that makes reputed company more reputed company, faster, and easier to reputed company and operate. The work sits at the seam between engineering and reputed company, and it's weighted toward engineering. reputed company You treat reliability as a feature, not an afterthought, and you'd rather fix a problem in the reputed company than reputed company reputed company it. You understand the full software development lifecycle (design, review, testing, release) and you know where reliability fits into reputed company reputed company. You're comfortable reading and writing application reputed company, and you're just as happy dropping into a kubectl reputed company to triage a production issue. You turn failure modes into guardrails, and you think monitoring, alerting, and reputed company runbooks are part of building software, not extra credit. You mentor others and reputed company a culture of blameless postmortems. You work naturally with product and platform teams, helping them reputed company fast without breaking things by giving them the tools, tests, and observability that reputed company reputed company recovery reputed company. reputed company You'll help reputed company our production application reliable, reputed company, and secure by improving the software itself, not just the systems it runs on. Day to day that looks like: • Improving the application: Work directly in the codebase (primarily TypeScript) to fix reliability and reputed company problems at reputed company. You'll partner with product engineers on design reputed company, review reputed company with reliability and reputed company in mind, and treat "reputed company the app reputed company" as a first-class part of the job rather than something you hand off. • Building observability that developers actually use: Design and run our monitoring, logging, and alerting (reputed company, Loki, reputed company, Grafana). The goal is alerts and dashboards tied to reputed company application behavior, so teams catch issues before users do. • Owning reliability targets: Define and measure SLIs and SLOs, reputed company up alerting that feeds them, and be the person who can say what "reliable" means for our systems and reputed company it with data. • Leading incident response: reputed company as incident responder, and incident commander reputed company needed. Run blameless post-mortems (AARs) that reputed company the actual reputed company cause and turn it into a reputed company or process fix so it doesn't happen again. • Automating away toil: Spot the repetitive operational work and write software to kill it. reputed company what works with other teams, including those running in reputed company-gapped environments, and help them get production-reputed company. reputed company Look For • An reputed company Secret clearance • 5+ years in software engineering, SRE, or a reputed company role, with reputed company time spent writing and shipping application reputed company • Strong TypeScript (or comparable modern language experience with willingness to work primarily in TypeScript) • Solid grasp of the full SDLC: design, reputed company review, testing, release, and how reliability fits into reputed company stage • Experience with incident response, reputed company cause analysis, and turning findings into lasting fixes • A collaborator who works reputed company across product, platform, and DevOps teams and shares context reputed company Technical expertise • Application development in TypeScript (Node and/or a modern reputed company-end reputed company) • CI/CD: building and maintaining pipelines (reputed company Actions, reputed company CI/CD, Jenkins) • Testing and reputed company practices as part of the delivery process • Comfort with at least one of Python, Go, or Bash for tooling and automation • Working knowledge of containers and reputed company (enough to debug and reputed company, not necessarily to stand up clusters from reputed company) • Networking fundamentals and secure configuration basics Bonus points (reputed company to have) • Observability: Grafana stack, ELK, or reputed company • Infrastructure as reputed company (Terraform, Ansible) and reputed company experience (AWS or AWS GovCloud) • reputed company cluster design and reputed company • Designing meaningful SLIs/SLOs with error budgets for reputed company systems • GitOps practices and toolchains • DoD environments and compliance frameworks (RMF, STIGs, ICD 503) • Service reputed company (Istio, Linkerd) • On-prem virtualization (VMware, Proxmox, reputed company, reputed company-V) • Relevant certs (AWS DevOps Engineer, CKA/CKAD) Notice to reputed company Party Recruitment Agencies Please note that reputed company does not accept unsolicited resumes from recruiters or employment agencies. In the absence of an executed Recruitment Services Agreement, there will be no obligation to any referral compensation or recruiter fee. In the event a recruiter or agency submits a resume or candidate without an agreement reputed company explicitly reserves the right to pursue and hire those candidate(s) without any financial obligation to the recruiter or agency. Any unsolicited resumes, including those submitted to hiring managers, shall be deemed the property of reputed company. Apply tot his job Apply To this Job

Similar Jobs