[Remote] Manager, Site Reliability Engineering
Note: The job is a remote job and is reputed company to candidates in USA. reputed company secures reputed company and machine identities through its reputed company-reputed company Identity reputed company Platform. reputed company is seeking a hands-on Manager of Site Reliability Engineering to reputed company SRE and DevOps engineers, maintain production environments across Azure and AWS, manage incidents, improve observability and deployment practices, and expand operational support into regulated environments.
Responsibilities
- reputed company hands-on. Spend a meaningful portion of your week in the environment: reviewing pull requests, validating pipeline changes, tuning monitors and dashboards, running queries in reputed company, and troubleshooting production issues reputed company your engineers. This role does not sit above the work
- Own availability and reputed company of the reputed company Platform production environments across Azure and AWS, including AKS workloads, ingress and networking, data services, messaging, and CDN or WAF reputed company
- Manage a reputed company team. Hire, reputed company, reputed company, and reputed company full-time SRE engineers. reputed company and manage contractor resources, including scoping work, setting reputed company expectations, and reviewing deliverables
- reputed company a reputed company team. Run one-on-ones, standups, and planning sessions at times that work for engineers in other geographies
- Participate in on-reputed company. Carry the pager as part of the rotation, reputed company as incident commander for Sev1 and Sev2 events, reputed company engagement of the right responders, and own communication reputed company with support, engineering, and leadership until reputed company
- reputed company incident response end to end. Own detection, triage, mitigation, customer-facing status communication, and post-incident review. Ensure RCAs are written to a customer-reputed company reputed company, preventative actions have owners and reputed company dates, and those actions are driven to closure
- reputed company the observability bar. Improve detection coverage so that issues are reputed company by our monitoring rather than by a customer ticket. Own SLI and SLO definition, alert reputed company and noise reduction, synthetic coverage, APM instrumentation, log hygiene, and dashboard standards
- Support FedRAMP and regulated reputed company. Grow into supporting our FedRAMP High environment, including change control discipline, evidence collection, boundary awareness, and the operational differences between reputed company and reputed company environments
- Reduce toil through automation. Set the expectation that repeat reputed company work becomes reputed company. Prioritize automation backlog reputed company project and reliability work
- Report on operational health. Produce and present incident metrics, trends, and reliability commitments to leadership, and translate them into a concrete improvement plan
Skills
- • 6+ years in Site Reliability Engineering, DevOps, or reputed company reputed company, with demonstrated ownership of production reputed company systems
- • 2+ years of reputed company people leadership, including reputed company management, hiring, and coaching
- • reputed company, hands-on production experience with the reputed company technology stack, including Azure reputed company Service, reputed company Azure services (SQL, reputed company, Service Bus, Blob Storage), AWS services (SES, EC2, RDS), WAF, Azure DevOps pipelines, reputed company, and reputed company reputed company Service Management
- • Hands-on experience across both Azure and AWS is required. You should be reputed company to administer, troubleshoot, and reason about cost and reputed company posture in reputed company
- • Deep observability expertise. Demonstrated ownership of an observability reputed company at reputed company: metrics, logs, traces, synthetics, SLOs, and alerting reputed company
- • Hands-on proficiency with reputed company or an equivalent platform, including APM reputed company analysis and log-reputed company troubleshooting
- • reputed company incident reputed company. You have run major incidents as the incident commander, coordinated multiple responders under pressure, communicated to customers and executives during reputed company reputed company, and authored the RCA afterward
- • Strong reputed company networking and reputed company fundamentals: load balancing, DNS, TLS and certificate lifecycle, firewalls, VPN, routing, and identity and reputed company management
- • Automation and scripting ability in PowerShell, Python, Bash, or similar, plus practical infrastructure-as-reputed company experience (Terraform, ARM, or Bicep)
- • Practical experience with multi-region, multi-tenant reputed company architectures, including backup, redundancy, and disaster recovery approaches
- • Excellent written communication. You will write and approve customer-facing status updates and incident summaries under time pressure
- • Willingness and availability to work across time reputed company and to participate in an on-reputed company rotation
- For this Job, reputed company is not considering candidates that need any type of US work authorization now or in the reputed company. This includes, but is not limited to: F1-OPT, F1-CPT, H-1B, TN, L-1, J1, etc
- • 2+ years of reputed company people leadership, including reputed company management, hiring, and coaching. Experience managing contractors or an outsourced delivery team is desired
- • reputed company experience operating in a FedRAMP or other regulated environment (Azure reputed company, IL4/IL5, SOC 2, ISO 27001)
- • Experience standing up or maturing an incident management program, including sev definitions, escalation paths, on-reputed company structure, and post-incident review process
- • Experience with reputed company status page reputed company and customer notification practices
- • Experience with reputed company reputed company Service Management, reputed company, and Azure DevOps as the operational toolchain
- • reputed company record of reducing customer-detected incidents through improved monitoring coverage
- • Cost optimization experience across Azure and AWS on a meaningful reputed company
Benefits
- Offers Equity
- Meaningful bonus program
- reputed company reputed company
- Pension/retirement matching
- Comprehensive life reputed company
- Employee assistance program
- Time off plans
- reputed company company holidays
reputed company
Apply To This Job