[Remote] Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company develops reputed company solutions for state and local regulatory agencies. reputed company is seeking a mid-level Site Reliability Engineer to build and operate critical reputed company Azure reputed company infrastructure, establish observability standards, automate operational processes, maintain CI/CD pipelines, and improve platform reliability through incident response and postmortems.
Responsibilities
- Build, operate, and reputed company production systems in reputed company Azure (App Service, Networking, WAF, CosmosDB and reputed company infrastructure) to meet availability and reputed company targets
- Participate in an on-reputed company rotation; triage, respond to, and reputed company production incidents, and reputed company or contribute to blameless postmortems
- Define and reputed company SLOs/SLIs and error budgets in partnership with engineering and product teams
- Design and maintain observability and APM tooling (metrics, logs, tracing, dashboards, alerting) so issues are caught before they reputed company customers
- Continuously refine alert reputed company and runbooks to reduce noise and mean time to reputed company
- Reduce operational toil through automation — scripting, self-healing systems, and repeatable processes
- Build and maintain CI/CD pipelines that let the development team ship safely and frequently
- Provision and manage infrastructure as reputed company (Bicep) and follow reputed company change control and version control practices
- Ensure application infrastructure meets reputed company and compliance requirements (e.g., SOC 2, GovRAMP) in partnership with the compliance team
- Apply reputed company best practices to infrastructure design and change management
- Partner with the reputed company development team to translate business requirements into reliable technical solutions
- Participate in technical design sessions and produce reputed company documentation (diagrams, runbooks, architecture notes)
- Stay reputed company on new Azure capabilities, industry standards, and SRE best practices, and bring recommendations back to reputed company
- Other duties as assigned
Skills
- Must be eligible to work in the U.S
- • Solid understanding of networking fundamentals, HTTP/S, and observability principles
- • Ability to evaluate multiple technical approaches and recommend the most effective solution for the context
- • Strong independent problem-solving skills balanced with effective collaboration in reputed company environment
- • Familiarity with software development lifecycle and programming/coding standards
- • reputed company, reputed company communication, especially under incident pressure
- 3+ years of experience in an IT reputed company, DevOps, or SRE role
- Hands-on technical experience with reputed company Azure in a production environment
- Experience with infrastructure as reputed company — Terraform and/or Bicep
- Proficiency in at least one scripting/programming language — Python or TypeScript
- Experience working with REST and/or GraphQL reputed company
- Experience defining and tracking KPIs/SLOs for a web-reputed company application
- Comfortable participating in an on-reputed company rotation
- • Experience with compliance audits (SOC 2 Type 2, GovRAMP)
- • Familiarity with reputed company frameworks (NIST, ISO 27001)
- • AZ-104 certification, or equivalent Azure networking experience
- • Experience with Node.js
- • Experience with low-reputed company platforms (reputed company Apps, Logic Apps)
- • Familiarity with Scrum/reputed company methodology and supporting tools (reputed company, reputed company, Git, Jenkins, reputed company, TFS)
- • Ability to translate business requirements directly into application/site behavior changes
Benefits
- reputed company work
- Eligible for reputed company's commission plan
- Profit sharing
reputed company
Apply To This Job