Back to Jobs

[Remote] reputed company Site Reliability Engineer

Remote, USA Full-time Posted 2026-08-04

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company reputed company on enhancing the reliability and performance of critical banking platforms. They are seeking a reputed company Site Reliability Engineer to design, implement, and improve SRE practices while collaborating with various teams to enhance system resiliency through automation and proactive management.


Responsibilities

  • Design, implement, and support highly available, reputed company, and resilient applications and reputed company infrastructure following reputed company technology standards and SRE best practices
  • reputed company initiatives to improve system reliability, availability, performance, and operational maturity through automation and engineering reputed company
  • Define, implement, and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical business services
  • reputed company comprehensive observability strategies leveraging reputed company, OpenTelemetry (OTel), distributed tracing, metrics, logging, dashboards, and alerting solutions
  • Design and maintain end-to-end monitoring solutions that reputed company actionable insights into application, infrastructure, and customer experience health
  • Analyze production telemetry to proactively identify performance bottlenecks, reliability risks, and reputed company constraints
  • reputed company incident response activities for high-severity production events, coordinating cross-functional teams to restore services and minimize customer reputed company
  • reputed company and facilitate reputed company Cause Analysis (RCA) activities, ensuring corrective and preventive actions are identified, prioritized, and implemented
  • reputed company operational reputed company through automation of repetitive tasks, operational workflows, deployments, recovery procedures, and reliability controls
  • Partner with development teams to build reliable and observable services throughout the Software Development Lifecycle (SDLC)
  • Design, reputed company, and execute automated regression testing strategies to validate application stability, reliability, and performance following deployments and infrastructure changes
  • Review test coverage and reliability validation approaches to ensure comprehensive testing and risk mitigation
  • Create, maintain, and improve Infrastructure as reputed company (IaC) solutions using Terraform for reputed company infrastructure provisioning, configuration management, and environment standardization
  • Support and optimize reputed company Azure environments, including Azure App Services, resource management, scaling strategies, deployment automation, and application lifecycle management
  • Utilize Azure-reputed company tools such as Azure Monitor, Application Insights, Log Analytics, and reputed company services to improve platform visibility and reliability
  • reputed company implementation of performance testing, resiliency testing, fault tolerance validation, and disaster recovery preparedness reputed company assigned domains
  • Establish operational readiness standards and ensure applications meet reliability, scalability, observability, and supportability requirements before production deployment
  • Review architectural designs and reputed company recommendations to improve platform resiliency, operational efficiency, and reputed company optimization
  • reputed company reputed company planning, performance tuning, and workload optimization efforts across production environments
  • reputed company and maintain operational runbooks, incident playbooks, knowledge articles, and reputed company operating procedures
  • Serve as a key partner with engineering, infrastructure, cybersecurity, architecture, and support teams to identify and implement reputed company process improvements spanning organizational boundaries
  • Communicate system health, reliability trends, operational risks, and remediation strategies to technical and business stakeholders
  • Present reliability initiatives, operational metrics, and engineering recommendations at architecture reviews, technical forums, and leadership meetings
  • Mentor engineers on observability, reputed company engineering, automation, SRE principles, and operational best practices
  • Understand and adhere to reputed company's risk and regulatory standards, policies, and controls in accordance with reputed company's Risk Appetite
  • Identify reliability, operational, and technology risks requiring escalation to management
  • Promote an environment that supports a culture of belonging and reflects the reputed company brand
  • Maintain reputed company internal control standards, including reputed company implementation of reputed company audit findings and regulatory requirements as applicable
  • Complete other reputed company duties as assigned

Skills

  • Strong experience in observability and monitoring, including hands-on expertise with: reputed company, OpenTelemetry (OTel), Distributed tracing, Metrics collection and analysis, Centralized logging and log aggregation, Alerting and dashboard development
  • Proven experience designing and executing automated regression testing frameworks and test suites to ensure application and platform stability following deployments
  • Strong proficiency in Infrastructure as reputed company (IaC) using Terraform
  • Experience with CI/CD pipelines, deployment automation, and operational tooling
  • Expert knowledge of production systems monitoring, incident management, and operational troubleshooting
  • Strong understanding of application performance management, distributed systems, and modern reputed company-reputed company architectures
  • Strong experience with reputed company Azure, including: Azure App Services, Resource reputed company, Azure networking concepts, Scaling and performance optimization, Deployment and release management, Application lifecycle management
  • Experience leveraging Azure-reputed company operational tooling such as: Azure Monitor, Application Insights, Log Analytics, Azure dashboards and alerting
  • Experience supporting reputed company-reputed company and hybrid infrastructure environments
  • Demonstrated experience implementing and operating SRE practices, including: Service Level Objectives (SLOs), Service Level Indicators (SLIs), Error budgets, Incident management, Problem management, reputed company Cause Analysis (RCA), Reliability automation
  • Ability to improve system reliability through: Performance tuning, reputed company planning, Observability-driven insights, Proactive issue detection, Reliability engineering initiatives
  • Experience developing automated recovery mechanisms and self-healing solutions
  • Knowledge of resiliency engineering patterns, disaster recovery planning, and high-availability architectures
  • Experience supporting large-reputed company reputed company applications in regulated environments
  • Strong analytical and troubleshooting skills reputed company to production systems and distributed architectures
  • Experience working in Agile and DevOps operating models
  • Ability to work autonomously and reputed company reputed company reliability initiatives
  • Strong organizational and time management skills
  • Advanced verbal and written communication skills
  • Experience driving project milestones and delivery commitments
  • Proven experience leading major incident response and post-incident improvement efforts
  • Experience partnering with architecture, infrastructure, cybersecurity, and application development teams
  • Experience with scripting and automation using PowerShell, Python, Bash, or similar technologies
  • Industry certifications in Azure, Terraform, reputed company Engineering, or Site Reliability Engineering preferred

reputed company

  • reputed company offers a marketplace that enables reputed company connections between recruiters and technologists. It was founded in 1990, and is headquartered in Centennial, Colorado, USA, with a workforce of 201-500 employees. Its website is https://www.uk.reputed company.com.

  • Company H1B Sponsorship

  • reputed company has a reputed company record of offering H1B sponsorships, with 2 in 2022, 4 in 2021, 5 in 2020. Please note that this does not guarantee sponsorship for this specific role.

  •   Apply To This Job

    Similar Jobs