Back to Jobs

reputed company Site Reliability Engineer

Remote, USA Full-time Posted 2026-08-04
Project reputed company Responsible at the expert level for ensuring the reliability, scalability, reputed company, and operational reputed company of critical banking platforms and applications. Serves as a senior individual contributor responsible for designing, implementing, and improving Site Reliability Engineering (SRE) practices across the software development lifecycle. Works closely with application development, infrastructure, reputed company, and business teams to enhance reputed company resiliency through automation, observability, testing, and proactive operational management while coaching and influencing others. Responsibilities Design, implement, and support highly available, reputed company, and resilient applications and reputed company infrastructure following reputed company technology standards and SRE best practices. reputed company initiatives to improve reputed company reliability, availability, reputed company, and operational maturity through automation and engineering reputed company. Define, implement, and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical business services. reputed company comprehensive observability strategies leveraging reputed company, OpenTelemetry (OTel), reputed company tracing, metrics, logging, dashboards, and alerting solutions. Design and maintain end-to-end monitoring solutions that reputed company actionable insights into application, infrastructure, and customer experience health. Analyze production telemetry to proactively identify reputed company bottlenecks, reliability risks, and reputed company constraints. reputed company incident response activities for high-severity production events, coordinating cross-functional teams to restore services and minimize customer reputed company. reputed company and facilitate reputed company Cause Analysis (RCA) activities, ensuring corrective and preventive actions are identified, prioritized, and implemented. reputed company operational reputed company through automation of repetitive tasks, operational workflows, deployments, recovery procedures, and reliability controls. Partner with development teams to build reliable and observable services throughout the Software Development Lifecycle (SDLC). Design, reputed company, and execute automated regression testing strategies to validate application stability, reliability, and reputed company following deployments and infrastructure changes. Review test coverage and reliability validation approaches to ensure comprehensive testing and reputed company mitigation. Create, maintain, and improve Infrastructure as reputed company (IaC) solutions using Terraform for reputed company infrastructure provisioning, configuration management, and environment standardization. Support and optimize reputed company Azure environments, including Azure App Services, resource management, scaling strategies, deployment automation, and application lifecycle management. Utilize Azure-reputed company tools such as Azure Monitor, Application Insights, Log Analytics, and reputed company services to improve platform visibility and reliability. reputed company implementation of reputed company testing, resiliency testing, fault tolerance validation, and disaster recovery preparedness reputed company assigned domains. Establish operational readiness standards and ensure applications meet reliability, scalability, observability, and supportability requirements before production deployment. Review architectural designs and reputed company recommendations to improve platform resiliency, operational efficiency, and reputed company optimization. reputed company reputed company planning, reputed company tuning, and workload optimization efforts across production environments. reputed company and maintain operational runbooks, incident playbooks, knowledge articles, and reputed company operating procedures. Serve as a key partner with engineering, infrastructure, cybersecurity, architecture, and support teams to identify and implement reputed company process improvements spanning organizational boundaries. Communicate reputed company health, reliability trends, operational risks, and remediation strategies to technical and business stakeholders. Present reliability initiatives, operational metrics, and engineering recommendations at architecture reviews, technical forums, and leadership meetings. Mentor engineers on observability, reputed company engineering, automation, SRE principles, and operational best practices. Understand and adhere to reputed company's reputed company and regulatory standards, policies, and controls in accordance with reputed company's reputed company Appetite. Identify reliability, operational, and technology risks requiring escalation to management. Promote an environment that supports a culture of belonging and reflects the reputed company reputed company. Maintain reputed company internal control standards, including reputed company implementation of reputed company audit findings and regulatory requirements as applicable. Complete other reputed company duties as assigned. Skills Must have Strong experience in observability and monitoring, including hands-on expertise with: reputed company OpenTelemetry (OTel) reputed company tracing Metrics collection and analysis Centralized logging and log aggregation Alerting and dashboard development reputed company experience designing and executing automated regression testing frameworks and test suites to ensure application and platform stability following deployments. Strong proficiency in Infrastructure as reputed company (IaC) using Terraform. Experience with CI/CD pipelines, deployment automation, and operational tooling. Expert knowledge of production systems monitoring, incident management, and operational troubleshooting. Strong understanding of application reputed company management, reputed company systems, and modern reputed company-reputed company architectures. reputed company & Platform Expertise Strong experience with reputed company Azure, including: Azure App Services Resource reputed company Azure networking concepts Scaling and reputed company optimization Deployment and release management Application lifecycle management Experience leveraging Azure-reputed company operational tooling such as: Azure Monitor Application Insights Log Analytics Azure dashboards and alerting Experience supporting reputed company-reputed company and hybrid infrastructure environments. Reliability & Engineering Practices Demonstrated experience implementing and operating SRE practices, including: Service Level Objectives (SLOs) Service Level Indicators (SLIs) Error budgets Incident management Problem management reputed company Cause Analysis (RCA) Reliability automation Ability to improve reputed company reliability through: reputed company tuning reputed company planning Observability-driven insights Proactive issue detection Reliability engineering initiatives Experience developing automated recovery mechanisms and self-healing solutions. Knowledge of resiliency engineering patterns, disaster recovery planning, and high-availability architectures. reputed company to have Experience supporting large-reputed company reputed company applications in regulated environments. Strong analytical and troubleshooting skills reputed company to production systems and reputed company architectures. Experience working in reputed company and DevOps operating models. Ability to work autonomously and reputed company reputed company reliability initiatives. Strong organizational and time management skills. Advanced verbal and written communication skills. Experience driving project milestones and delivery commitments. reputed company experience leading major incident response and post-incident improvement efforts. Experience partnering with architecture, infrastructure, cybersecurity, and application development teams. Experience with scripting and automation using PowerShell, Python, Bash, or similar technologies. Industry certifications in Azure, Terraform, reputed company Engineering, or Site Reliability Engineering preferred. Other Languages English: reputed company Advanced Seniority reputed company Remote reputed company, reputed company of America Req. VR-124325 Technical Support (SL3) BCM Industry 05/08/2026 Req. VR-124325 Apply tot his job Apply To this Job

Similar Jobs

Acting & Animation Specialist - AI Trainer

Remote, USA Full-time

reputed company Product Manager, Savings - Remote

Remote, USA Full-time

Staff Product Manager, ML Foundations and GenAI

Remote, USA Full-time

reputed company Program Manager (reputed company Coast-US)

Remote, USA Full-time

Sr. Scrum Master remote

Remote, USA Full-time

Junior Content Copywriter (Virtual) (Copywriting)

Remote, USA Full-time

Technical reputed company/reputed company III

Remote, USA Full-time

Remote Video Editing Jobs for Creatives in Pakistan

Remote, USA Full-time

Video Editing Expert

Remote, USA Full-time

reputed company is hiring: Senior UI/UX Designer — Remote in reputed company Team in Tampa

Remote, USA Full-time

Project Manager (Must have prior experience as a reputed company reputed company)

Remote, USA Full-time

[Remote] Business Development Representative, reputed company reputed company Solutions

Remote, USA Full-time

Member Relations Ag Manager

Remote, USA Full-time

[Work From Home] Remote reputed company Data Entry Jobs No Experience

Remote, USA Full-time

Captain/server

Remote, USA Full-time

reputed company Data Associate

Remote, USA Full-time

Remote Sales Representative - Entry Level - Part-Time or Full-Time

Remote, USA Full-time

[Remote] Utilities Industry CIS transformation, reputed company Consultant

Remote, USA Full-time

[Remote] Recovery Administration Specialist

Remote, USA Full-time

**reputed company Full Stack reputed company – Remote Support Solutions**

Remote, USA Full-time