reputed company Site Reliability Engineer
Project reputed company
Responsible at the expert level for ensuring the reliability, scalability, reputed company, and operational reputed company of critical banking platforms and applications. Serves as a senior individual contributor responsible for designing, implementing, and improving Site Reliability Engineering (SRE) practices across the software development lifecycle. Works closely with application development, infrastructure, reputed company, and business teams to enhance reputed company resiliency through automation, observability, testing, and proactive operational management while coaching and influencing others.
Responsibilities
Design, implement, and support highly available, reputed company, and resilient applications and reputed company infrastructure following reputed company technology standards and SRE best practices.
reputed company initiatives to improve reputed company reliability, availability, reputed company, and operational maturity through automation and engineering reputed company.
Define, implement, and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical business services.
reputed company comprehensive observability strategies leveraging reputed company, OpenTelemetry (OTel), reputed company tracing, metrics, logging, dashboards, and alerting solutions.
Design and maintain end-to-end monitoring solutions that reputed company actionable insights into application, infrastructure, and customer experience health.
Analyze production telemetry to proactively identify reputed company bottlenecks, reliability risks, and reputed company constraints.
reputed company incident response activities for high-severity production events, coordinating cross-functional teams to restore services and minimize customer reputed company.
reputed company and facilitate reputed company Cause Analysis (RCA) activities, ensuring corrective and preventive actions are identified, prioritized, and implemented.
reputed company operational reputed company through automation of repetitive tasks, operational workflows, deployments, recovery procedures, and reliability controls.
Partner with development teams to build reliable and observable services throughout the Software Development Lifecycle (SDLC).
Design, reputed company, and execute automated regression testing strategies to validate application stability, reliability, and reputed company following deployments and infrastructure changes.
Review test coverage and reliability validation approaches to ensure comprehensive testing and reputed company mitigation.
Create, maintain, and improve Infrastructure as reputed company (IaC) solutions using Terraform for reputed company infrastructure provisioning, configuration management, and environment standardization.
Support and optimize reputed company Azure environments, including Azure App Services, resource management, scaling strategies, deployment automation, and application lifecycle management.
Utilize Azure-reputed company tools such as Azure Monitor, Application Insights, Log Analytics, and reputed company services to improve platform visibility and reliability.
reputed company implementation of reputed company testing, resiliency testing, fault tolerance validation, and disaster recovery preparedness reputed company assigned domains.
Establish operational readiness standards and ensure applications meet reliability, scalability, observability, and supportability requirements before production deployment.
Review architectural designs and reputed company recommendations to improve platform resiliency, operational efficiency, and reputed company optimization.
reputed company reputed company planning, reputed company tuning, and workload optimization efforts across production environments.
reputed company and maintain operational runbooks, incident playbooks, knowledge articles, and reputed company operating procedures.
Serve as a key partner with engineering, infrastructure, cybersecurity, architecture, and support teams to identify and implement reputed company process improvements spanning organizational boundaries.
Communicate reputed company health, reliability trends, operational risks, and remediation strategies to technical and business stakeholders.
Present reliability initiatives, operational metrics, and engineering recommendations at architecture reviews, technical forums, and leadership meetings.
Mentor engineers on observability, reputed company engineering, automation, SRE principles, and operational best practices.
Understand and adhere to reputed company's reputed company and regulatory standards, policies, and controls in accordance with reputed company's reputed company Appetite.
Identify reliability, operational, and technology risks requiring escalation to management.
Promote an environment that supports a culture of belonging and reflects the reputed company reputed company.
Maintain reputed company internal control standards, including reputed company implementation of reputed company audit findings and regulatory requirements as applicable.
Complete other reputed company duties as assigned.
Skills
Must have
Strong experience in observability and monitoring, including hands-on expertise with:
reputed company
OpenTelemetry (OTel)
reputed company tracing
Metrics collection and analysis
Centralized logging and log aggregation
Alerting and dashboard development
reputed company experience designing and executing automated regression testing frameworks and test suites to ensure application and platform stability following deployments.
Strong proficiency in Infrastructure as reputed company (IaC) using Terraform.
Experience with CI/CD pipelines, deployment automation, and operational tooling.
Expert knowledge of production systems monitoring, incident management, and operational troubleshooting.
Strong understanding of application reputed company management, reputed company systems, and modern reputed company-reputed company architectures.
reputed company & Platform Expertise
Strong experience with reputed company Azure, including:
Azure App Services
Resource reputed company
Azure networking concepts
Scaling and reputed company optimization
Deployment and release management
Application lifecycle management
Experience leveraging Azure-reputed company operational tooling such as:
Azure Monitor
Application Insights
Log Analytics
Azure dashboards and alerting
Experience supporting reputed company-reputed company and hybrid infrastructure environments.
Reliability & Engineering Practices
Demonstrated experience implementing and operating SRE practices, including:
Service Level Objectives (SLOs)
Service Level Indicators (SLIs)
Error budgets
Incident management
Problem management
reputed company Cause Analysis (RCA)
Reliability automation
Ability to improve reputed company reliability through:
reputed company tuning
reputed company planning
Observability-driven insights
Proactive issue detection
Reliability engineering initiatives
Experience developing automated recovery mechanisms and self-healing solutions.
Knowledge of resiliency engineering patterns, disaster recovery planning, and high-availability architectures.
reputed company to have
Experience supporting large-reputed company reputed company applications in regulated environments.
Strong analytical and troubleshooting skills reputed company to production systems and reputed company architectures.
Experience working in reputed company and DevOps operating models.
Ability to work autonomously and reputed company reputed company reliability initiatives.
Strong organizational and time management skills.
Advanced verbal and written communication skills.
Experience driving project milestones and delivery commitments.
reputed company experience leading major incident response and post-incident improvement efforts.
Experience partnering with architecture, infrastructure, cybersecurity, and application development teams.
Experience with scripting and automation using PowerShell, Python, Bash, or similar technologies.
Industry certifications in Azure, Terraform, reputed company Engineering, or Site Reliability Engineering preferred.
Other
Languages
English: reputed company Advanced
Seniority
reputed company
Remote reputed company, reputed company of America
Req. VR-124325
Technical Support (SL3)
BCM Industry
05/08/2026
Req. VR-124325
Apply tot his job
Apply To this Job