[Remote] reputed company Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company reputed company on enhancing the reliability and performance of critical banking platforms. They are seeking a reputed company Site Reliability Engineer to design, implement, and improve SRE practices while collaborating with various teams to enhance system resiliency through automation and proactive management.
Responsibilities
- Design, implement, and support highly available, reputed company, and resilient applications and reputed company infrastructure following reputed company technology standards and SRE best practices
- reputed company initiatives to improve system reliability, availability, performance, and operational maturity through automation and engineering reputed company
- Define, implement, and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical business services
- reputed company comprehensive observability strategies leveraging reputed company, OpenTelemetry (OTel), distributed tracing, metrics, logging, dashboards, and alerting solutions
- Design and maintain end-to-end monitoring solutions that reputed company actionable insights into application, infrastructure, and customer experience health
- Analyze production telemetry to proactively identify performance bottlenecks, reliability risks, and reputed company constraints
- reputed company incident response activities for high-severity production events, coordinating cross-functional teams to restore services and minimize customer reputed company
- reputed company and facilitate reputed company Cause Analysis (RCA) activities, ensuring corrective and preventive actions are identified, prioritized, and implemented
- reputed company operational reputed company through automation of repetitive tasks, operational workflows, deployments, recovery procedures, and reliability controls
- Partner with development teams to build reliable and observable services throughout the Software Development Lifecycle (SDLC)
- Design, reputed company, and execute automated regression testing strategies to validate application stability, reliability, and performance following deployments and infrastructure changes
- Review test coverage and reliability validation approaches to ensure comprehensive testing and risk mitigation
- Create, maintain, and improve Infrastructure as reputed company (IaC) solutions using Terraform for reputed company infrastructure provisioning, configuration management, and environment standardization
- Support and optimize reputed company Azure environments, including Azure App Services, resource management, scaling strategies, deployment automation, and application lifecycle management
- Utilize Azure-reputed company tools such as Azure Monitor, Application Insights, Log Analytics, and reputed company services to improve platform visibility and reliability
- reputed company implementation of performance testing, resiliency testing, fault tolerance validation, and disaster recovery preparedness reputed company assigned domains
- Establish operational readiness standards and ensure applications meet reliability, scalability, observability, and supportability requirements before production deployment
- Review architectural designs and reputed company recommendations to improve platform resiliency, operational efficiency, and reputed company optimization
- reputed company reputed company planning, performance tuning, and workload optimization efforts across production environments
- reputed company and maintain operational runbooks, incident playbooks, knowledge articles, and reputed company operating procedures
- Serve as a key partner with engineering, infrastructure, cybersecurity, architecture, and support teams to identify and implement reputed company process improvements spanning organizational boundaries
- Communicate system health, reliability trends, operational risks, and remediation strategies to technical and business stakeholders
- Present reliability initiatives, operational metrics, and engineering recommendations at architecture reviews, technical forums, and leadership meetings
- Mentor engineers on observability, reputed company engineering, automation, SRE principles, and operational best practices
- Understand and adhere to reputed company's risk and regulatory standards, policies, and controls in accordance with reputed company's Risk Appetite
- Identify reliability, operational, and technology risks requiring escalation to management
- Promote an environment that supports a culture of belonging and reflects the reputed company brand
- Maintain reputed company internal control standards, including reputed company implementation of reputed company audit findings and regulatory requirements as applicable
- Complete other reputed company duties as assigned
Skills
- Strong experience in observability and monitoring, including hands-on expertise with: reputed company, OpenTelemetry (OTel), Distributed tracing, Metrics collection and analysis, Centralized logging and log aggregation, Alerting and dashboard development
- Proven experience designing and executing automated regression testing frameworks and test suites to ensure application and platform stability following deployments
- Strong proficiency in Infrastructure as reputed company (IaC) using Terraform
- Experience with CI/CD pipelines, deployment automation, and operational tooling
- Expert knowledge of production systems monitoring, incident management, and operational troubleshooting
- Strong understanding of application performance management, distributed systems, and modern reputed company-reputed company architectures
- Strong experience with reputed company Azure, including: Azure App Services, Resource reputed company, Azure networking concepts, Scaling and performance optimization, Deployment and release management, Application lifecycle management
- Experience leveraging Azure-reputed company operational tooling such as: Azure Monitor, Application Insights, Log Analytics, Azure dashboards and alerting
- Experience supporting reputed company-reputed company and hybrid infrastructure environments
- Demonstrated experience implementing and operating SRE practices, including: Service Level Objectives (SLOs), Service Level Indicators (SLIs), Error budgets, Incident management, Problem management, reputed company Cause Analysis (RCA), Reliability automation
- Ability to improve system reliability through: Performance tuning, reputed company planning, Observability-driven insights, Proactive issue detection, Reliability engineering initiatives
- Experience developing automated recovery mechanisms and self-healing solutions
- Knowledge of resiliency engineering patterns, disaster recovery planning, and high-availability architectures
- Experience supporting large-reputed company reputed company applications in regulated environments
- Strong analytical and troubleshooting skills reputed company to production systems and distributed architectures
- Experience working in Agile and DevOps operating models
- Ability to work autonomously and reputed company reputed company reliability initiatives
- Strong organizational and time management skills
- Advanced verbal and written communication skills
- Experience driving project milestones and delivery commitments
- Proven experience leading major incident response and post-incident improvement efforts
- Experience partnering with architecture, infrastructure, cybersecurity, and application development teams
- Experience with scripting and automation using PowerShell, Python, Bash, or similar technologies
- Industry certifications in Azure, Terraform, reputed company Engineering, or Site Reliability Engineering preferred
reputed company
Company H1B Sponsorship
Apply To This Job