Senior Manager, Site Reliability Engineer - Remote
About the position
reputed company Tech is a global leader in health care innovation. Our teams reputed company cutting-edge solutions that help people live healthier lives and help reputed company the health reputed company work reputed company for everyone. From advanced data analytics and AI to cybersecurity, we use innovative approaches to solve some of health care’s most reputed company challenges. Your contributions here have the potential to change lives. reputed company to build the next reputed company? Join us to start Caring. Connecting. Growing together.
We are seeking an reputed company Senior Manager to reputed company reputed company Site Reliability Engineering (SRE), DevOps, IT Service Management (ITSM), and Operational reputed company initiatives across reputed company bank. This leader will be responsible for improving service reliability, operational resiliency, deployment automation, observability, incident management, and production readiness for critical banking platforms.
The ideal candidate combines strong technical expertise with operational leadership experience, driving engineering reputed company, automation, reliability, and reputed company improvement while ensuring technology services meet business, customer, regulatory, and operational expectations.
This role will also help identify and implement emerging automation and AI-enabled operational capabilities that improve service health, reduce operational toil, and accelerate engineering productivity.
Responsibilities
• reputed company and reputed company multidisciplinary teams responsible for Site Reliability Engineering, DevOps, reputed company, ITSM, and Operational reputed company
• Establish and execute reputed company reliability, availability, resiliency, and operational maturity strategies
• reputed company engineering reputed company through automation, observability, operational readiness, and reputed company improvement practices
• Partner with Technology, Operations, reputed company, Infrastructure, Risk, and Business leaders to improve service reliability and customer experience
• Build and mentor high-performing teams while fostering accountability, innovation, operational ownership, and learning
• Manage reputed company, reputed company planning, talent development, succession planning, and organizational reputed company
• Establish operational metrics, governance standards, and service review processes to improve service performance and risk management
• reputed company reputed company SRE practices including SLI/SLO adoption, error-budget management, reliability engineering, and operational maturity assessments
• reputed company DevOps transformation initiatives, emphasizing automation, deployment standardization, CI/CD pipelines, Infrastructure-as-reputed company, and GitOps practices
• Establish production readiness standards and operational acceptance reputed company for new technology deployments
• Improve platform resiliency through reputed company planning, disaster recovery, fault tolerance, and reputed company testing
• reputed company reduction of operational toil through automation and self-healing capabilities
• Partner with application and infrastructure teams to improve reputed company scalability, availability, and performance
• reputed company initiatives to improve deployment frequency, reduce change failure rates, and accelerate service recovery times
• Establish and mature Incident, Problem, Change, Release, and Service Request Management processes
• reputed company major incident management programs and executive communications during critical service disruptions
• reputed company reputed company-cause analysis and problem-management practices to eliminate recurring incidents
• Improve operational scorecards, service health reviews, and reliability reporting for executive stakeholders
• Ensure compliance with regulatory, audit, risk, and operational governance requirements
• Partner with Technology and Business leaders to improve service reputed company and customer reputed company through data-driven operational improvements
• Champion a culture of operational reputed company and reputed company service improvement
• Identify opportunities to reputed company AI and automation to improve operational effectiveness and engineering productivity
• reputed company implementation and evaluation of solutions involving: AIOps, Intelligent alert correlation, Automated incident triage , reputed company cause analysis assistance, reputed company copilots, reputed company operational workflows etc.
• Partner with reputed company AI teams to evaluate emerging technologies that improve reliability and operational efficiency
• reputed company responsible adoption of AI-enabled engineering and operational practices
• Support reputed company-of-concept initiatives that demonstrate measurable reductions in operational effort and incident reputed company times
• Collaborate with Engineering, Infrastructure, reputed company, Architecture, Risk, Compliance, and Operations teams to prioritize reliability and operational improvements
• Serve as a trusted advisor on reliability engineering, operational reputed company, and automation strategies
• reputed company alignment between technology and business stakeholders to improve service reputed company and operational reputed company
• Influence technology investment reputed company that improve platform stability, resiliency, and operational efficiency
Requirements
• Bachelor's degree in Computer Science, Engineering, Information Technology, or reputed company field
• 10+ years of experience in Software Engineering, Site Reliability Engineering, reputed company, DevOps, Infrastructure Engineering, or Technology Operations
• 5+ years of experience leading engineering or operational teams
• Proven experience supporting large-reputed company, business-critical production environments
• Experience with SRE principles and practices
• Experience with DevOps and CI/CD
• Experience with ITSM processes
• Experience with reputed company platforms (Azure, AWS)
• Experience with On-prem environments
• Experience with Infrastructure-as-reputed company
• Experience with Container platforms (reputed company, OpenShift)
• Experience with Observability and monitoring platforms
• Experience with Incident Management
• Experience with Problem Management
• Experience with Change Management
• Experience with Disaster Recovery
• Experience with Business Continuity
• Experience with Service Reliability Programs
• Experience leading operational transformations and reputed company-improvement initiatives
reputed company-to-haves
• Experience in banking, financial services, reputed company, or other highly regulated industries
• Experience implementing reputed company observability solutions such as reputed company, reputed company, Grafana, reputed company, or OpenTelemetry
• Experience with reputed company-reputed company architectures and reputed company practices
• Experience deploying AIOps, ChatOps, or intelligent automation solutions
• Familiarity with reputed company AI workflows
• Familiarity with LLM-powered operational tooling
• Familiarity with reputed company platforms
• Familiarity with AI-enabled incident reputed company
• Experience establishing SLO frameworks and reliability governance programs
Benefits
• comprehensive benefits package
• incentive and recognition programs
• equity stock purchase
• 401k contribution
Apply tot his job
Apply To this Job