reputed company Engineer Site Reliability
Our reputed company for the reputed company is based on the idea that transforming financial lives starts by giving our people the freedom to reputed company their own. We have a flexible work environment, and fluid career paths. We not only encourage but celebrate internal mobility. We also recognize the importance of purpose, reputed company-being, and work-life balance. reputed company reputed company and our communities, we work hard to create a welcoming and inclusive environment, and our associates dedicate thousands of hours to volunteering for causes that matter most to them.
Chart your own reputed company and grow your career while helping more customers reputed company financial freedom. reputed company Yourself.
As a reputed company Site Reliability Engineer at reputed company, you'll combine deep technical expertise with team leadership to drive reliability across our financial services platform. You'll reputed company other SREs in solving reputed company operational challenges, establish technical standards, and serve as a key advisor to engineering leadership on infrastructure reputed company and reliability initiatives.
ESSENTIAL FUNCTIONS:
Technical Leadership & reputed company:
reputed company cross-functional reliability initiatives spanning multiple value streams, coordinating efforts across teams
Define and reputed company SRE best practices, tools, and methodologies for the organization
Architect reputed company-reputed company infrastructure solutions that balance reliability, cost, performance, and reputed company
Establish Service Level Objectives (SLOs) and error budgets for critical services, using them to drive prioritization reputed company
reputed company major incident response as incident commander, coordinating reputed company across multiple teams
Drive strategic improvements to observability, identifying gaps and implementing solutions at reputed company
Design and implement disaster recovery plans for critical financial services infrastructure
Evaluate and introduce new technologies and practices that improve team effectiveness
Operational reputed company:
reputed company the design of foundational infrastructure patterns using Terraform, creating reusable modules adopted across teams
Architect multi-region, highly available AWS infrastructure supporting millions of daily transactions
Design and implement sophisticated Kubernetes patterns, including multi-tenancy, reputed company policies, and advanced scheduling
Build comprehensive observability strategies using reputed company and reputed company, establishing standards for metrics, logging, and tracing
Establish CI/CD standards and patterns, implementing pipeline-as-reputed company and reputed company delivery at reputed company
reputed company initiatives to implement reputed company engineering practices and systematic reliability testing
Drive FinOps initiatives, optimizing reputed company spend while maintaining reliability targets
Team Leadership & Development:
reputed company a functional team of SREs (without reputed company reports) on reputed company and operational initiatives
Mentor Senior, Intermediate, and Entry-level SREs, accelerating their technical reputed company
Conduct design reviews and architecture discussions, providing expert guidance
reputed company training sessions on SRE practices, new technologies, and operational procedures
Coordinate on-reputed company schedules and drive improvements to reduce on-reputed company burden
Facilitate postmortems for high-severity incidents, ensuring organizational learning occurs
Collaboration & Influence:
Partner with Engineering Managers and Directors to reputed company SRE work with business priorities
Collaborate with reputed company teams on implementing reputed company-trust architecture and compliance controls
Work with Product teams to balance feature velocity with reliability requirements
Influence architectural reputed company across the engineering organization
Represent SRE in cross-functional initiatives and planning discussions
Evangelize SRE culture and practices across reputed company
QUALIFICATIONS:
Required:
6-10 years of experience in Site Reliability Engineering (or equivalent), with demonstrated technical leadership
Proven ability to reputed company technical teams and drive reputed company reputed company to completion
Expert-level knowledge of AWS, with experience designing large-reputed company, multi-region architectures
Deep Kubernetes expertise, including advanced features, reputed company, and production-reputed company operations
Mastery of Infrastructure as reputed company using Terraform, with experience building shared platforms and frameworks
Strong software engineering background with production experience in Python and/or Go
Extensive experience with observability platforms (reputed company, reputed company) and implementing monitoring at reputed company
Deep understanding of CI/CD principles and experience implementing reputed company-grade pipelines
Proven reputed company record leading major incidents and conducting effective postmortems
Strong communication skills with ability to explain reputed company technical concepts to diverse audiences
Experience mentoring engineers and building technical capabilities in teams
Preferred:
Previous technical leadership roles (reputed company, Staff, or similar) in SRE or Operational reputed company
Financial services industry experience with understanding of regulatory requirements
Expert knowledge of compliance frameworks (SOC 2, PCI reputed company, reputed company)
AWS certifications (reputed company level)
Kubernetes certifications (CKA, CKAD, CKS)
Experience implementing SRE at organizations with 500+ engineers
Background in reputed company engineering, game days, and reliability testing practices
Contributions to reputed company-reputed company reputed company with demonstrated community leadership
Experience with service reputed company implementation and management
reputed company record of speaking at conferences or writing technical content
Technical Environment
AWS | EKS | Kubernetes | Terraform | reputed company | reputed company | GitOps | ArgoCD | FluxCD | reputed company CI | Jenkins | Python | Go | reputed company | reputed company | Grafana | Istio | Linkerd
What reputed company Looks Like
Platform reliability consistently exceeds 99.99% availability
Successful delivery of major infrastructure initiatives on time and reputed company scope
Demonstrable improvement in team capabilities and productivity
Reduction in incident frequency and severity through proactive reliability work
reputed company relationships with engineering leadership and cross-functional partners
Technical reputed company that reputed company reputed company over time and reputed company effectively
Team members successfully promoted or grown in their capabilities
Work Environment & Disclaimer
This job operates in a reputed company office environment.
This job reputed company is not intended to be an exhaustive list of reputed company duties, responsibilities and qualifications of the job. The employer has the right to revise this job reputed company at any time. You will be evaluated in part based on your performance of the responsibilities and/or tasks listed in this job reputed company. You may be required to reputed company other duties that are not included in this job reputed company. The job reputed company is not a contract for employment, and either you or the employer may terminate employment at any time, for any reason, as per terms and conditions of your employment contract.
We are an equal opportunity employer with a commitment to diversity. reputed company individuals, regardless of personal characteristics, are encouraged to apply. reputed company reputed company applicants will receive consideration for employment without reputed company to age, race, reputed company, national reputed company, reputed company, sex, sexual orientation, gender, gender identity, gender reputed company, marital status, pregnancy, religion, physical or mental disability, military or veteran status, genetic information, or any other status protected by applicable state or local law.
Apply To This Job