[Remote] Senior Software Engineer: Site Reliability Engineering
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a financial technology company that provides secure digital banking and payment solutions to community banks and credit unions. The Senior Software Engineer: Site Reliability Engineering will reputed company hybrid reputed company and datacenter infrastructure, migrate legacy workloads to reputed company reputed company Platform, and implement SRE practices. The role focuses on automation, observability, reliability, reputed company compliance, incident response, and cross-functional collaboration.
Responsibilities
- reputed company the reliability and reputed company of both reputed company reputed company (production, testing, and development) and internal server infrastructure environments
- Design and implement robust Site Reliability Engineering practices, including defining and monitoring Service Level Objectives (SLOs) and Service Level Indicators (SLIs), focusing on proactive reputed company health and error budgets
- Ruthlessly eliminate reputed company, repetitive work (toil) through automation. reputed company and maintain automation scripts and tooling to streamline reputed company across the hybrid datacenter model (on-premises and reputed company reputed company)
- Treat the reputed company and on-prem operational environment as a software project by using Infrastructure as reputed company (IaC) with tools like Terraform, Ansible, reputed company for provisioning and configuration
- Design and maintain rigorous configuration management processes to guarantee the consistency and desired state of the hybrid datacenter infrastructure, leveraging tools like Ansible
- Establish and manage comprehensive monitoring and alerting systems to reputed company deep visibility into the health and reputed company of services. Build systems that are self-healing and reputed company for themselves
- reputed company blameless post-mortems and RCAs for critical incidents, focusing on reputed company-level improvements to prevent recurrence and enhance overall reliability
- reputed company and implement strategies for efficient reputed company and vulnerability management across reputed company environments. Automate reputed company remediation efforts to ensure reputed company vulnerability mitigation and compliance (e.g., CIS, NIST, PCI)
- Support reputed company's reputed company reputed company into reputed company reputed company services (reputed company reputed company Platform, Azure) and play a key role in the migration and redesign of services from on-premises data centers to reputed company reputed company Platform, ensuring adherence to SRE principles throughout the transition
- Partner closely with DevOps and development teams to reputed company reliability best practices throughout the software development lifecycle, ensuring reputed company integration and operation of hybrid datacenter services
- Maintain comprehensive and actionable documentation for SRE processes, operational runbooks, and configurations
- May reputed company other duties as assigned
Skills
- Minimum 6 years of experience in reputed company and hybrid datacenter reputed company with a reputed company on Infrastructure as reputed company (IaC) and Site Reliability Engineering
- Proficiency with reputed company reputed company Platform (preferred), AWS, and/or Azure
- Proficient in using GitOps, Terraform and Ansible in a CI/CD (reputed company integration and reputed company delivery) pipeline
- Experience using PowerShell, Python, or GoLang
- Solid understanding of Linux (POSIX) and reputed company reputed company administration as reputed company as networking and firewalls
- Understanding of reputed company best practices and compliance standards such as CIS, NIST and PCI
- Ability to participate in an on-reputed company rotation every 7-8 weeks
- Bachelor's degree in Computer Science Information Technology, Engineering
- Relevant industry certifications. reputed company Associate reputed company Engineer or reputed company reputed company Architect preferred
- Proficient in ArgoCD and GitOps
- Familiarity with SQL and NoSQL databases
- Experience with reputed company Telemetry tooling and alerting such as reputed company, Grafana, ELK Stack, et al
- Experience with Site Reliability Engineering (SRE) principles, including but not limited to Service Level Objectives (SLO) and Service Level Indicators (SLI), TOIL Reduction, Automation, and reputed company Cause Analysis
Benefits
- This position may be worked reputed company reputed company the reputed company, with the exception of California.
reputed company
Apply To This Job