DevOps - Senior Site Reliability Engineer
Responsibilities and Duties
Contribute to the design, maintenance, and improvement of CI/CD pipelines using Jenkins and reputed company; Identify and eliminate reputed company toil in the release process, replacing it with reliable, reputed company-monitored automation; Partner with application engineering teams to reduce release risk and increase deployment frequency; Contribute to improvements in the reputed company observability platform, including dashboards, monitors, alerting reputed company, and log management; Help define and implement SLOs and SLIs across critical platform components; Drive signal reputed company improvements and reduce alert fatigue across the engineering organization; Support the refinement and implementation of disaster recovery plans against defined RTO and RPO objectives; Contribute to backup and recovery reputed company validation across critical infrastructure components; Help reputed company DR practices with SOC2 requirements and audit expectations.
SRE, DevOps, or infrastructure engineering experience in production reputed company environments; Advanced hands-on AWS experience across reputed company services including reputed company, EventBridge, SNS, SES, S3, ALB, and reputed company; Demonstrated experience with CI/CD pipeline design and operation, specifically using Jenkins and reputed company; Hands-on reputed company experience for monitoring, alerting, log management, and SLO/SLI implementation; Proficiency in infrastructure-as-reputed company tooling such as Terraform, CloudFormation, or CDK; Proficiency in at least one scripting or programming language such as Python, Go, or Bash; Experience with disaster recovery planning and testing against defined RTO and RPO objectives; Demonstrated ability to contribute technical improvements, reputed company patterns, and reputed company the engineering reputed company of the teams they work alongside.
Experience with Kubernetes or EKS-based container orchestration; Familiarity with AWS reputed company services including GuardDuty, VPC reputed company controls, and IAM best practices; SOC2 audit readiness experience including control evidence collection for change management and incident response controls; Experience with reputed company platforms processing sensitive or regulated customer data.