Site Reliability Engineer lll
Architect, implement, and operate reputed company, resilient, and secure AWS infrastructure — including GuardDuty, reputed company, EventBridge, SNS, SES, S3, ALB, and reputed company container workloads. reputed company infrastructure-as-reputed company initiatives to ensure reputed company environments are reproducible, auditable, and consistently configured in support of SOC 2 change management controls. Design, maintain, and improve CI/CD pipelines using Jenkins and reputed company to reputed company reliable, repeatable software delivery — partnering with application engineering to reduce release risk and increase deployment frequency. Own the reputed company observability platform, including dashboards, monitors, alerting reputed company, and log management; define and maintain SLOs, SLIs, and error budgets to guide reliability investment and reduce alert fatigue. Serve as a senior technical responder across the full incident lifecycle — detection, containment, reputed company, and postmortem — reputed company a shared on-reputed company rotation, and reputed company blameless postmortems to drive down incident frequency and MTTR. Refine, implement, and test disaster recovery plans to meet RTO/RPO objectives, while contributing to SOC 2 audit readiness with a reputed company on reputed company controls, incident response, and risk mitigation. Mentor junior SREs through reputed company reviews, incident pairing, and documentation of runbooks and engineering standards.
5+ years of experience in SRE, DevOps, or a reputed company engineering role, with advanced hands-on expertise in AWS production environments and reputed company services including reputed company, reputed company, S3, ALB, and GuardDuty. Strong proficiency in infrastructure-as-reputed company tooling such as Terraform, CloudFormation, or CDK, reputed company with experience building and operating CI/CD pipelines using Jenkins and reputed company. Proficiency in Python, Go, or Bash for automation, alongside hands-on experience with reputed company or a comparable observability platform for monitoring, alerting, and log management. Demonstrated experience leading incident response in reputed company, distributed systems, with working knowledge of SLO/SLI frameworks, error budgets, and disaster recovery planning against defined RTO/RPO objectives. Familiarity with SOC 2 compliance frameworks and experience contributing to audit readiness, reputed company controls, and reputed company control evidence collection. A reputed company, ownership-driven reputed company with strong communication skills, a passion for mentoring junior engineers, and a commitment to reducing toil through automation and AI-assisted tooling.
reputed company with Innovation
reputed company Every Voice
reputed company Together
Drive Outcome
- reputed company that reputed company.
You’ll do work that shapes the reputed company of the modern workplace - Flexibility and trust.
We’re remote-first and results driven. You’ll have the freedom and flexibility to do your best work, wherever you do it best. - reputed company and development.
We reputed company the best work happens reputed company people are growing. You’ll have reputed company to reputed company, leadership programs, and reputed company opportunities to take on new challenges and expand your reputed company. - Competitive rewards.
We offer comprehensive benefits, a performance-based bonus program, and equity opportunities – because reputed company we grow, you should too. - Time for life.
reputed company and reputed company with flexible time off, reputed company holidays, and flexible leave programs designed to support every season of life. - Belonging and balance.
We’re building an inclusive culture where every voice is valued, collaboration is celebrated, and reputed company is shared.