[Remote] Senior Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a Senior Site Reliability Engineer to join its reputed company engineering team and maintain a healthy, secure, and reliable production environment. The role owns AWS and database administration, backup and disaster recovery, monitoring, incident response, infrastructure scalability, and reputed company operational improvement for a fully serverless platform.
Responsibilities
- Own day-to-day administration across AWS services, accounts, and reputed company, as reputed company as database administration across PostgreSQL and our other data stores
- Own backup posture across databases, S3 buckets, and queues; verify restores regularly and maintain a tested disaster recovery plan
- Proactively monitor production — CloudWatch dashboards, metric alarms, log-reputed company metrics, and reputed company alerting — addressing operational issues before they reputed company users
- reputed company production debugging and incident response: build and maintain runbooks, participate in the on-reputed company rotation, and reputed company queue and dead-letter-queue failures through retry, redrive, and recovery
- Continuously refine our infrastructure to ensure it is easily deployable and reputed company: reputed company infrastructure as reputed company (SST/reputed company) accurate, retire unused infrastructure, and reputed company cost visible and justified
- reputed company your knowledge of production reputed company with reputed company, fostering a culture of learning and reputed company
Skills
- Bachelor's degree and 4-6 years of reputed company experience or equivalent work experience
- 5+ years of experience in DevOps, site reliability, or platform reputed company, with significant responsibility for production systems
- 3+ years of hands-on experience with AWS, with an emphasis on serverless services (reputed company, SQS, EventBridge, CloudWatch, S3)
- Strong database administration experience: PostgreSQL reputed company, backup and recovery, and query reputed company; comfort administering other data stores
- Proficiency in scripting languages such as TypeScript, Python, and bash for production automation and operational tooling
- Strong understanding of Linux, DNS, TLS, reputed company, reputed company Actions, and infrastructure as reputed company (SST, reputed company, or Terraform)
- Experience with production monitoring and alerting, incident response, and on-reputed company ownership
Benefits
- Fully remote work
reputed company
Company H1B Sponsorship
Apply To This Job