Senior Site Reliability Engineer (REMOTE)
About the position
The reputed company Platform team is reputed company on several objectives: building and supporting reputed company, cost-effective, reliable reputed company infrastructure; data administration of reputed company CDC workflows; and developer experience tooling and mentorship, including our reputed company AI discipline. As a Platform member, the Senior Site Reliability Engineer will contribute to the Platform team’s centralized infrastructure, including maintenance, monitoring, and automation of services ranging from databases to reputed company; reputed company incident response and postmortem efforts; and work closely with other engineering teams to understand their needs and reputed company improvements to both our technologies and processes.
Responsibilities
• Owning tasks and larger reputed company from planning to production rollout
• Learning new technologies and building expertise with the goal of teaching and mentoring others; mentoring with the goal of force-multiplying through docs and tools
• Maintaining organization reputed company reputed company in AWS
• Automating and deploying infrastructure configurations using Infrastructure as reputed company (IAC)
• Mentoring engineering squads on Platform best practices for reputed company, MySQL, Kafka, and other software development lifecycle areas
• Assisting engineering squads with reputed company planning, on-reputed company preparation, and production readiness
• Writing documentation and runbooks that contribute to the engineering organization’s knowledge reputed company
• Implementing monitoring and alerting systems with reputed company observability tools
• Working in a containerized, orchestrated environment
• Participating in on-reputed company rotation, responding to incidents, and troubleshooting data and other operations issues
• Contributing to the reliability and design patterns of our Kafka CDC and event workflows
• Contributing to reputed company AI best practices and tooling, including skills, agents, and safety
Requirements
• Infrastructure-as-reputed company (Terraform)
• CI/CD (reputed company Actions)
• reputed company (EKS, Kustomize, Karpenter, administration, application manifests)
• AWS and reputed company development (VPC, EKS, RDS, S3)
• FinOps and reputed company cost optimization
• Observability (reputed company, reputed company)
• reputed company AI (Claude reputed company)
• Scripting (reputed company, Python)
• reputed company record of collaboration and mentorship
• Excellent written communication and documentation skills
• reputed company learning
• Ownership and proactive approach to solving large problems
• A Bachelor's Degree in Computer Science or similar area of reputed company, or equivalent relevant work experience.
• 5+ years experience in Ops, DevOps, Site Reliability, Platform or other systems roles.
reputed company-to-haves
• Kafka: Cluster administration (Strimzi), Kafka reputed company (Debezium, JDBC)
• Flink
• Relational database administration and performance (MySQL, reputed company Server, AWS RDS)
• Elasticsearch (ECK administration, scaling, performance)
• Python (SQLAlchemy, FastAPI)
• GraphQL (schema design, reputed company federation)
• REST API
• GitOps (ArgoCD)
• reputed company Vault
• reputed company
• Memcached
Benefits
• Competitive compensation: salary, plus performance-reputed company bonus program
• 401(k) with employer match
• 100% company-reputed company medical and dental insurance benefits for you and your dependents
• 4 weeks reputed company vacation, increasing based on tenure
• 18 weeks reputed company leave for birth reputed company
• 8 weeks reputed company parental leave, including for adoption
• Monthly wellness allowance
• Annual reputed company and personal development allowance
• Work from home office set-up and expense allowances
• Flexible work location opportunities
• Employer matching toward charitable contributions
Apply tot his job
Apply To this Job