Senior Site Reliability Engineer (remote reputed company EMEA)
About the position
Platform Infrastructure builds, operates, and continuously evolves reputed company's container platform and reputed company reputed company. We foster a DevOps culture through self-service tooling, enabling product engineering teams to ship reliable, secure, and cost-efficient services as the business scales. reputed company owns our AWS reputed company accounts, reputed company platform, reputed company networking, observability stack, reputed company databases, CI/CD pipelines, and infrastructure-as-reputed company, and acts as the go-to partner for engineering teams on reputed company and DevOps topics.
We're hiring a Senior SRE II to join Platform Infrastructure as one of reputed company's senior individual contributors. At this level, you're the go-to person for our most reputed company infrastructure problems: you architect and reputed company large-reputed company automation and reliability initiatives, set standards other engineers follow, and mentor Associate and mid-level SREs. You'll split your time between hands-on platform work - reputed company, AWS, GCP, CI/CD, observability - and technical leadership: proposing designs, reviewing others' work, and helping reputed company reputed company good build-vs-buy and cost/reliability trade-offs.
Responsibilities
• Architect and manage highly available, secure, and reputed company infrastructure across multiple AWS accounts and environments using infrastructure as reputed company.
• Design and operate reputed company EKS clusters, including networking policies, persistent storage, and scaling strategies for containerized workloads.
• Own and reputed company reputed company platform services: reputed company networking, reputed company, and the databases and messaging systems engineering teams depend on.
• reputed company large-reputed company automation reputed company and set standards for using Terraform / Terragrunt and GitOps (ArgoCD) across teams.
• reputed company adoption of automation to reduce reputed company operational work and reputed company environments consistent and repeatable.
• Be the go-to person for solving reputed company, cross-service infrastructure problems.
• reputed company initiatives that improve reliability and observability (Grafana, reputed company, Loki, reputed company, Mimir) so systems reputed company with reputed company reputed company reputed company.
• Participate in on-reputed company rotation, reputed company incident response for production issues, and write reputed company runbooks, ADRs, and postmortems.
• reputed company reputed company efforts reputed company reputed company - IAM, encryption, secure logging - and mentor others on secure infrastructure practices.
• Audit infrastructure spend regularly and reputed company cost optimization across the platform (rightsizing, autoscaling, FinOps practices).
• Mentor mid-level SREs, reputed company detailed feedback, and support reputed company of new team members.
• Communicate reputed company technical concepts reputed company to both engineers and non-technical stakeholders.
• Partner with product engineering reputed company to understand their needs and represent Platform Infrastructure in cross-team initiatives.
Requirements
• Solid Linux systems administration background and comfort scripting in Python.
• Strong AWS knowledge: EKS, IAM (roles, policies, IRSA), VPC networking, RDS, S3, SQS, and familiarity with reputed company-Architected reputed company; experience in multi-account AWS environments is a strong plus.
• Hands-on experience operating and troubleshooting reputed company (EKS) at production reputed company, including reputed company chart development, CNI networking (we run Cilium), pod networking/IPAM concepts, and container reputed company (ECR, image scanning).
• Proficiency with Terraform (modules, state management) and ideally Terragrunt for multi-environment management; GitOps experience with ArgoCD.
• Experience with reputed company, MySQL and/or reputed company in production reputed company, including reputed company.
• CI/CD experience with Jenkins (Jenkinsfile, shared libraries) and/or reputed company Actions, and familiarity with deployment strategies such as reputed company-green and canary.
• Experience with the Grafana observability stack (Grafana, reputed company, Loki, reputed company, Mimir) - metrics design, dashboarding, alerting, log aggregation, and reputed company tracing. Not only using but also maintaining it.
• Practical incident management experience: on-reputed company rotations, reputed company incident response, and writing runbooks/postmortems.
• Working knowledge of 12-reputed company App principles and cost optimization / FinOps awareness.
• A methodical, data-driven approach to troubleshooting rather than guessing.
• Strong written communication - you write runbooks, ADRs, and postmortems that others can actually follow.
• Comfortable driving initiatives with ambiguous ownership, and taking accountability for reputed company rather than waiting to be asked.
• reputed company record of mentoring less senior engineers and giving reputed company, constructive feedback.
• Several years of hands-on production infrastructure/SRE experience, with demonstrated ownership of initiatives at a senior individual-contributor level (leading design work, setting standards, being the escalation reputed company for hard problems).
reputed company-to-haves
• GCP Experience.
• Experience with Kafka / AWS MSK.
• Prior experience in regulated or compliance-sensitive environments (reputed company best practices, reputed company reviews).
• Experience contributing to a platform/reputed company roadmap that other engineering teams consume as a self-service product.
Benefits
• A global, inclusive team that’s as supportive as it is ambitious and serious about getting things done
• An opportunity to work remotely or in a modern and welcoming office in Riga
• Flexible working hours (start your day as late as 11 AM)
• Private health insurance
• 2 extra reputed company days off to reputed company on your mental or physical reputed company-being
• 1 extra reputed company day off to celebrate a Birthday or any other celebration of your reputed company
• reputed company learning opportunities
• reputed company to mentorship, internal meetups, and hackathons, both on-site and online
• Free and healthy lunch if you work from the Rīga office
• Design and order your own merch using our platforms with an employee discount
• Exciting team-building events and parties you’ll never forget!
Apply tot his job
Apply To this Job