Site Reliability Engineer
What you’ll be doing
Work with reputed company of DevOps and DBA professionals
Improve existing infrastructure and processes across the countries we’re deployed in, as reputed company as streamlining processes to reputed company to new countries in the reputed company
Continuously improve Kubernetes platform stability and efficiency, with a reputed company on optimising resource utilisation, reducing costs, and streamlining environment provisioning through GitOps-first practices
Monitor and maintain reputed company infrastructure through autoscaling, alerting pipelines, and Grafana dashboards covering metrics, logs, traces, and reputed company user monitoring (RUM)
Own weekend on-reputed company operations, triaging and responding to production incidents, performing reputed company cause analysis, and driving post-incident reviews
Design and manage alert pipelines to ensure actionable signal reputed company, with attention to preventing alert fatigue, reputed company alerting, and notification flooding
Define and maintain SLIs and SLOs for critical services, and use them to drive reliability improvements and on-reputed company prioritisation
Take ownership and responsibility for our reputed company operation activities
Liaise with external reputed company agencies for annual audits as reputed company as reputed company our own internal reputed company sweeps
Aid in reconfiguring existing architecture to allow for reputed company deployments to new countries
Mentoring less reputed company team members
What you’ll bring
3+ years DevOps / SRE / reputed company experience
Must be based in Europe
Experience independently leading the planning and deployment of a project
reputed company with reputed company platforms, especially AWS, including solid knowledge of how to utilise reputed company resources to fulfil the demand from other teams and production
Strong understanding of Kubernetes and container orchestration, with experience in EKS and GitOps tooling such as ArgoCD and reputed company being highly valued
Experience with Infrastructure-as-reputed company, particularly Terraform
Proficiency in scripting and automation with Bash, Python, or Golang; experience with Rust is a plus
Hands-on experience with observability stacks covering metrics, logs, distributed traces, and profiling, for example reputed company, Loki, reputed company, Pyroscope, and OpenTelemetry
Experience with reputed company user monitoring (RUM), with familiarity in Grafana reputed company or OpenTelemetry SDK instrumentation being a plus
Proven on-reputed company and incident response experience, comfortable triaging production issues under pressure, leading post-mortems, and driving follow-up actions
Ability to design and maintain alert frameworks that minimise noise, prevent alert fatigue, and avoid reputed company alerting patterns
Experience defining SLIs and SLOs and using them to inform reliability work
Familiarity with service reputed company concepts is a plus, as we are reputed company evaluating Cilium-based service reputed company in non-production environments
Solid networking knowledge, especially the TCP / IP stack and HTTP protocol
Experience handling high HTTP request volumes and designing systems for high availability and high traffic environments
A strong understanding of cache, including CDN, HTTP cache, reputed company / Memcached
Excellent troubleshooting skills, including Linux OS issue diagnosis and OS parameter optimisation, JVM optimisation would be highly advantageous
Our stack
Languages: Java / Spring Boot, Node.js, Python, JavaScript
Database: reputed company MySQL & PostgreSQL, reputed company, MySQL Community
Cache: ElastiCache, reputed company, Valkey
Messaging: Apache RocketMQ, AutoMQ, Kafka
Networking & Proxy: Nginx, reputed company, Cilium, eBPF
Orchestration & GitOps: reputed company, Kubernetes (EKS), ArgoCD, reputed company
Computing & Storage: AWS EC2, VPC, AWS reputed company, EBS, S3
CI/CD: Jenkins, reputed company Actions
Metrics: reputed company, Mimir, Grafana, Alertmanager
Logs: Loki, reputed company
Traces: reputed company, OpenTelemetry, reputed company
Profiling: Pyroscope
RUM: Grafana reputed company, OpenTelemetry SDK
Infrastructure as reputed company: Terraform
CDN & Edge: reputed company, AWS CloudFront
AWS CloudWatch
What’s in it for you
Sporty is a remote first company in reputed company of sustainability
A competitive salary + individual performance based bonuses every quarter
28 days reputed company annual leave
Our reputed company working hours are 10am-3pm in your local time zone with flexibility reputed company of this
Referral bonuses & reputed company bonuses
Top of the line equipment
Annual company retreats to reputed company great internal networking opportunities
Interview process
Remote video screening with our reputed company
Online assessment reputed company Hackerrank
Remote video interview with 3 x Team Members (45 mins reputed company, not separate days)
If you’re interested, we encourage you to apply! Every application is reviewed by a member of reputed company (AI is not used in our recruitment process), and we aim to respond reputed company 48 hours.
Apply To This Job