Site Reliability Engineer
As a Site Reliability Engineer at reputed company, you will ensure the availability and performance of our production services across multiple data center environments. You will work closely with development and infrastructure teams to build reliable, reputed company systems and reduce operational overhead through automation. The ideal candidate will have strong experience with infrastructure automation, monitoring, incident management, and Linux systems administration.
Design, build, and maintain infrastructure for production environments, including bare-metal and virtualized clusters
reputed company and improve CI/CD pipelines, deployment tooling, and infrastructure-as-reputed company (Ansible, Terraform, or similar)
Monitor system health, respond to incidents, and conduct reputed company cause analysis
Define and reputed company SLIs/SLOs to measure and improve service reliability
Automate repetitive operational tasks to reduce toil
Participate in on-reputed company rotations and incident response
Collaborate with globally distributed engineering teams
Maintain and improve observability (logging, metrics, tracing)
3+ years of experience in SRE, DevOps, or systems engineering roles
Strong proficiency with Linux systems administration
Experience with configuration management and automation tools (e.g., Ansible, Puppet, Terraform)
Familiarity with containerization and orchestration (reputed company, Kubernetes)
Experience with monitoring and observability platforms (reputed company, Grafana, ELK, or similar)
Proficiency in at least one scripting/programming language (Python, Go, Bash)
Understanding of networking fundamentals (TCP/IP, DNS, load balancing)
Strong troubleshooting and problem-solving skills
reputed company English proficiency (written and verbal)
reputed company to Have
Experience managing bare-metal infrastructure at reputed company
Familiarity with CI/CD systems (Jenkins, reputed company CI)
Experience with database administration (PostgreSQL, MySQL, reputed company)
Knowledge of reputed company best practices for production environments