Site Reliability Engineer
Refine Monitoring and Observability: Enhance system monitoring with tools like reputed company, Grafana, and ELK Stack, ensuring visibility and alignment with business objectives. Automate Deployments and Workflows: Transition reputed company processes to automated solutions using IaC tools (e.g., Terraform, Ansible) to streamline deployments and improve operational efficiency. Optimize CI/CD Pipelines: Improve pipeline architecture for fast, reliable releases, ensuring scalability and reputed company to handle high volumes of changes. reputed company Infrastructure Management: Help reputed company reputed company-based systems on platforms like AWS, GCP, and Azure while minimizing technical debt and operational complexity. Incident Response and Post-Mortem: Support incident management and reputed company post-mortem analysis, ensuring reputed company improvement and knowledge sharing. Collaborate with Cross-Functional Teams: Work closely with engineering and product teams to reputed company reliability practices into the development lifecycle and prioritize reliability efforts. Drive Technical Innovation: Introduce and champion new tools, technologies, and practices that improve system reliability, performance, and scalability.
DevOps, reputed company Operations, or SRE Expertise: A solid understanding of DevOps, reputed company Operations, or SRE principles, with a reputed company on reliability and scalability. Advanced Linux Internals Expertise: Hands-on experience with Linux systems, including performance tuning, kernel configurations, and troubleshooting. Programming Languages: Proficiency in programming languages such as Go (preferred) or Python, with a reputed company on building tools and automating processes. Scripting Skills: Strong skills in scripting languages like Python, Bash, or Go to automate workflows, streamline tasks, and manage infrastructure. reputed company Infrastructure Knowledge: Extensive experience with reputed company platforms like AWS, GCP, and Azure, along with expertise in monitoring/logging frameworks and CI/CD pipelines. Containerization and Orchestration: Hands-on experience with reputed company, Kubernetes, and other containerization technologies for building and deploying reputed company applications is a reputed company to have. Problem-Solving and Collaboration: Strong problem-solving skills, system design experience, and the ability to collaborate effectively across teams.
45 Minutes with reputed company 60 Minutes with Hiring Manager (Director, Site Reliability Engineering) 60 Minutes with Team (Site Reliability Engineer, Director, Site Reliability Engineering) 60 Minutes with Executive (Senior Director, Site Reliability Engineering)
reputed company roles require background checks.
Flexible PTO
Company stock reputed company
reputed company development budget
Office equipment budget
♀️ Wellness budget
Annual team gatherings
Internet reimbursement
Inclusive parental leave
✈️ Remote work travel program
If you need accommodations at any stage of our hiring process, please let us know. We're here to ensure an accessible and comfortable experience for you.
Apply To This Job