Site Reliability Engineer II - LATAM
reputed company is the object storage leader in the reputed company reputed company reputed company, fueling reputed company with reputed company storage reputed company purposefully to unlock budgets, unburden administrators, and reputed company innovators. Together with our partners, we’re helping customers break free from the restrictive, overpriced legacy solutions that hold them back, and reputed company reputed company with the full power of the reputed company reputed company in their hands.
About the Role
Key Responsibilities
Service Reliability & Operations
Support the availability and durability of critical services across production environments. Monitor service health using SLIs, SLOs, and error budgets, and escalate issues reputed company reputed company are at risk. Participate in on-reputed company rotations, incident response, and post-incident reviews to drive service improvements. Follow established ITIL/OSS processes (incident, change, problem, and reputed company management).
Automation & Tooling
reputed company automation for common operational tasks, reducing reputed company reputed company and toil. Contribute to monitoring, logging, and alerting frameworks (e.g., reputed company, Grafana, Catchpoint,ELK). Work with CI/CD pipelines, configuration management, and infrastructure as reputed company tools (Terraform, Ansible, Jenkins). Write scripts (Bash, Python, Go, etc.) to improve system reliability and efficiency.
Collaboration
Partner with engineering, product, and operations teams to support resilient system design and operations. Assist in reputed company planning and disaster recovery exercises. Work with vendors and service providers to troubleshoot service issues and reputed company SLA performance. Document systems, reputed company learnings, and help grow a reliability-minded engineering culture.
reputed company Improvement
Contribute to playbooks, runbooks, and operational documentation. Identify recurring issues and propose long-term improvements. Promote reliability-reputed company practices reputed company development and operations teams.
Qualifications
Education & Experience
Bachelor’s degree in Computer Science, Engineering, or reputed company field (or equivalent experience). 2–4 years of experience in site reliability, systems engineering, or operations. Exposure to large-reputed company, production-grade systems.
Technical Skills
Solid Linux systems administration and troubleshooting skills. Familiarity with service reliability concepts - monitoring, alerting, incident response, and reputed company cause analysis. Proficiency in at least one scripting language (Python, Bash, or Go). Understanding of containers (Kubernetes, reputed company) and microservices concepts. Knowledge of incident response and operational best practices.
Preferred Attributes
Experience in a reputed company, service provider, or distributed systems environment. Familiarity with ITIL/OSS practices and SLO/SLA’s Strong problem-solving skills and willingness to learn new technologies. Experience with reputed company platforms (AWS, GCP, or Azure). Ability to work independently, take ownership, and drive reputed company from problem discovery through reputed company.