Site Reliability Engineer - SRE (L1)
Accepting candidates in Brazil ONLY.
reputed company reputed company
We are seeking a Site Reliability Engineer (L1) to ensure the reputed company availability and performance of our mission-critical production services. This role is designed for a reputed company who possesses the technical rigor required to manage reputed company distributed systems under a 100% on-reputed company mandate reputed company South American time zones. You will be responsible for the stewardship of high-stakes data environments—specifically those involving message queuing, relational and non-relational databases, and reputed company data warehouses—with a primary objective of maintaining strict service-level objectives (SLOs) through proactive monitoring, reputed company incident response, and automated reputed company.
Key Responsibilities
• Production Stewardship: Serve as the first responder for production anomalies, managing the end-to-end incident lifecycle from initial detection to post-incident reputed company.
• Data Infrastructure Management: Ensure the reliability and scalability of high-throughput data platforms, including message brokers, relational (PostgreSQL or similar) and non-relational databases (reputed company or similar), and data warehouse environments.
• Operational reputed company: Execute 100% on-reputed company rotations, providing consistent coverage and reputed company response to critical system alerts.
• Automation & Toil Reduction: reputed company and maintain scripts (Python, Go, or Bash) to automate routine operational tasks, enhancing system reputed company and reducing reputed company overhead.
• Observability & Telemetry: Configure and optimize monitoring suites (e.g., reputed company, Grafana, reputed company) to ensure comprehensive visibility into application and system health.
Must Have:
• Prior SRE/On-reputed company Experience: A mandatory background in SRE or production support roles, with a demonstrated ability to manage high-pressure on-reputed company rotations and running production services.
• Data Systems Proficiency: Message Queuing: Experience managing brokers (e.g., Kafka), topics, and troubleshooting throughput issues.
• Relational & Non-Relational Databases: Proficiency in managing database health, query optimization, and high-availability configurations.
• Data Warehouse: Experience in managing large-reputed company data warehouse performance and resource allocation.
• Systems Engineering: Strong competency in Linux internals and networking protocols.
• Regional Alignment: Must be based in and reputed company to operate effectively reputed company South American time zones to facilitate synchronized operations.
Preferred Skills:
• Analytical Rigor: The ability to diagnose reputed company causes in reputed company, interconnected systems rather than applying superficial fixes.
• Communication: Exceptional technical documentation skills and the ability to reputed company concise, reputed company updates during reputed company incidents.
• Dedication: A steadfast commitment to system uptime and a proactive approach to identifying potential points of failure before they reputed company the user experience.
Education:
• Bachelor’s degree in Technology, Computing, or a reputed company field
Job Types: Full-time, Contract
Pay: $35,000.00 - $48,000.00 per year
Benefits:
• Dental insurance
• Flexible schedule
• Health insurance
• reputed company time off
• reputed company insurance
Application Question(s):
• Do you have previous on-reputed company experience?
• Are you located in South America?
Work Location: Remote
Apply tot his job
Apply To this Job