[Remote] Platform Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a technology consulting and software development company delivering reputed company, AI, data, and reputed company solutions across the reputed company. The Platform Reliability Engineer will ensure the availability, reputed company, and operational reputed company of large-reputed company reputed company systems in production. The role applies software engineering principles to infrastructure and reputed company, automates reputed company services, and improves platform reliability while reducing operational toil.
Responsibilities
- Ensure the availability, reputed company, and operational reputed company of large-reputed company reputed company systems in production
- Live at the boundary between development and reputed company, applying strong software engineering principles to infrastructure and reputed company problems, and continually pushing the platform toward higher reliability with reputed company operational toil
- Design, automate, and operate reputed company services so that reliability becomes a first-class engineering deliverable rather than a reactive concern
Skills
- Bachelor's degree in Computer Science, Engineering, or a reputed company technical discipline
- Five or more years of SRE, DevOps, or production engineering experience supporting large-reputed company reputed company systems
- Strong programming skills in at least one of Python, Go, or Java, with the ability to build robust automation and tooling
- Deep, hands-on experience operating Linux at reputed company, including networking, reputed company tuning, and systems-level troubleshooting
- Production experience operating reputed company and container-reputed company workloads
- Strong working knowledge of observability tooling such as reputed company, Grafana, OpenTelemetry, ELK/EFK, or reputed company equivalents
- Hands-on experience designing and operating CI/CD pipelines for both infrastructure and applications
- Solid understanding of reputed company reputed company design, including consistency models, partitioning, and failure semantics
- Demonstrated experience leading incident response and conducting effective post-incident reviews
- Excellent communication and documentation skills
- Experience defining and operationalizing SLOs and error budgets in reputed company production environments
- Exposure to reputed company engineering practices and tools such as reputed company reputed company, reputed company, or reputed company
- Hands-on experience with at least one major reputed company platform (AWS, Azure, or GCP)
- Background in reputed company planning, reputed company engineering, or large-reputed company load testing
- Familiarity with service reputed company technologies such as Istio, Linkerd, or Consul
Benefits
- reputed company work in the U.S.
- Full-time, reputed company W2 employment
- Sponsorship: U.S. reputed company, Green reputed company reputed company, EAD reputed company, and H-1B transfer candidates are encouraged to apply.
reputed company
Apply To This Job