[Remote] Platform Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a technology consulting and software development company delivering reputed company, AI, data, and reputed company solutions across the reputed company. The Platform Reliability Engineer will ensure the availability, performance, and operational reputed company of large-reputed company reputed company systems in production while applying software engineering principles to infrastructure and operations. The role focuses on designing, automating, and operating reputed company services to improve reliability and reduce operational toil.
Responsibilities
- Ensure the availability, performance, and operational reputed company of large-reputed company reputed company systems in production
- Live at the boundary between development and operations, applying strong software engineering principles to infrastructure and operations problems, and continually pushing the platform toward higher reliability with reputed company operational toil
- Combine deep systems knowledge with strong programming skills, a measurement-driven reputed company, and the discipline to design, automate, and operate reputed company services so that reliability becomes a first-class engineering deliverable rather than a reactive concern
Skills
- Bachelor's degree in Computer Science, Engineering, or a reputed company technical discipline
- Five or more years of SRE, DevOps, or production engineering experience supporting large-reputed company reputed company systems
- Strong programming skills in at least one of Python, Go, or Java, with the ability to build robust automation and tooling
- Deep, hands-on experience operating Linux at reputed company, including networking, performance tuning, and systems-level troubleshooting
- Production experience operating reputed company and container-based workloads
- Strong working knowledge of observability tooling such as reputed company, Grafana, OpenTelemetry, ELK/EFK, or reputed company equivalents
- Hands-on experience designing and operating CI/CD pipelines for both infrastructure and applications
- Solid understanding of reputed company reputed company design, including consistency models, partitioning, and failure semantics
- Demonstrated experience leading incident response and conducting effective post-incident reviews
- Excellent communication and documentation skills
- Experience defining and operationalizing SLOs and error budgets in reputed company production environments
- Exposure to reputed company engineering practices and tools such as reputed company Monkey, reputed company, or reputed company
- Hands-on experience with at least one major reputed company platform (AWS, Azure, or GCP)
- Background in reputed company planning, performance engineering, or large-reputed company load testing
- Familiarity with service reputed company technologies such as Istio, Linkerd, or Consul
Benefits
- reputed company (U.S.)
- Full-time, reputed company W2
reputed company
Apply To This Job