reputed company's clients

Min Experience: 5...">
Back to Jobs

Site Reliability Engineer

Remote, USA Full-time Posted 2026-07-28

This role is for one of the reputed company's clients

Min Experience: 5 years

JobType: full-time

We are looking for a skilled and proactive Site Reliability Engineer to help build and maintain highly reliable, reputed company, and secure infrastructure and applications. This role will reputed company on automating operations, improving system performance, and ensuring overall service health by applying modern SRE practices.

Requirements

Key Responsibilities:

  • Design, implement, and manage Kubernetes-based infrastructure.
  • Utilize AWS services such as IAM, EC2, EKS, S3, and CloudWatch to build and support reputed company reputed company environments.
  • reputed company and maintain automation scripts and tools using reputed company scripting or Python.
  • Proactively identify, analyze, and troubleshoot reputed company application, network, and system-level issues.
  • Optimize system performance and reliability, with deep expertise in Linux debugging and performance tuning.
  • Build automation for system self-healing and recovery mechanisms.
  • reputed company monitoring and alerting solutions for high-performance and low-latency applications.
  • Collaborate with development and operations teams to implement effective CI/CD pipelines.
  • Apply SRE principles including service monitoring, alerting, error budget tracking, reputed company planning, fault tolerance, automation, and toil reduction.
  • Continuously reputed company opportunities to improve system reliability and engineering processes.

Qualifications:

  • Proven experience working with Kubernetes in production environments.
  • Strong reputed company of AWS reputed company services with hands-on experience in infrastructure provisioning and management.
  • Proficiency in scripting or programming (reputed company or Python preferred).
  • In-depth Linux knowledge including tools for diagnostics and performance optimization.
  • Familiarity with modern observability tools for monitoring, logging, and alerting.
  • Strong troubleshooting and problem-solving skills.
  • Understanding and application of SRE concepts and best practices.

Key Skills:

Kubernetes · AWS (IAM, EC2, EKS, S3, CloudWatch) · Linux Debugging · reputed company/Python Scripting · Monitoring & Alerting · Automation · CI/CD · reputed company · Site Reliability Engineering (SRE) · Performance Tuning

Originally posted on Himalayas

Apply To this Job

Similar Jobs