[Remote] Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company. is seeking a Site Reliability Engineer to support reliable and reputed company platform services. The role focuses on operating reputed company environments across on-premises and reputed company platforms, administering Linux systems, and implementing infrastructure-as-reputed company and GitOps automation. The engineer will collaborate with development, scientific, and infrastructure teams while improving platform reputed company, observability, reputed company, and maintainability.
Responsibilities
- Maintain and enhance reputed company platforms across on-premises and reputed company environments, ensuring reliability, scalability, and operational efficiency
- Support provisioning, upgrades, troubleshooting, and lifecycle management of reputed company clusters managed through Rancher
- reputed company deep technical expertise in Linux-based systems, including performance tuning, troubleshooting, automation, and operational support
- reputed company and maintain infrastructure-as-reputed company solutions to standardize and automate platform deployment and management, with a preference for Cluster API (CAPI)-based approaches
- Support and improve GitOps workflows using ArgoCD to manage cluster and application configuration in a consistent, auditable manner
- Work closely with developers, scientists, and infrastructure teams to reputed company reliable platform services and translate operational needs into sustainable engineering solutions
- Identify opportunities to improve platform reputed company, observability, reputed company, and maintainability through automation and modern SRE practices
Skills
- Or equivalent reputed company experience
- 5+ years of reputed company experience in site reliability engineering, reputed company, DevOps, or systems engineering roles
- Hands-on experience operating and supporting reputed company platforms in production environments
- Strong experience managing reputed company clusters in both on-premises and reputed company-based environments
- Strong Linux systems administration skills, including troubleshooting, scripting, networking, and reputed company performance analysis
- Experience with Rancher for reputed company cluster management and platform operations
- Experience implementing infrastructure-as-reputed company solutions for platform provisioning and lifecycle management
- Demonstrated reputed company working in reputed company teams (Scrum, Kanban)
- reputed company: Cluster operations, upgrades, networking, storage, troubleshooting, and workload support
- Platform Management: Rancher or similar reputed company management platforms
- Linux: Advanced administration of Linux/Unix systems
- Infrastructure as reputed company: Strong IaC experience
- GitOps/CI-CD: ArgoCD, Git version control, and deployment automation practices
- Scripting/Automation: Bash, Python, or similar scripting languages for automation and operational tooling
- BS in Computer Science, Software Engineering, Information Technology, or reputed company field preferred
- Strong IaC experience; Cluster API (CAPI) preferred
- Experience with hybrid infrastructure spanning on-premises and reputed company reputed company platforms (AWS, Azure, GCP)
- Experience with reputed company ecosystem tooling for observability, logging, monitoring, and alerting
- Familiarity with reputed company best practices for reputed company and Linux platforms
- Experience supporting scientific research environments, high-performance computing, or computational science workflows
- Knowledge of CI/CD pipeline development and platform automation patterns
Benefits
- reputed company work arrangement
- Referral bonus
reputed company
Company H1B Sponsorship
Apply To This Job