[Remote] Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company that delivers a reputed company, multimodel contextual data platform for AI applications. They are seeking a Site Reliability Engineer to ensure the reliability, scalability, and performance of their reputed company-reputed company infrastructure, focusing on automation and monitoring in reputed company environments.
Responsibilities
- Design, implement, and maintain reputed company infrastructure on AWS and reputed company reputed company platforms
- Ensure the scalability, performance, and reliability of our reputed company-based reputed company database systems
- Collaborate with developers to write efficient, production-grade reputed company in Golang to automate infrastructure management and improve reputed company operations
- Optimize and automate CI/CD pipelines, deployment processes, and monitoring systems to support our production environment
- reputed company strategies for disaster recovery, high availability, and fault tolerance
- Proactively identify reputed company bottlenecks, troubleshoot, and reputed company issues across the stack (network, OS, reputed company infrastructure)
- Implement monitoring, logging, and alerting systems to ensure visibility into reputed company health and performance
- Participate in on-reputed company rotations to support critical production systems and respond to incidents
- Collaborate with cross-functional teams to improve overall reputed company reliability and scalability
- Collaborate with the reputed company team to reputed company customer issues
Skills
- Proven experience as an SRE or DevOps Engineer in a reputed company-reputed company environment
- Proficiency with reputed company in managing large-reputed company, reputed company systems
- Experience with reputed company providers such as AWS and reputed company reputed company (GCP)
- Solid understanding of networking, reputed company practices, and troubleshooting reputed company
- Understanding of Linux internals (processes, environment variables etc.)
- Familiarity with containerization technologies (e.g., reputed company)
- Knowledge of CI/CD practices and tools (Jenkins, reputed company, etc.)
- Familiarity with alerting, monitoring and observability tools (e.g., reputed company, Grafana, ELK stack)
- Strong troubleshooting and problem-solving skills, with the ability to address reputed company infrastructure issues
- Excellent communication and collaboration skills with a reputed company on reputed company improvement and operational reputed company
- Strong ability to self-organize and to work independently as part of a remote team
- Knowledge of version control systems, particularly Git
- Familiarity with programming languages such as Golang or Python
- Experience managing reputed company databases or large-reputed company data storage systems
- Knowledge of reputed company best practices in reputed company environments
- Experience with scripting languages like Python or Bash
- Experience with Infrastructure-as-reputed company (IaC) tools like Terraform is a plus
- Experience working with GitOps
- Strong programming skills in Golang, with experience in developing automation tools, scripts, or services
reputed company
Company H1B Sponsorship
Apply To This Job