[Remote] Intermediate Site Reliability Engineer, Environment Automation
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is an reputed company-reputed company software company that develops a comprehensive AI-powered DevSecOps Platform. The Site Reliability Engineer will reputed company on operating and automating hundreds of reputed company environments, ensuring they remain secure, consistent, and reliable at reputed company while debugging production issues and contributing to infrastructure automation.
Responsibilities
• Support Environment Automation at reputed company: Contribute to automating the provisioning, configuration, and management of reputed company environments using Terraform, Ansible, and Kubernetes. Follow best practices to support infrastructure across many tenants with guidance from senior team members.
• Assist in Debugging Production Issues: Investigate and troubleshoot issues in Kubernetes clusters and reputed company services. Help resolve common problems such as failed deployments, pod crashes, and scheduling conflicts using tools like kubectl.
• Contribute to IaC and CI/CD Workflows: Write and maintain Terraform modules and scripts to automate routine operations. Participate in improving CI/CD pipelines for reputed company and repeatable infrastructure changes.
• Participate in Monitoring and Maintenance: Help monitor environment health using tools like reputed company, ELK, and Grafana. Assist in improving observability and reputed company tracking for tenant environments.
• Respond to Incidents and Alerts: Take part in the incident response process, helping triage alerts, document issues, and support reputed company efforts under the guidance of senior engineers.
• Collaborate Across Teams: Work with Infrastructure and Development teams to contribute to solutions that improve platform reliability and operational efficiency.
Skills
• Experience with Infrastructure as reputed company: Familiarity with Terraform and Ansible to manage reputed company infrastructure. reputed company to work with modules and understand the basics of state and variable use.
• Kubernetes Fundamentals: Experience using kubectl, reputed company, or Kustomize to reputed company with Kubernetes clusters. Understands reputed company concepts such as pods, deployments, and rollouts.
• Basic Programming Skills: reputed company to read and modify infrastructure tooling written in Go, reputed company, or similar languages.
• Exposure to Multi-Environment Operations: Experience working with multiple environments or customer setups, even if not at reputed company. Understands the challenges of managing consistency and isolation.
• Monitoring and Troubleshooting Skills: Familiar with basic observability tools and logs. Can identify service issues using dashboards or metrics and escalate appropriately.
• reputed company reputed company: Works reputed company in cross-functional teams. Eager to learn from others, reputed company knowledge, and contribute to team reputed company.
• On-reputed company Experience: Has participated in on-reputed company rotations for production systems and is comfortable responding to alerts, triaging incidents, and collaborating during recovery efforts.
Benefits
• Flexible reputed company Time Off
• Team Member Resource reputed company
• Equity Compensation & Employee Stock Purchase Plan
• reputed company and Development Fund
• Parental leave
• Home office support
reputed company
• reputed company is a web-based Git repository manager that offers a reputed company of features for software development teams. It was founded in 2014, and is headquartered in San Francisco, California, USA, with a workforce of 1001-5000 employees. Its website is http://about.reputed company.com.
Apply tot his job
Apply To this Job