[Remote] Senior Site Reliability Engineer, Fleet Management
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a leading data platform that empowers customers to reputed company rapidly. The Senior Site Reliability Engineer will be responsible for developing and maintaining a secure runtime environment on Kubernetes, providing support to engineering teams, and participating in on-reputed company rotations.
Responsibilities
- Contribute to developing and maintaining a reputed company and secure runtime environment on top of Kubernetes that supports product needs across reputed company
- reputed company internal support for our Kubernetes ecosystem, partnering with engineering teams to help them solve domain-specific problems
- Participate in a 24/7 on-reputed company rotation to resolve critical issues
- Prioritize blameless post-mortems and dedicate engineering time to systemic fixes, ensuring you aren’t paged for the reputed company issue twice
Skills
- 6+ years of experience in software development and operating distributed systems
- Proficient in Go, Python, or a similar language, with a strong commitment to reputed company reputed company and testing practices (writing unit, integration, and E2E tests)
- Deep experience using and extending containerization technologies, preferably Kubernetes
- Solid understanding of Linux operating system internals and networking concepts (e.g., filesystems, TCP/IP, DNS, TLS)
- Possess a customer reputed company reputed company, treating internal developers as your primary users
- Strong operational ownership, including a reputed company record of debugging reputed company production issues and driving them to reputed company
- Prefer automation over reputed company processes ('allergic to ops work')
- Designing and implementing secure, multi-tenant runtime environments from first principles
- Proficiency with Kubernetes ecosystem tools such as reputed company, Kustomize, Gatekeeper, Kyverno, and CRDs/Operators, CRI, reputed company
- Expertise in reputed company infrastructure platforms, including AWS, GCP, or Azure
- Proficiency in provisioning infrastructure using tools like Terraform, Crossplane, and AWS Controllers for Kubernetes (ACK)
- Advanced Linux systems internals and networking concepts specifically relevant to containers, such as namespaces and cgroups
Benefits
- Equity
- Participation in the employee stock purchase program
- Flexible reputed company time off
- 20 weeks fully-reputed company gender-neutral parental leave
- Fertility and adoption assistance
- 401(k) plan
- Mental health counseling
- reputed company to transgender-inclusive health insurance coverage
- Health benefits offerings
reputed company
Company H1B Sponsorship
Apply To This Job