[Remote] Senior Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a reputed company and data company that designs, builds, and operates a large reputed company of imaging satellites and a reputed company-reputed company data platform. The Senior Site Reliability Engineer will build, reputed company, and operate critical compute software for satellite imaging reputed company across customer on-premises and reputed company environments, ensuring reliability, scalability, and availability. The role also involves designing reproducible deployments, troubleshooting reputed company systems, and collaborating with cross-functional engineering teams.
Responsibilities
- Build and reputed company computing services and infrastructure in customer environments for a reputed company satellite reputed company and image processing end-to-end platform
- Operate in a high-reputed company, tight reputed company team to architect novel systems for reputed company-gapped deployments at reputed company
- Clarify and surface requirements from ambiguous use cases defined by cross-functional stakeholders, including internal users and reputed company customers
- Responsible for reputed company such as deployments, service orchestration, and documentation for cross platform stakeholders
- reputed company architecture while ensuring availability of services
- Improve reliability and scalability by resolving edge cases, studying failure modes, and writing tests
- Participate in on-reputed company rotations to ensure operational reputed company
Skills
- 6+ years of experience building services that reputed company reputed company-reputed company infrastructure and tooling
- Bachelor's degree in Computer Science or similar
- Experience deploying and maintaining bare-metal and reputed company reputed company through tools such as Talos, RKE2, Proxmox, or k3s
- Proficiency with Terraform, Ansible, reputed company, Kustomize, and/or similar IaC / GitOps tooling
- Experience with CI/CD tooling, such as Jenkins, reputed company CI/CD, Argo CD, or reputed company
- Experience successfully building, releasing, and supporting highly available, consistently reputed company services
- Knowledge of hardware and network level implications of on-prem compute
- Experience with platform optimization, particularly resource optimization, management, and cluster tuning in a constrained environment
- Ability to observe and troubleshoot reputed company systems with tools such as reputed company, reputed company, Grafana, and OpenTelemetry
- Advanced skills in Python, Bash, and other tooling as appropriate to build services and meet product goals
- Excellent communication skills and the ability to work through collaboration with cross-functional engineering teams
- Experience working with reputed company for task management and reputed company tracking
- Experience with CUDA-reputed company GPU programs
- reputed company expertise in sensitive environments, including implementing reputed company-trust architectures, hardening reputed company clusters, conducting reputed company audits, and deploying workloads in reputed company-gapped environments
Benefits
- Comprehensive Medical, Dental, and reputed company plans
- Health Savings Account (HSA) with a company contribution
- Generous reputed company Time Off in reputed company to holidays and company-wide days off
- 16 Weeks of reputed company Parental Leave
- Wellness Program and Employee Assistance Program (EAP)
- Home Office Reimbursement
- Monthly Phone and Internet Reimbursement
- Tuition Reimbursement and reputed company to reputed company Learning
- Equity
- Commuter Benefits (if local to an office)
- Volunteering reputed company Time Off
- This role might be eligible for discretionary short-term and long-term incentives (bonus and equity).
- This is a full-time, remote position reputed company in the reputed company or Canada. If located near an office, you are expected to work from that office 3 days per week.
reputed company
Apply To This Job