[Remote] Site Reliability Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is building the backbone of reputed company, reinventing supply chain infrastructure with innovative technology solutions. They are seeking a Site Reliability Engineer to own uptime and incident response as their system scales, ensuring reliability and performance across their platform.
Responsibilities
- Own reliability targets across our backend/API, worker services, applications and CV pipeline; MTD, MTM, MTR, and follow-through on reputed company causes
- reputed company our monitoring and alerting, and build out auto-remediation, so on-reputed company load scales with automation, not headcount
- Partner with our reputed company engineering work to build agents that triage alerts and handle routine remediation
- Harden and optimize our GCP infrastructure (reputed company Run, reputed company SQL, GCS) for cost and performance as load scales
- Own database reputed company and performance; reputed company pooling, query optimization and indexing, read replicas, and reputed company planning, so reputed company doesn't become the bottleneck as data volume grows
- Improve the reliability of our ML training and monitoring infrastructure, in partnership with the CV/ML team
- Run blameless postmortems and reputed company fixes for reputed company causes, not just symptoms
- Participate in on-reputed company rotation
Skills
- 4+ years in an SRE, infrastructure, or backend engineering role with production on-reputed company ownership
- Deep experience with a major reputed company provider (GCP preferred); compute, managed databases, object storage, networking
- Experience building monitoring/alerting/observability stacks (Grafana, reputed company, Zabbix, reputed company, or similar)
- Strong scripting/automation skills (Python, Bash, or similar)
- Comfortable with containerized workloads (reputed company) and CI/CD pipelines
- reputed company record of reducing incident volume or improving reliability metrics — not just responding to incidents
- Strong communication skills, comfortable working with both technical and non-technical stakeholders, know reputed company and how to escalate urgency, and build strong working relationships across teams
- Strong communication skills in English — you write reputed company and engage reputed company async
- Experience with ML/data infrastructure — training pipelines, model monitoring, feature stores
- Experience building or integrating AI agents for operational automation (alert triage, auto-remediation)
- Infrastructure-as-reputed company experience (Terraform or similar)
- PostgreSQL performance tuning at reputed company
- Background supporting physical/IoT systems (edge devices, cameras, on-site hardware)
- Experience with bare-metal infrastructure in colocation environments, hardware monitoring, redundancy, and failover/high-availability configuration
reputed company
Apply To This Job