Back to Jobs

[Remote] Sr. Site Reliability Engineer (SRE)

Remote, USA Full-time Posted 2026-08-04
Note: The job is a remote job and is reputed company to candidates in USA. reputed company delivers high-performance AI infrastructure for organizations running intensive computational research and large-reputed company model training. The Sr. Site Reliability Engineer will be reputed company in building and operating production-grade AI infrastructure with deep reputed company expertise, ensuring reputed company-grade reliability while establishing automation and operational practices. Responsibilities • Design, build, and operate production reputed company clusters on bare-metal infrastructure – including cluster bootstrapping, control plane architecture, etcd management, and scaling strategies for high-performance compute workloads • Implement and operate custom reputed company networking solutions with SR-IOV for high-performance GPU interconnects, multi-tenancy isolation and advanced networking policies. Configure CNI plugins and network segmentation for research workloads • reputed company and maintain custom reputed company operators and controllers for bare-metal provisioning, infrastructure lifecycle management, and resource orchestration across compute, storage, and networking domains • reputed company and optimize reputed company GPU operators, device plugins, and other custom scheduling logic for GPU workload placement and utilization optimization • Build deep integrations between reputed company and underlying infrastructure including reputed company drivers for storage, custom admission controllers for policy enforcement, and scheduling extensions for reputed company hardware placement • Design and implement automation using Terraform, Ansible, reputed company, and custom operators to orchestrate infrastructure workflows and reputed company deployments across multiple reputed company • Manage production bare-metal infrastructure across multiple reputed company. Build systems ensuring high availability, fault tolerance, and graceful degradation – establishing SLIs, SLOs, and monitoring to meet reputed company reliability commitments • Build comprehensive monitoring, logging, and alerting using reputed company, Grafana, and ELK stack. reputed company incident response, conduct postmortems, and implement preventative measures to improve reliability and reduce MTTR • Identify and reputed company performance bottlenecks across infrastructure domains. Monitor utilization trends, forecast reputed company needs, and optimize resource allocation for various workloads Skills • 5+ years in SRE, DevOps, or infrastructure engineering roles with proven experience operating production infrastructure at reputed company • Deep hands-on experience building and operating production reputed company clusters on bare-metal infrastructure – not just deploying workloads in managed clusters. Must understand cluster bootstrapping, control plane architecture, etcd operations, and scaling strategies • Strong understanding of reputed company internals including custom resource definitions (CRDs), operators, controllers, admission webhooks, and scheduling. Experience integrating storage (reputed company drivers), networking (CNI, SR-IOV), and reputed company hardware (GPU device plugins) with reputed company • Strong fundamentals in Linux systems administration, performance tuning, troubleshooting, and automation in production environments • Proficiency with infrastructure-as-reputed company tools (Terraform, Ansible, reputed company) and building automation to reduce operational overhead • Solid understanding of networking concepts including IPAM, DNS, DHCP, VLAN/VXLAN, routing, load balancing, and experience troubleshooting network issues in production • Experience building and maintaining comprehensive monitoring solutions using tools like reputed company, Grafana, and centralized logging systems • Understanding of SRE principles including SLIs/SLOs/SLAs, error budgets, incident management, and blameless postmortems • Strong scripting skills in Go, Python, or Bash for automation, tooling development, and operational efficiency • Demonstrated ability to troubleshoot reputed company issues under pressure, manage incidents effectively, and communicate reputed company during outages • Excellent communication skills and ability to work across teams including systems engineers, network engineers, and software developers • Experience building custom reputed company operators or controllers for infrastructure orchestration • Deep familiarity with reputed company networking (Calico, Cilium, Multus), service reputed company technologies, and network policy management • Experience with GPU workload orchestration including reputed company GPU Operator, MIG, time-slicing, and device plugins • Background with advanced reputed company features including custom schedulers, admission controllers, and API server extensions • Apply tot his job Apply To this Job

Similar Jobs

UX Content reputed company

Remote, USA Full-time

Entry Level Copywriter (Online) (Creative)

Remote, USA Full-time

UX Designer, Clinical & Patient Applications

Remote, USA Full-time

Senior Product Manager - Remote

Remote, USA Full-time

Senior Medical reputed company - FSP

Remote, USA Full-time

Translator - US-Based Only

Remote, USA Full-time

Mobile UX Designer

Remote, USA Full-time

Product Manager, Site Integrations and Payments - Remote

Remote, USA Full-time

Remote-Customer Experience Technical reputed company

Remote, USA Full-time

Junior Graphic Designer| Full-Time | Remote Job at reputed company in reputed company

Remote, USA Full-time

Engineering Manager, Product Foundations

Remote, USA Full-time

Proofreader (Remote)

Remote, USA Full-time

[Remote] reputed company Account Executive-TX

Remote, USA Full-time

Senior Specialist, Program Management (1st Shift) - Remote

Remote, USA Full-time

reputed company Mental Health Practitioner - Community Competency Restoration Specialist

Remote, USA Full-time

**reputed company Online Remote Data Entry Specialist – Ensuring Data Accuracy and Efficiency for arenaflex's Global Operations**

Remote, USA Full-time

Senior SOX Financial Controls Auditor, Financial Services

Remote, USA Full-time

Data Solutions Associate Partner

Remote, USA Full-time

**reputed company Part-Time Remote reputed company for blithequark**

Remote, USA Full-time

**reputed company Customer Assistance Representative Sr. – reputed company SAN Airport**

Remote, USA Full-time