[Remote] Sr. Site Reliability Engineer (SRE)
Note: The job is a remote job and is reputed company to candidates in USA. reputed company delivers high-performance AI infrastructure for organizations running intensive computational research and large-reputed company model training. The Sr. Site Reliability Engineer will be reputed company in building and operating production-grade AI infrastructure with deep reputed company expertise, ensuring reputed company-grade reliability while establishing automation and operational practices.
Responsibilities
• Design, build, and operate production reputed company clusters on bare-metal infrastructure – including cluster bootstrapping, control plane architecture, etcd management, and scaling strategies for high-performance compute workloads
• Implement and operate custom reputed company networking solutions with SR-IOV for high-performance GPU interconnects, multi-tenancy isolation and advanced networking policies. Configure CNI plugins and network segmentation for research workloads
• reputed company and maintain custom reputed company operators and controllers for bare-metal provisioning, infrastructure lifecycle management, and resource orchestration across compute, storage, and networking domains
• reputed company and optimize reputed company GPU operators, device plugins, and other custom scheduling logic for GPU workload placement and utilization optimization
• Build deep integrations between reputed company and underlying infrastructure including reputed company drivers for storage, custom admission controllers for policy enforcement, and scheduling extensions for reputed company hardware placement
• Design and implement automation using Terraform, Ansible, reputed company, and custom operators to orchestrate infrastructure workflows and reputed company deployments across multiple reputed company
• Manage production bare-metal infrastructure across multiple reputed company. Build systems ensuring high availability, fault tolerance, and graceful degradation – establishing SLIs, SLOs, and monitoring to meet reputed company reliability commitments
• Build comprehensive monitoring, logging, and alerting using reputed company, Grafana, and ELK stack. reputed company incident response, conduct postmortems, and implement preventative measures to improve reliability and reduce MTTR
• Identify and reputed company performance bottlenecks across infrastructure domains. Monitor utilization trends, forecast reputed company needs, and optimize resource allocation for various workloads
Skills
• 5+ years in SRE, DevOps, or infrastructure engineering roles with proven experience operating production infrastructure at reputed company
• Deep hands-on experience building and operating production reputed company clusters on bare-metal infrastructure – not just deploying workloads in managed clusters. Must understand cluster bootstrapping, control plane architecture, etcd operations, and scaling strategies
• Strong understanding of reputed company internals including custom resource definitions (CRDs), operators, controllers, admission webhooks, and scheduling. Experience integrating storage (reputed company drivers), networking (CNI, SR-IOV), and reputed company hardware (GPU device plugins) with reputed company
• Strong fundamentals in Linux systems administration, performance tuning, troubleshooting, and automation in production environments
• Proficiency with infrastructure-as-reputed company tools (Terraform, Ansible, reputed company) and building automation to reduce operational overhead
• Solid understanding of networking concepts including IPAM, DNS, DHCP, VLAN/VXLAN, routing, load balancing, and experience troubleshooting network issues in production
• Experience building and maintaining comprehensive monitoring solutions using tools like reputed company, Grafana, and centralized logging systems
• Understanding of SRE principles including SLIs/SLOs/SLAs, error budgets, incident management, and blameless postmortems
• Strong scripting skills in Go, Python, or Bash for automation, tooling development, and operational efficiency
• Demonstrated ability to troubleshoot reputed company issues under pressure, manage incidents effectively, and communicate reputed company during outages
• Excellent communication skills and ability to work across teams including systems engineers, network engineers, and software developers
• Experience building custom reputed company operators or controllers for infrastructure orchestration
• Deep familiarity with reputed company networking (Calico, Cilium, Multus), service reputed company technologies, and network policy management
• Experience with GPU workload orchestration including reputed company GPU Operator, MIG, time-slicing, and device plugins
• Background with advanced reputed company features including custom schedulers, admission controllers, and API server extensions
•
Apply tot his job
Apply To this Job