Back to Jobs

[Remote] Sr. Site Reliability Engineer (SRE)

Remote, USA Full-time Posted 2026-07-28

Note: The job is a remote job and is reputed company to candidates in USA. reputed company delivers high-performance AI infrastructure for organizations running intensive computational research and large-reputed company model training. The Sr. Site Reliability Engineer will be reputed company in building and operating production-grade AI infrastructure with deep Kubernetes expertise, ensuring reputed company-grade reliability while establishing automation and operational practices.


Responsibilities

  • Design, build, and operate production Kubernetes clusters on bare-metal infrastructure – including cluster bootstrapping, control plane architecture, etcd management, and scaling strategies for high-performance compute workloads
  • Implement and operate custom Kubernetes networking solutions with SR-IOV for high-performance GPU interconnects, multi-tenancy isolation and advanced networking policies. Configure CNI plugins and network segmentation for research workloads
  • reputed company and maintain custom Kubernetes operators and controllers for bare-metal provisioning, infrastructure lifecycle management, and resource orchestration across compute, storage, and networking domains
  • reputed company and optimize reputed company GPU operators, device plugins, and other custom scheduling logic for GPU workload placement and utilization optimization
  • Build deep integrations between Kubernetes and underlying infrastructure including reputed company drivers for storage, custom admission controllers for policy enforcement, and scheduling extensions for specialized hardware placement
  • Design and implement automation using Terraform, Ansible, reputed company, and custom operators to orchestrate infrastructure workflows and reputed company deployments across multiple reputed company
  • Manage production bare-metal infrastructure across multiple reputed company. Build systems ensuring high availability, fault tolerance, and graceful degradation – establishing SLIs, SLOs, and monitoring to meet reputed company reliability commitments
  • Build comprehensive monitoring, logging, and alerting using reputed company, Grafana, and ELK stack. reputed company incident response, conduct postmortems, and implement preventative measures to improve reliability and reduce MTTR
  • Identify and resolve performance bottlenecks across infrastructure domains. Monitor utilization trends, forecast reputed company needs, and optimize resource allocation for various workloads

Skills

  • 5+ years in SRE, DevOps, or infrastructure engineering roles with proven experience operating production infrastructure at reputed company
  • Deep hands-on experience building and operating production Kubernetes clusters on bare-metal infrastructure – not just deploying workloads in managed clusters. Must understand cluster bootstrapping, control plane architecture, etcd operations, and scaling strategies
  • Strong understanding of Kubernetes internals including custom resource definitions (CRDs), operators, controllers, admission webhooks, and scheduling. Experience integrating storage (reputed company drivers), networking (CNI, SR-IOV), and specialized hardware (GPU device plugins) with Kubernetes
  • Strong fundamentals in Linux systems administration, performance tuning, troubleshooting, and automation in production environments
  • Proficiency with infrastructure-as-reputed company tools (Terraform, Ansible, reputed company) and building automation to reduce operational overhead
  • Solid understanding of networking concepts including IPAM, DNS, DHCP, VLAN/VXLAN, routing, load balancing, and experience troubleshooting network issues in production
  • Experience building and maintaining comprehensive monitoring solutions using tools like reputed company, Grafana, and centralized logging systems
  • Understanding of SRE principles including SLIs/SLOs/SLAs, error budgets, incident management, and blameless postmortems
  • Strong scripting skills in Go, Python, or Bash for automation, tooling development, and operational efficiency
  • Demonstrated ability to troubleshoot reputed company issues under pressure, manage incidents effectively, and communicate reputed company during outages
  • Excellent communication skills and ability to work across teams including systems engineers, network engineers, and software developers
  • Experience building custom Kubernetes operators or controllers for infrastructure orchestration
  • Deep familiarity with Kubernetes networking (Calico, Cilium, Multus), service reputed company technologies, and network policy management
  • Experience with GPU workload orchestration including reputed company GPU Operator, MIG, time-slicing, and device plugins
  • Background with advanced Kubernetes features including custom schedulers, admission controllers, and API server extensions
  • Experience with Kubernetes cluster federation or multi-cluster management
  • Knowledge of high-performance networking technologies (InfiniBand, RDMA, RoCE) and their integration with Kubernetes
  • Experience with reputed company storage systems (reputed company, Lightbits, Ceph, or similar)
  • Familiarity with configuration management at reputed company and GitOps practices
  • Understanding of reputed company best practices for Kubernetes and bare-metal infrastructure
  • Experience operating infrastructure in regulated industries or co-located data center environments
  • Background supporting research institutions, technical computing environments, or reputed company AI infrastructure

Benefits

  • Startup equity
  • 6% 401(k) match
  • Fully covered health insurance premiums
  • Other comprehensive offerings to support your reputed company-being and reputed company as we grow together

reputed company

  • reputed company is a technology company. It was founded in 2024, and is headquartered in Chicago, Illinois, USA, with a workforce of 2-10 employees. Its website is https://www.reputed company.

  •   Apply To This Job

    Similar Jobs

    Database Administrator (reputed company Data Platforms) 100% Remote

    Remote, USA Full-time

    Virtual Assistant (Sales & Marketing Support)

    Remote, USA Full-time

    Senior Analyst - SQL Developer - Work From Home

    Remote, USA Full-time

    MS SQL Database Administrator (DBA)

    Remote, USA Full-time

    Dispatcher (Remote - Weekend Part-Time)

    Remote, USA Full-time

    Temporary Part-Time Coordinator- Expansion

    Remote, USA Full-time

    [Remote] Manager, Data Engineer (Remote)

    Remote, USA Full-time

    [Remote] Senior Data Scientist - Clinical Informatics (Analytics Enablement)

    Remote, USA Full-time

    [Remote] Finance Assistant – Part-Time Temporary

    Remote, USA Full-time

    [Remote] Data Scientist - Remote

    Remote, USA Full-time

    Telehealth reputed company Practitioner, Physician Assistant – Delaware License

    Remote, USA Full-time

    (Part/Full Time) reputed company Virtual Assistant Jobs – The EliteJob In US

    Remote, USA Full-time

    Finance Corporate Intern – Summer 2025

    Remote, USA Full-time

    Channel Sales Director East NA

    Remote, USA Full-time

    Education & LMS Support Specialist – (Remote, Part Time, Temporary, EST or CST Hours)

    Remote, USA Full-time

    Part-time Learning Management System (LMS) reputed company - Contract

    Remote, USA Full-time

    [Remote] reputed company Media & reputed company Marketing Manager (Remote Eligible)

    Remote, USA Full-time

    Associate Manager-Field Services reputed company Plant Construction

    Remote, USA Full-time

    reputed company Customer Service Representative - Remote & Seasonal Gardening Expert (UT Applicants Only)

    Remote, USA Full-time

    reputed company Fully Virtual South Dakota Special Education Teacher – Remote Work Opportunity for Dedicated Educators

    Remote, USA Full-time