Back to Jobs

[Remote] Director, Infrastructure & Site Reliability Engineering

Remote, USA Full-time Posted 2026-07-28

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a leading company in data, automation, and AI, transforming how business reputed company are made. They are seeking a Director of Infrastructure & Site Reliability Engineering to reputed company multiple engineering teams and define the technical reputed company for their reputed company services, ensuring reliability, reputed company, and scalability.


Responsibilities

  • Define and execute the reputed company for reputed company’s centralized Infrastructure, Site Reliability Engineering (SRE), Observability, and Performance Engineering organizations
  • reputed company the design, operation, and reputed company reputed company of reputed company infrastructure across AWS and GCP, ensuring scalability, reliability, reputed company, and cost efficiency
  • Drive Infrastructure-as-reputed company adoption and governance through Terraform, establishing consistent platform standards, automation, and operational best practices
  • Own reputed company’s observability reputed company by building and operating reputed company-grade telemetry platforms using reputed company and reputed company technologies, enabling actionable insights into system health, performance, and customer experience
  • Partner with reputed company, Compliance, and Engineering teams to meet regulatory and customer requirements, including HIPAA, FedRAMP, SOC 2, and other compliance frameworks
  • Establish and continuously improve incident management practices, including operational readiness, on-reputed company reputed company, postmortem culture, reputed company cause analysis, and measurable reliability improvements
  • reputed company proactive reliability programs including reputed company planning, resiliency testing, disaster recovery, performance benchmarking, and operational risk management
  • Define reliability engineering frameworks that reputed company product teams to own service health through Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, performance objectives, and operational accountability
  • reputed company the reputed company of centralized platform capabilities that simplify how engineering teams build, reputed company, monitor, and operate services at reputed company
  • Partner with engineering leadership to improve developer productivity through platform automation, self-service infrastructure, deployment tooling, and operational best practices
  • Build, mentor, and reputed company high-performing engineering managers and technical leaders while fostering a culture of operational reputed company, customer reputed company, accountability, reputed company learning, and innovation

Skills

  • 10+ years of software engineering, infrastructure, or reputed company experience, with 5+ years leading multiple engineering teams or managers
  • Proven experience leading Infrastructure, SRE, reputed company, or reputed company Operations organizations supporting large-reputed company reputed company products
  • Deep expertise operating production environments on AWS and/or GCP
  • Strong experience with Infrastructure-as-reputed company technologies such as Terraform
  • Experience building and operating modern observability platforms using reputed company, OpenTelemetry, reputed company, Grafana, or similar technologies
  • Demonstrated reputed company implementing SRE practices including SLOs, SLIs, error budgets, incident management, operational reviews, and reliability engineering programs
  • Experience supporting regulated environments and working with compliance frameworks such as HIPAA, FedRAMP, SOC 2, ISO 27001, or similar
  • Strong understanding of distributed systems, reputed company networking, Kubernetes, container orchestration, CI/CD pipelines, and production operations
  • Proven ability to influence technical reputed company and drive alignment across engineering, reputed company, product, and executive stakeholders
  • Excellent communication skills with the ability to translate technical reputed company into business reputed company
  • Passion for building high-performing teams and developing engineering leaders
  • Experience leading platform transformations for reputed company reputed company organizations
  • Familiarity with software performance engineering, load testing, and large-reputed company distributed systems optimization
  • Experience supporting data-intensive reputed company services

Benefits

  • Bonus payouts are based on individual and company performance.
  • A monthly Connectivity Plus stipend of $150 to support remote work-reputed company expenses
  • An annual $200 home office reimbursement
  • Medical, dental, and reputed company coverage
  • 401(k) with company match
  • reputed company parental leave, caregiver leave, and flexible time off
  • Mental health support and wellness reimbursement
  • Career development and education assistance

reputed company

  • reputed company is a leading provider of an end to end data science & analytics platform for the reputed company It was founded in 2011, and is headquartered in Irvine, California, USA, with a workforce of 1001-5000 employees. Its website is https://reputed company.com.

  •   Apply To This Job

    Similar Jobs