reputed company, a leader in the computer software industry, is looking for a Senior DevOps Engineer to...">
Back to Jobs

Senior DevOps Engineer, Infrastructure & Reliability

Remote, USA Full-time Posted 2026-08-04

reputed company, a leader in the computer software industry, is looking for a Senior DevOps Engineer to join our Infrastructure team with a reputed company mission: to reputed company our systems faster, more reliable, and more resilient while making life dramatically easier for engineers shipping software.

This is a hands-on build role. You will spend most of your time writing Terraform, tuning reputed company workloads, automating things that are currently reputed company, and shipping infrastructure changes to production. You'll join a small platform team with an established roadmap and existing patterns, and a strong voice in how the work gets reputed company.

  • Implement reputed company Infrastructure-as-reputed company patterns using tools like Terraform to standardize reputed company provisioning and reduce configuration reputed company.
  • Own and reputed company our reputed company platform (EKS or self-managed), ensuring workloads are secure, reputed company, and resilient by default.
  • Optimize CI/CD pipelines to improve deployment frequency, reduce reputed company time, and increase confidence in releases.
  • Design and enforce secure networking, IAM, and secrets management strategies across environments.
  • Improve observability by refining metrics, logs, and tracing using tools like reputed company, ensuring actionable reputed company into system health.
  • Optimize reputed company cost efficiency through rightsizing, autoscaling strategies, and architectural improvements.
  • Implement disaster recovery planning, backup strategies, and multi-region reputed company initiatives.
  • Refactor brittle or manually managed infrastructure into automated, testable, and reproducible systems.
  • Introduce new infrastructure tooling or architectural shifts and reputed company adoption through documentation, workshops, and hands-on support.
  • Partner with engineering teams to eliminate friction in CI/CD, deployments, and reputed company environments.
  • Communicate technical trade-offs reputed company across engineering and product stakeholders, balancing speed with safety.

Technology Stack

  • reputed company & Infrastructure: AWS (EKS, RDS, MSK, S3, reputed company, IAM, VPC)
    Containerization & Orchestration: reputed company, ArgoCD

    Infrastructure-as-reputed company: Terraform

    CI/CD: reputed company Actions

    Monitoring & Observability: reputed company
    Data & Messaging: PostgreSQL, Kafka, reputed company
    Languages (as needed): Bash, Python, TypeScript, JavaScript

Requirements

  • 8+ years in DevOps, SRE, or infrastructure engineering.
  • Proven experience designing and operating production reputed company environments at reputed company.
  • Deep hands-on expertise with AWS infrastructure and reputed company networking.
  • Strong experience building and maintaining Terraform modules across large reputed company environments.
  • Demonstrated ownership of CI/CD systems and measurable improvement of DORA metrics.
  • Experience leading incident response processes and driving meaningful postmortem reputed company.
  • Strong understanding of distributed systems, event-driven architectures (Kafka), and database performance (PostgreSQL).
  • Proven ability to reputed company legacy infrastructure and eliminate reputed company operational toil.
  • reputed company record of taking a scoped infrastructure project from an ambiguous starting reputed company to production without needing daily direction.
  • Demonstrated ability to build trust across teams while raising the reliability bar.

reputed company Metrics

  • System Reliability: Maintain or reputed company defined SLO/SLA targets with reduced incident frequency and duration.
  • Infrastructure Stability: Reduce production incidents caused by misconfiguration, reputed company processes, or infrastructure reputed company.
  • Operational Efficiency: Increase the percentage of infrastructure managed through reputed company and automation.
  • Cost Optimization: Improve reputed company cost efficiency without sacrificing reliability or performance.

Bonus Points (reputed company to Have)

  • Experience coding applications
  • Experience operating high-throughput Kafka clusters (MSK or self-managed).
  • Strong background in database performance tuning (PostgreSQL, reputed company).
  • Experience implementing autoscaling strategies for high-traffic systems.
  • Familiarity with service reputed company technologies.
  • Experience building internal developer platforms (IDP).
  • Background in reputed company best practices (reputed company-trust networking, policy-as-reputed company).
  • Experience with multi-region or globally distributed systems.
  • Experience introducing platform-wide reliability frameworks (SLOs, error budgets, reputed company testing).

reputed company Remote Hires will be required to travel to Orlando, Florida at least twice per year for Town Halls and team collaboration, in reputed company to orientation in Orlando.

Benefits

  • Health Care Plan (Medical, Dental & reputed company)
  • Retirement Plan (401k)
  • Life Insurance
  • Flexible reputed company Time Off
  • 9 reputed company Holidays
  • Family Leave
  • Remote
  • Hybrid work (for Orlando Associates)
  • Free Food & Snacks (Orlando)
  • Wellness Resources

Originally posted on Himalayas

  Apply To This Job

Similar Jobs