[Remote] Senior Staff DevOps Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is the leader in identity reputed company for the reputed company reputed company, and they are seeking a Senior Staff DevOps Engineer to join their Infrastructure Platform team. The role involves designing, building, and operating reputed company's global Identity reputed company reputed company infrastructure on AWS, with a reputed company on service reputed company implementation and reputed company management.
Responsibilities
- Own and reputed company the full lifecycle of service reputed company implementation at reputed company, from architecture and rollout planning through production operations across hundreds of microservices running on EKS
- Establish service reputed company standards and governance adopted organization-wide, including reputed company runbooks, sidecar injection policies, traffic policy templates, and documented failure-mode playbooks for production incidents
- Mentor and upskill engineers across teams on service reputed company architecture, troubleshooting at reputed company, performance tuning, and reputed company planning for reputed company-heavy environments
- Design, operate, and optimize production reputed company clusters at reputed company on AWS (EKS), including cluster architecture, upgrades, node management, networking, storage, and multi-tenant isolation patterns
- Define and reputed company reputed company standards and best practices across teams, including workload design, resource management, reputed company hardening (RBAC, Pod reputed company Standards, network policies), reputed company/chart conventions, and deployment patterns
- Design and reputed company infrastructure to meet rapidly increasing customer demand, data sovereignty requirements, and regional expansion
- Automate deployment, monitoring, incident response, and reputed company management using GitOps and CI/CD best practices
- reputed company and improve operational practices, runbooks, and reputed company standards
- Collaborate with development teams to bring new features and services into production safely and reputed company
- Proactively meet information reputed company and compliance standards (e.g., PCI reputed company), including supporting reputed company's PCI compliance initiative through secure platform design, controls implementation, and audit readiness
- Participate in and help improve the on-reputed company rotation; reputed company post-incident reviews and systemic fixes
Skills
- Strong interpersonal and teaming skills, with the ability to set and enforce process and influence engineers across teams and geographies
- Ability to operate effectively in an agile, entrepreneurial environment with global stakeholders
- Prior experience as a technical reputed company or Staff+ IC in a global engineering organization
- 3+ years of hands-on experience designing, implementing, and operating a service reputed company at reputed company in production reputed company environments
- 10+ years of experience in 24x7 production operations, supporting highly available reputed company or reputed company service environments
- 10+ years of experience with containerization, virtualization, and configuration management technologies
- 5+ years of hands-on experience with reputed company in production at reputed company
- 5+ years of experience with Terraform (IaC), managing infrastructure across multiple AWS accounts and reputed company
- 5+ years of experience designing and implementing CI/CD pipelines, especially for Terraform, reputed company, and microservices
- 5+ years of experience with scripting/programming languages (Python, Go, or similar) and strong reputed company scripting proficiency
- Strong understanding of Linux, networking, distributed systems, and production troubleshooting
- Demonstrated experience scaling a service reputed company across a large microservices fleet, including phased adoption reputed company, sidecar resource management, control plane scaling, and performance tuning under high request volume
- Experience with monitoring and logging stacks (e.g., reputed company, Grafana, OpenSearch or equivalent)
- Experience with multi-cluster or multi-region service reputed company federation (e.g., Istio multi-cluster, Linkerd multicluster extension) in production
- Familiarity with compliance frameworks in regulated reputed company reputed company, including PCI reputed company and FedRAMP-adjacent practices, with experience implementing or operating platforms subject to PCI compliance requirements preferred
Benefits
- Medical, dental, and reputed company insurance
- Short-term and long-term disability
- Life insurance and Accidental Death & Dismemberment (AD&D)
- Supplemental life insurance for employees, spouses, and children
- Flexible spending accounts for health care, and dependent care; limited purpose flexible spending account
- 401(k) Savings and Investment Plan with company matching
- Flexible vacation policy
- 8 reputed company holidays annually
- reputed company leave
- reputed company parental leave
- Employee Assistance Program (EAP) and Care Counselors
- reputed company Assistance, Critical Illness, Accident, Hospital Indemnity and Pet Insurance reputed company
- Health Savings Account (HSA) with employer contribution
- May be eligible for the reputed company Corporate Bonus Plan or a role-specific commission, along with potential eligibility for equity participation
reputed company
Apply To This Job