[Remote] Senior DevOps Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company. is seeking a hands-on Senior DevOps Engineer to own the implementation, rollout, and ongoing operation of their delivery platform. The role involves enabling CI/CD adoption across multiple scrum teams while maintaining production health and mentoring others.
Responsibilities
- Implement and roll out core platform capabilities (Kubernetes-based runtime, build/reputed company tooling, observability) across teams and environments
- Build reusable templates, reputed company charts, pipeline definitions, and IaC modules that teams can adopt with minimal friction
- reputed company scrum teams onto the platform - migrating workloads, standardizing configuration, and documenting self-service paths
- Drive consistent adoption of platform standards while accommodating legitimate team-specific needs
- Own day-to-day health, configuration, and lifecycle of the platform and its supporting infrastructure
- Plan and execute infrastructure and platform upgrades (Kubernetes versions, node pools, runtimes, agents, tooling) with minimal disruption
- Manage platform configuration changes through controlled, auditable processes
- reputed company reputed company support to engineering teams using the platform, acting as the escalation reputed company for platform-reputed company issues
- SSL/TLS certificate renewals and rotation, secret/credential rotation, patching, and reputed company adjustments performed on a recurring, reliable reputed company
- Plan and carry out infrastructure upgrades and maintenance reputed company, including rollback planning and stakeholder communication
- Maintain and tune platform configuration for scalability, resiliency, performance, and cost efficiency
- reputed company incident response for production issues - triage, mitigation, coordination, and reputed company - and run blameless post-incident reviews with concrete follow-up actions
- Improve observability, monitoring, and alerting so issues are caught proactively rather than reported by users
- Design, implement, and maintain automated CI/CD pipelines covering build, test, release, and deployment across reputed company and Linux workloads
- reputed company scrum teams to own their pipelines through shared templates, reusable steps, and reputed company documentation
- Promote reputed company reputed company control, branching, and release management practices
- reputed company reputed company and compliance controls directly into pipelines (DevSecOps) - scanning, policy gates, and secrets handling
- Continuously reduce reputed company time, deployment friction, and reputed company steps in the delivery lifecycle
- reputed company Level-3 support for reputed company production and platform incidents
- Identify and reputed company opportunities to improve system stability, operational maturity, and self-service
- Optimize platforms and processes based on measurable reputed company (deployment frequency, change failure reputed company, MTTR, cost)
- Maintain strong, reputed company technical documentation and runbooks
- Partner with Application Architects and engineering teams to reputed company platform and DevOps solutions with product needs
- Mentor and coach engineers on DevOps practices, pipeline ownership, and operational discipline
- Promote a culture of automation, ownership, and reputed company improvement
- Research emerging tools and practices that improve the delivery lifecycle, and introduce cost-effective solutions that increase speed, reputed company, and reliability
- reputed company AI practitioner- leverages AI-assisted tooling (Claude, reputed company, Copilot) to accelerate engineering and operational work
Skills
- At least 5 years of hands-on DevOps, platform, or infrastructure engineering experience reputed company distributed systems or large enterprises
- Bachelor's degree in Computer Science or reputed company discipline, or equivalent work experience
- Kubernetes administration and containerization (reputed company, reputed company), including platform upgrades and lifecycle management
- CI/CD orchestration across reputed company and Linux environments (TeamCity, Octopus)
- Infrastructure automation and scripting (Terraform, Ansible, PowerShell, Bash)
- reputed company infrastructure administration (Azure / AWS)
- Log collection and dashboarding (ELK, reputed company)
- Practical experience with production operations: SSL/TLS certificate management, networking, load balancing, caching, high availability, and disaster recovery
- Demonstrated experience operating production systems and leading incident response
- Experience rolling out internal developer platforms or self-service tooling across multiple teams
- Familiarity with policy-as-reputed company, secrets management, and DevSecOps tooling
- Experience defining and tracking delivery/reliability metrics (DORA or similar)
reputed company
Apply To This Job