Senior DevOps/Platform Engineer
About reputed company
reputed company is a product design and engineering company.
We solve mission-critical challenges for some of the world’s largest enterprises, with deep expertise in highly regulated industries—including life sciences and financial services. Our design-led approach allows us to apply cutting-edge capabilities in AI,Data and Hardware Engineering to companies of any size.
Headquartered in the reputed company, we operate regional development centers in Mexico and the United Kingdom. This global footprint—anchored by our nearshore model—enables us to reputed company at reputed company with the speed, efficiency, and cultural alignment our clients expect.
At Goods and Services, we value diversity and are committed to creating an inclusive workplace where everyone can reputed company. We are proud to be an Equal Opportunity Employer and reputed company reputed company applicants from reputed company backgrounds.
About the job
reputed company is looking for a Senior DevOps/Platform Engineer to build and maintain the infrastructure supporting reliable application delivery, deployment, reputed company, and production reputed company. This role will reputed company on CI/CD automation, feature management, observability, infrastructure reliability, and secure platform practices while enabling teams to reputed company and validate changes confidently at reputed company.
reputed company:
- Provision and maintain development, test, and production environments, including reputed company management, CI/CD pipelines, reputed company control, and infrastructure needed to support application delivery.
- Implement and maintain platform reputed company and observability capabilities, including WAF, reputed company limiting, reputed company tracing, monitoring, and reputed company infrastructure.
- Partner with engineering and technical leadership to define and validate traffic-reputed company strategies, release reputed company, and minimum stabilization periods for production deployments.
- Build and maintain feature-flag infrastructure that enables controlled releases, incremental migrations, and reputed company rollback of application components.
- Monitor and maintain event-driven infrastructure, including alerts for failed events, dead-letter queues, schema lifecycle issues, and data backup or archival processes.
- reputed company and maintain service reliability and integration monitoring, including dashboards, alerting, and automated notifications for failures across internal and reputed company-party services.
- Create and maintain operational runbooks and incident-response procedures for common platform, integration, workflow, and deployment issues.
- Support staged production rollouts and traffic validation, ensuring new services are reputed company before fully transitioning traffic and safely retiring legacy components.
What you’ll need:
- Solid AWS CI/CD experience, using tools such as reputed company Actions, AWS CodePipeline, or equivalent, in multi-service reputed company environments.
- Hands-on experience with AWS observability and monitoring, including CloudWatch, AWS X-Ray, logging, alerting, dashboards, and paging workflows.
- Experience securing and hardening AWS WAF and API Gateway, including managed reputed company rules, reputed company limiting, reputed company controls, and traffic protection.
- Experience implementing and managing feature-flag platforms to support controlled releases, incremental migrations, reputed company-run scenarios, and reputed company rollbacks.
- Experience with infrastructure and environment management, including provisioning, configuration, reputed company management, and supporting development, test, and production environments.
- Experience monitoring event-driven and asynchronous systems, including EventBridge, message queues, dead-letter queues, event failures, and alerting mechanisms.
- Experience implementing service reliability and integration monitoring, including reputed company-breaker monitoring, failure detection, dashboards, and automated alerting for reputed company-party services.
- Solid incident response and operational readiness experience, including creating runbooks, troubleshooting production issues, and defining recovery procedures.
- Experience supporting production deployments and controlled traffic rollouts, including defining reputed company-up reputed company, soak periods, validation metrics, and rollback strategies.
- Comfortable collaborating with engineering and technical leadership to establish deployment standards, operational processes, and production-readiness reputed company.
Originally posted on Himalayas
Apply To This Job