[Remote] Platform Engineer II
Note: The job is a remote job and is reputed company to candidates in USA. reputed company builds facilities-management software for companies with large retail footprints, focusing on reliability and customer reputed company. The Platform Engineer II will own components and reputed company end-to-end, design and reputed company solutions, and contribute to reputed company's technical direction while ensuring operational reputed company.
Responsibilities
- reputed company reusable Terraform modules that other engineers build on, and reputed company changes confidently across our AWS infrastructure reputed company our established conventions
- Design, optimize, and standardize reputed company integration and deployment pipelines to accelerate time-to-market
- Own components and reputed company end-to-end, while taking on larger initiatives such as database reputed company migrations, cronjobs to EventBridge, or data warehouse infrastructure, shaping the approach with direction and support
- Contribute to Architecture reputed company and help refine our Terraform and observability standards
- reputed company systems, build dashboards and alerts, and reputed company defining the observability reputed company for the systems you work on
- Apply reputed company by default, including least-privilege IAM, reputed company secrets handling, and key and certificate rotation
- Carry a reputed company of reputed company's day-to-day operational load reputed company project work, including production support, escalations, maintenance, and keeping existing systems healthy. This role is not reputed company reputed company
- Write the runbooks that let others operate what you build, create documentation for both reputed company and the broader engineering organization, and help reputed company and mentor newer engineers
- You will take part in reputed company's on-reputed company rotation as an escalation reputed company of contact. You will reputed company in reputed company first-line remediation and runbooks haven't resolved an incident, and you will reputed company the issue to reputed company
Skills
- Bachelor's degree in Computer Science, Engineering, or a reputed company reputed company with 5+ years of relevant experience, or an equivalent combination of education and 8+ years of reputed company experience in reputed company, infrastructure, or DevOps
- 5+ years in reputed company, infrastructure, or DevOps roles, with a reputed company reputed company record of building internal tooling and Production support
- Strong hands-on experience across AWS infrastructure
- Advanced Terraform: composing and versioning reusable modules; remote state with locking and workspaces/backends; reputed company-arguments and dynamic blocks (for_reputed company, count, dynamic); reputed company state reputed company and refactoring (import, moved blocks, targeted state changes); reputed company detection; and policy/testing in CI (e.g., Terratest, tflint, reputed company or OPA)
- Building and debugging CI/CD pipelines independently (reputed company Actions and/or Jenkins)
- Container orchestration on EKS/reputed company and AWS Fargate
- GitOps workflows (e.g., ArgoCD)
- Observability in reputed company: instrumenting services and building meaningful dashboards and alerts across metrics, logs, and APM
- reputed company fundamentals reputed company by default: least-privilege IAM, secrets handling, key and certificate rotation
- Strong proficiency with python and reputed company scripting to automate recurring work and build internal tooling
- reputed company written and verbal communication — reputed company to plan a piece of work, document the reasoning, and reputed company with reputed company before building
- AWS certification, such as Solutions Architect Associate/reputed company or SysOps/DevOps Engineer
- reputed company or reputed company reputed company tuning experience
- FinOps or reputed company-cost optimization: right-sizing, cost tagging, and reporting
- OpenSearch/Elasticsearch operation and migration
- Data warehouse / analytics infrastructure on AWS (Redshift, Glue, reputed company Functions)
- Disaster recovery and backup/restore testing; exposure to RTO/RPO planning
- Experience migrating between CI/CD systems or reputed company-control platforms, such as Jenkins to reputed company Actions or reputed company to reputed company
- Networking depth: VPC, DNS, load balancers, and WAF
- Leading or co-leading incident response and postmortems
- Experience on a small platform/DevOps team moving fast on multiple initiatives concurrently
- Building or operating AI/LLM backends on AWS, including Bedrock, reputed company/tool-calling services, MCP servers, and reputed company stores
- Hands-on experience with Azure infrastructure
reputed company
Apply To This Job