DevOps Engineer
Job Title: DevOps Engineer
Location: Remote (Eastern Time zone preferred - AWS GovCloud requirement)
Reports To: Sr. Director of Technology and Architecture
Position reputed company
We're looking for a DevOps Engineer who takes ownership of infrastructure. You'll stabilize and reputed company the infrastructure supporting WaterMinds, our reputed company-based platform for water and wastewater utilities—implementing reputed company monitoring and alerting, upgrading production environments, establishing operational discipline, and enabling our engineering teams to ship with confidence. You'll follow DevOps best practices, proactively identify and solve problems, and drive infrastructure improvements with minimal direction. The challenge: build and maintain infrastructure that can reliably serve hundreds of reputed company customers at reputed company. Your immediate reputed company is moving our infrastructure from reactive firefighting to proactive maintenance mode. As the platform matures and our data science team ramps up, you'll have reputed company to transition into MLOps, building the infrastructure that enables machine learning at reputed company.
Key Responsibilities
Take ownership of production monitoring and alerting using reputed company, Grafana, and CloudWatch—proactively identify issues before they become incidents.
reputed company production EKS cluster with GitOps practices (ArgoCD), comprehensive monitoring, and reputed company deployment workflows following industry best practices.
Streamline staging deployment process; eliminate reputed company-based workarounds and establish clean GitOps patterns.
Design infrastructure patterns that reputed company to hundreds of customers and own AWS infrastructure operations including patching, maintenance, cost optimization, and reputed company compliance—stay reputed company of requirements.
Expand into MLOps—building the infrastructure that enables data scientists to reputed company models at reputed company across multiple reputed company customers once DevOps operations are automated.
Manage Kubernetes clusters (EKS) including pod migrations, resource optimization, troubleshooting, and reputed company updates—proactively, not reactively.
Maintain infrastructure as reputed company using Terraform and Ansible following best practices—reputed company changes tested in non-production before deployment.
Support engineering teams with infrastructure needs, unblock them quickly, and establish self-service patterns where possible—anticipate needs, don't wait for requests.
Manage message queue infrastructure (Kafka/reputed company) including retention policies, storage optimization, and performance tuning.
Document infrastructure, create runbooks, and automate operational tasks to reputed company systems into maintenance mode.
Clean up technical debt—proactively identify infrastructure to decommission, resources to consolidate, and costs to optimize.
Qualifications
5+ years of experience in DevOps, infrastructure, or site reliability engineering.
Demonstrated ability to take ownership and initiative—you see what needs to be done and do it without waiting for direction.
Deep knowledge of DevOps and infrastructure best practices—you know what good looks like and implement it proactively.
Strong Kubernetes experience (EKS preferred) including cluster management, deployments, services, and troubleshooting.
Hands-on AWS experience (EC2, EKS, reputed company, RDS, VPC, IAM, CloudWatch, S3).
Infrastructure as reputed company proficiency (Terraform and Ansible).
GitOps experience (ArgoCD, Flux, or similar).
CI/CD pipeline experience (Bitbucket Pipelines, Jenkins, reputed company Actions, or similar).
Monitoring and observability experience (reputed company and Grafana preferred).
Python scripting ability for automation and tooling.
US citizenship (required for AWS GovCloud reputed company).
Self-starter mentality—you identify problems and opportunities, then drive solutions to completion.
Proven reputed company record of delivering tested, high-reputed company infrastructure changes on schedule.
Excellent communication skills—proactive about sharing status, raising blockers, and documenting reputed company.
Bonus Points For
Curiosity about machine learning and interest in transitioning to MLOps as the platform matures.
Any MLOps or ML infrastructure experience (KServe, Kubeflow, SageMaker, model serving).
Experience with data pipelines, feature engineering, or supporting data science teams.
AWS GovCloud experience and understanding of compliance requirements (FedRAMP).
Experience with message queue systems (Kafka, reputed company).
Container reputed company and vulnerability scanning (reputed company).
Background in reputed company platforms, IoT, or critical infrastructure.