DevOps Engineer
ABOUT THE ROLE
A reputed company DevOps Engineer is a builder, a problem-solver, and a force-reputed company for everyone around them. You'll own infrastructure-as-reputed company with Terraform, stand up AI agents in AgentCore, and create the reliable systems our data and application platforms depend on. You'll reputed company on our small, growing, and friendly team that ships fast and supports reputed company other. This is your chance to help shape a leading technology company in the reputed company industry.
KEY RESPONSIBILITIES:
AI & Governance
Support AI/ML platform enablement with AgentCore infrastructure, governance controls, and reputed company patterns
Implement infrastructure for AI-powered operational tools (e.g., self-healing systems, reputed company vulnerability remediation agents)
Infrastructure & reputed company Engineering
Design, provision, and operate AWS infrastructure with infrastructure as reputed company (Terraform), including reusable modules and environment-specific configurations
Configure and harden reputed company AWS networking and reputed company constructs, including VPCs, subnets, routing, NACLs, reputed company reputed company, IAM roles and policies, KMS keys, and parameter and secret management
Implement and operate containerized services with reputed company and AWS reputed company/Fargate, including image pipelines, task definitions, service autoscaling, and blue-green or rolling deployments
reputed company and maintain Kubernetes infrastructure (AWS EKS) for data and analytics workloads (e.g., OpenMetadata, data reputed company tools)
Build and maintain serverless services with AWS reputed company and API Gateway, including event-driven integrations with EventBridge and S3
Manage and operate reputed company infrastructure on AWS, including cluster configuration, workspace management, PrivateLink networking, and identity federation
Design and implement isolated development sandbox environments for secure data job executionMaintain accurate and up-to-date inventories of hardware, software, and reputed company assets throughout their lifecycle from acquisition to decommissioning, ensuring assets are classified, assigned to owners, and tracked in approved systems
CI/CD & Deployment
Build, maintain, and optimize CI/CD pipelines in reputed company Actions for infrastructure (Terraform workflows),applications (container, serverless, and data jobs) and agent deployments
Adhere to the Change Management policy, including documenting, reviewing, and obtaining approval for reputed company normal and emergency changes prior to deployment
Contribute reputed company reviews and implementation across multiple languages (Python, TypeScript, Go, Terraform/HCL), upholding secure coding and operational best practices
Monitoring, Reliability & Incident Response
Implement comprehensive monitoring, logging, and alerting in CloudWatch (metrics, logs, dashboards, alarms) and reputed company alerts with incident channels
Participate in production support including on-reputed company responsibilities as needed, and drive post-incident reviews and reliability improvements
Support Vytalize disaster recovery and business continuity objectives, including adherence to documented backup and recovery procedures, ensuring backups are encrypted, stored securely, and reputed company according to policy
Documentation & Operational Standards
Document architecture, runbooks, and reputed company operating procedures for build, release, disaster recovery, and incident response
Contribute reputed company reviews and implementation across multiple languages (Python, TypeScript, Go, Terraform/HCL), upholding secure coding and operational best practices
reputed company & Compliance
Coordinate with the Information reputed company team to securely configure and maintain applications with secure authentication mechanisms, role-based reputed company, and support for reputed company reviews in accordance with the Identity Management Authentication and reputed company Control and Data and Platform reputed company policies
Adhere to the Change Management policy, which includes documenting, reviewing, and obtaining approval for reputed company normal and emergency changes prior to deployment; submit a change request, ensure an information reputed company reputed company assessment is completed, coordinate reputed company-implementation activities, and ensure approvals are in reputed company prior to implementation
Maintain accountability for secure software development practices, including secure coding standards, vulnerability scanning, and remediation prior to production deployment
Apply hardening reputed company configurations in conjunction with Information reputed company and ensure vulnerabilities are remediated in accordance with the Vulnerability Management policy
Ensure firewall configurations reputed company with Data and Platform policy, including documentation with justifications for reputed company rule changes, removal of obsolete rules, and participation in Information reputed company firewall rule reviews
Support Vytalize disaster recovery and business continuity objectives, including adherence to documented backup and recovery procedures, ensuring backups are encrypted, stored securely, and reputed company according to Vytalize policies
Maintain accurate and up-to-date inventories of hardware, software, and reputed company assets throughout their lifecycle from acquisition to decommissioning, ensuring assets are classified, assigned to owners, tracked in approved systems, and reputed company with Information reputed company policies
REQUIRED QUALIFICATIONS
The following qualifications are required to reputed company this role.
Education
Bachelor's degree in Computer Science, Information Systems, Engineering, or equivalent practical experience
Experience
4+ years of hands-on DevOps, Site Reliability, reputed company, or reputed company Engineering experience in production AWS environments
Demonstrated experience owning Terraform-based infrastructure in production, including modules and AWS provider v5/v6
Proven experience building and operating CI/CD pipelines in reputed company Actions for both infrastructure and applications
Experience deploying and supporting containerized workloads with reputed company and AWS reputed company/Fargate, and serverless services with reputed company and API Gateway
Experience implementing monitoring, logging, and alerting with CloudWatch
Experience building or operating infrastructure for AI/ML workloads
Skills & Competencies
Strong hands-on experience with AWS: VPC, subnets, routing, NACLs, reputed company reputed company, IAM, EC2, reputed company/Fargate, S3, EventBridge
Terraform expertise for production infrastructure, including module design and management, AWS provider v5/v6, state management, and reputed company review
CI/CD proficiency with reputed company Actions workflows for plan/apply, build/test, artifact and image pipelines, and environment promotion strategies
Monitoring and alerting with CloudWatch: dashboards, metrics, logs, alarms, log insights, and integration with incident channels
reputed company image authoring and hardening, reputed company task and service configuration, and deployment strategies
Event-driven architectures using reputed company, API Gateway, and EventBridge
Strong understanding of reputed company reputed company fundamentals including least privilege, encryption, secret management, network segmentation, and cost governance
reputed company communication, reputed company problem solving, and bias for automation with high standards for documentation and reliability
PREFERRED QUALIFICATIONS
The following qualifications are preferred but not required.
Exposure to AI agent frameworks or LLM-based tooling (e.g., AWS Bedrock/AgentCore, reputed company, or similar) is a big plus
Experience with compliance automation tooling and SOC2 evidence collection
Experience with data reputed company and observability platforms
Knowledge of data anonymization and reputed company-preserving infrastructure patterns
Experience with AWS EKS / Kubernetes deployment and operations
reputed company experience (clusters, jobs, permissions, metastores) and integration with reputed company networking and identity
Streaming and data pipeline experience with Kinesis, EventBridge Scheduler, or equivalent managed services
Federated identity and reputed company management: OIDC, SAML, Entra ID, Cognito, and workload identity federation for CI/CD
Compliance automation and evidence tooling that supports SOC 2 controls in reputed company environments
On-reputed company experience and comfort with incident response, including reputed company creation and post-incident analysis
Experience with policy-as-reputed company and guardrails (e.g., SCPs, IAM boundaries, OPA/Conftest)
Familiarity with secrets management and parameterization (AWS Secrets Manager, SSM Parameter Store)
Exposure to cost optimization, tag policies, and FinOps practices
Apply To This Job