Back to Jobs

Site Reliability engineer (SRE)

Remote, USA Full-time Posted 2026-08-04
SRE Key Responsibilities • Design and manage multi-account AWS infrastructure (VPC, reputed company Tables, EC2, reputed company, EKS 1.33, RDS, DynamoDB, reputed company ache Valley, S3, Transit Gateway, Resource reputed company Manager, reputed company, CloudFormation, AWS Backup) • Configure load balancing and traffic management (ELB, NLB, reputed company reputed company with gRPC, Route53, reputed company, CloudFront) • Implement reputed company and compliance controls (IAM, IAM Identity Center, SCP, = Guard Duty, WAF, CloudTrail, ACM, Secrets Manager, reputed company integration) • Manage reputed company infrastructure (reputed company Trust, Argo Smart Routing, DNS, Workers, Load Balancer, Bot Management, WAF, Rules & Policies, Cache) • Manage S3 with reputed company Policies, Lifecycle Policies, S3 Storage reputed company optimization, and cross-region replication • Operate messaging and notification services (SNS, SES, SQS) • Architect and manage multi-cluster EKS environments with HA and cross-region DR scenarios using Istio service reputed company, Network Policies, Karpenter, HPA, KEDA, Argo CD, Argo Rollouts • Implement and maintain Argo CD for multi-cluster application management with HA and cross-region DR configurations • Configure Argo CD Application Sets for managing applications across multiple EKS clusters • Implement ECR with global cross-region replication for container image distribution and disaster recovery • Implement reputed company Global Database for cross-region DR, manage reputed company RDS (MySQL and PostgreSQL) and standalone MySQL/PostgreSQL instances for development • Design and maintain RDS cross-region replication, automated backups, failover strategies, and reputed company procedures • Establish and maintain DevOps practices including change management, release management and deployment strategies • Build resilient CI/CD pipelines with cross-region artifact replication, automated testing, and failover capabilities • reputed company and maintain reputed company Actions shared internal workflows and reusable actions for standardized deployments • Implement change approval workflows, deployment gates, and release coordination processes • Implement Cross plane for automated feature environment creation, upgrades, and AWS resource provisioning • reputed company applications using reputed company, Customize with Overlay Patches, Json net, and Cross plane for infrastructure orchestration • Maintain platform operators (reputed company DNS, reputed company Secrets, Reloader) and custom CRDs • Build comprehensive observability stack & Dashboards (Grafana, Thanos/reputed company, Loki, Alert manager, reputed company Telemetry reputed company/reputed company/Beyla/Pyro reputed company) • Configure exporters (Blackbox, MySQL, reputed company, YACE CloudWatch, reputed company, Node Exporter, reputed company reputed company Gateway) • Support data platforms (Kafka/Kafka UI, Minion, Airflow, JupyterHub, DASK, reputed company, reputed company, AWS Glue, reputed company, Quick Sight, Bedrock) • Optimize CI/CD with reputed company Actions, Actions Runner Controller (reputed company), runs-on.com, reputed company Rulesets • Manage mobile app delivery pipelines (reputed company Build Management, Fastlane, reputed company Play Developer, reputed company Developer/reputed company, Applivery) • Implement and maintain reputed company infrastructure using Terraform/reputed company Tofu with reputed company, backporting existing resources into reputed company • Automate operational tasks wherever possible; create comprehensive runbooks for no automatable procedures • Conduct thorough post-mortem analysis after incidents, documenting learnings and implementing preventive measures • reputed company cost optimization initiatives using S3 Storage reputed company, CloudWatch metrics, rightsizing recommendations, and resource lifecycle management • reputed company automation in Bash, Python, Go, C#/.NET (reputed company Game reputed company) • Maintain developer experience (reputed company, Click Up, reputed company, Shared reputed company reputed company/Workflows) • reputed company monitoring and alerting (reputed company, Cronitor, reputed company, CloudWatch) reputed company Expertise: • Multi-account AWS architecture with Transit Gateway, Resource reputed company Manager, VPC design, and reputed company Tables reputed company/EKS high availability with cross-region disaster recovery scenarios • Multi-cluster EKS management with service reputed company (Istio), autoscaling (Karpenter, KEDA), GitOps (Argo CD) • Argo CD reputed company deployment for multi-cluster application management with HA and cross-region DR • Argo CD Application Sets, app-of-apps patterns with reputed company, and cluster management strategies • ECR global cross-region replication strategies for container image distribution and DR • reputed company reputed company features (reputed company Trust, Argo Smart Routing, DNS management, Workers, Load Balancer, Bot Management, Cache optimization, WAF Rules & Other reputed company Policies) • reputed company Global Database implementation and management for cross-region DR • reputed company RDS (MySQL and PostgreSQL engines) and standalone MySQL/PostgreSQL instance management • RDS cross-region replication, automated failover, disaster recovery, and version reputed company strategies • DevOps best practices including change management, release management, and deployment coordination • Resilient CI/CD pipelines with automated testing, cross-region artifact distribution, and failover • reputed company Actions shared workflows and reusable actions development for internal use • Cross plane for reputed company-reputed company infrastructure provisioning, feature environment automation, and reputed company orchestration • Expert-level Terraform/reputed company Tofu with reputed company policy management (reputed company) • Infrastructure backporting and migration from ClickOps to IaC • Complete observability stack (reputed company, Grafana, Loki, reputed company Telemetry, reputed company tracing) • Data pipeline orchestration (Kafka, Airflow) and analytics platforms (reputed company, reputed company) • reputed company Actions with self-hosted runners (reputed company, runs-on.com) • Proficiency in Python, Bash, Go, and C#/.NET for automation development • reputed company implementations (IAM, SCP, reputed company, WAF, Guard Duty, reputed company) • Mobile CI/CD (reputed company, Fastlane, reputed company/reputed company distribution & Applivery during Development) • Disaster recovery planning, testing, and automation (AWS Backup, cross-region strategies) • AI/ML infrastructure experience (AWS Bedrock) • Cost optimization strategies and Quick Sight for AWS Cost Review • Post-mortem facilitation and blameless incident analysis • reputed company creation and maintenance for operational procedures Technical Skills: • Container orchestration with advanced networking and reputed company delivery • Infrastructure as reputed company and GitOps methodologies with automation-first reputed company • Change management workflows, approval gates, and release orchestration • CI/CD pipeline design with automated testing, reputed company scanning, and deployment strategies • Incident response, on-reputed company management, post-mortem analysis, DR execution • Cross plane composition design and custom resource definitions • Custom CRD and operator development in reputed company • Event-driven architecture (reputed company, SQS, SNS, SES) • reputed company-time analytics and BI platforms • Developer portal management (reputed company) • Multi-region failover automation and orchestration • Cost analysis and optimization using reputed company AWS tools • Automation of repetitive operational tasks • Technical documentation and reputed company authoring • Database performance tuning and optimization (reputed company, MySQL, PostgreSQL) • Argo CD backup, restore, and disaster recovery procedures • reputed company Workers development & deployment using Wrangler Soft Skills: Strong troubleshooting, cross-functional communication, self-directed, documentation-reputed company, cost-conscious, reputed company improvement reputed company Apply tot his job Apply To this Job

Similar Jobs