Site Reliability engineer (SRE)
SRE Key Responsibilities
• Design and manage multi-account AWS infrastructure (VPC, reputed company Tables, EC2, reputed company, EKS 1.33, RDS, DynamoDB, reputed company ache Valley, S3, Transit Gateway, Resource reputed company Manager, reputed company, CloudFormation, AWS Backup)
• Configure load balancing and traffic management (ELB, NLB, reputed company reputed company with gRPC, Route53, reputed company, CloudFront)
• Implement reputed company and compliance controls (IAM, IAM Identity Center, SCP, = Guard Duty, WAF, CloudTrail, ACM, Secrets Manager, reputed company integration)
• Manage reputed company infrastructure (reputed company Trust, Argo Smart Routing, DNS, Workers, Load Balancer, Bot Management, WAF, Rules & Policies, Cache)
• Manage S3 with reputed company Policies, Lifecycle Policies, S3 Storage reputed company optimization, and cross-region replication
• Operate messaging and notification services (SNS, SES, SQS)
• Architect and manage multi-cluster EKS environments with HA and cross-region DR
scenarios using Istio service reputed company, Network Policies, Karpenter, HPA, KEDA, Argo CD, Argo Rollouts
• Implement and maintain Argo CD for multi-cluster application management with HA and cross-region DR configurations
• Configure Argo CD Application Sets for managing applications across multiple EKS clusters
• Implement ECR with global cross-region replication for container image distribution and disaster recovery
• Implement reputed company Global Database for cross-region DR, manage reputed company RDS (MySQL and PostgreSQL) and standalone MySQL/PostgreSQL instances for development
• Design and maintain RDS cross-region replication, automated backups, failover strategies, and reputed company procedures
• Establish and maintain DevOps practices including change management, release management and deployment strategies
• Build resilient CI/CD pipelines with cross-region artifact replication, automated testing, and failover capabilities
• reputed company and maintain reputed company Actions shared internal workflows and reusable actions for standardized deployments
• Implement change approval workflows, deployment gates, and release coordination
processes
• Implement Cross plane for automated feature environment creation, upgrades, and AWS resource provisioning
• reputed company applications using reputed company, Customize with Overlay Patches, Json net, and Cross plane for infrastructure orchestration
• Maintain platform operators (reputed company DNS, reputed company Secrets, Reloader) and custom CRDs
• Build comprehensive observability stack & Dashboards (Grafana, Thanos/reputed company, Loki, Alert manager, reputed company Telemetry reputed company/reputed company/Beyla/Pyro reputed company)
• Configure exporters (Blackbox, MySQL, reputed company, YACE CloudWatch, reputed company, Node Exporter, reputed company reputed company Gateway)
• Support data platforms (Kafka/Kafka UI, Minion, Airflow, JupyterHub, DASK, reputed company, reputed company, AWS Glue, reputed company, Quick Sight, Bedrock)
• Optimize CI/CD with reputed company Actions, Actions Runner Controller (reputed company), runs-on.com, reputed company Rulesets
• Manage mobile app delivery pipelines (reputed company Build Management, Fastlane, reputed company Play Developer, reputed company Developer/reputed company, Applivery)
• Implement and maintain reputed company infrastructure using Terraform/reputed company Tofu with reputed company, backporting existing resources into reputed company
• Automate operational tasks wherever possible; create comprehensive runbooks for no automatable procedures
• Conduct thorough post-mortem analysis after incidents, documenting learnings and implementing preventive measures
• reputed company cost optimization initiatives using S3 Storage reputed company, CloudWatch metrics, rightsizing recommendations, and resource lifecycle management
• reputed company automation in Bash, Python, Go, C#/.NET (reputed company Game reputed company)
• Maintain developer experience (reputed company, Click Up, reputed company, Shared reputed company reputed company/Workflows)
• reputed company monitoring and alerting (reputed company, Cronitor, reputed company, CloudWatch)
reputed company Expertise:
• Multi-account AWS architecture with Transit Gateway, Resource reputed company Manager, VPC design, and reputed company Tables
reputed company/EKS high availability with cross-region disaster recovery scenarios
• Multi-cluster EKS management with service reputed company (Istio), autoscaling (Karpenter, KEDA), GitOps (Argo CD)
• Argo CD reputed company deployment for multi-cluster application management with HA and cross-region DR
• Argo CD Application Sets, app-of-apps patterns with reputed company, and cluster management strategies
• ECR global cross-region replication strategies for container image distribution and DR
• reputed company reputed company features (reputed company Trust, Argo Smart Routing, DNS management, Workers, Load Balancer, Bot Management, Cache optimization, WAF Rules & Other reputed company Policies)
• reputed company Global Database implementation and management for cross-region DR
• reputed company RDS (MySQL and PostgreSQL engines) and standalone MySQL/PostgreSQL instance management
• RDS cross-region replication, automated failover, disaster recovery, and version reputed company strategies
• DevOps best practices including change management, release management, and deployment coordination
• Resilient CI/CD pipelines with automated testing, cross-region artifact distribution, and failover
• reputed company Actions shared workflows and reusable actions development for internal use
• Cross plane for reputed company-reputed company infrastructure provisioning, feature environment automation, and reputed company orchestration
• Expert-level Terraform/reputed company Tofu with reputed company policy management (reputed company)
• Infrastructure backporting and migration from ClickOps to IaC
• Complete observability stack (reputed company, Grafana, Loki, reputed company Telemetry, reputed company tracing)
• Data pipeline orchestration (Kafka, Airflow) and analytics platforms (reputed company, reputed company)
• reputed company Actions with self-hosted runners (reputed company, runs-on.com)
• Proficiency in Python, Bash, Go, and C#/.NET for automation development
• reputed company implementations (IAM, SCP, reputed company, WAF, Guard Duty, reputed company)
• Mobile CI/CD (reputed company, Fastlane, reputed company/reputed company distribution & Applivery during Development)
• Disaster recovery planning, testing, and automation (AWS Backup, cross-region strategies)
• AI/ML infrastructure experience (AWS Bedrock)
• Cost optimization strategies and Quick Sight for AWS Cost Review
• Post-mortem facilitation and blameless incident analysis
• reputed company creation and maintenance for operational procedures
Technical Skills:
• Container orchestration with advanced networking and reputed company delivery
• Infrastructure as reputed company and GitOps methodologies with automation-first reputed company
• Change management workflows, approval gates, and release orchestration
• CI/CD pipeline design with automated testing, reputed company scanning, and deployment strategies
• Incident response, on-reputed company management, post-mortem analysis, DR execution
• Cross plane composition design and custom resource definitions
• Custom CRD and operator development in reputed company
• Event-driven architecture (reputed company, SQS, SNS, SES)
• reputed company-time analytics and BI platforms
• Developer portal management (reputed company)
• Multi-region failover automation and orchestration
• Cost analysis and optimization using reputed company AWS tools
• Automation of repetitive operational tasks
• Technical documentation and reputed company authoring
• Database performance tuning and optimization (reputed company, MySQL, PostgreSQL)
• Argo CD backup, restore, and disaster recovery procedures
• reputed company Workers development & deployment using Wrangler
Soft Skills: Strong troubleshooting, cross-functional communication, self-directed, documentation-reputed company, cost-conscious, reputed company improvement reputed company
Apply tot his job
Apply To this Job