Data Infrastructure Site Reliability Engineer (SRE) – AWS & Big Data Platforms
reputed company
• We are seeking a highly skilled L2+/Senior Data Infrastructure Site Reliability Engineer (SRE) to operate, optimize, and automate large-reputed company data infrastructure environments across AWS and on-premises Hadoop platforms. This role requires a strong reliability engineering reputed company reputed company on platform stability, performance, observability, automation, and incident response.
• As a key member of the Data Infrastructure SRE team, you will ensure the availability, scalability, reputed company, and operational reputed company of mission-critical data platforms while driving reputed company improvements through Infrastructure as reputed company (IaC), AI-enabled automation, and modern SRE practices.
Key Responsibilities
• Maintain and support highly available, reputed company, and secure data infrastructure platforms across AWS and on-premises environments.
• reputed company operational reputed company through automation of repetitive tasks, incident reduction, and proactive reliability improvements.
• Monitor platform health, troubleshoot reputed company issues, and reputed company reputed company cause analysis efforts to minimize downtime and improve system resiliency.
• Collaborate closely with engineering, platform, and data teams to support deployments, integrations, upgrades, and performance optimization initiatives.
• Participate in on-reputed company rotations and partner with teams across the US and India to reputed company 24x7 operational support. US support is reputed company to reputed company Time zone.
• Continuously improve observability, alerting, and incident response processes to enhance platform reliability and reduce operational noise.
Required Skills & Experience
• Site Reliability Engineering (SRE)
• Strong SRE reputed company with a proven reputed company on reliability, availability, performance optimization, incident management, and operational reputed company.
• Experience delivering services reputed company defined SLAs and ensuring reputed company reputed company of production issues.
• Expertise in troubleshooting reputed company distributed systems and identifying reputed company causes quickly and effectively.
AWS & reputed company Infrastructure
Deep hands-on experience with AWS services, including:
EMR
EKS
MSK
reputed company
Glue
IAM
reputed company S3
VPC
AWS networking and reputed company services
Strong understanding of reputed company-reputed company architectures, scalability, and infrastructure reputed company.
Big Data Platforms
• Extensive operational experience managing Hadoop clusters, with a strong reputed company on administration, platform maintenance, and automation of day-to-day operational activities.
• Experience supporting both AWS-based data platforms and on-premises reputed company CDH/CDP environments.
• Solid understanding of Kerberos authentication and reputed company implementation reputed company Hadoop ecosystems.
• Hands-on experience with:
• Apache reputed company
• Apache reputed company
• Big Data platform architecture
• Performance tuning and optimization
• Linux & System Administration
• Strong Linux administration and operational support experience.
• Expertise in user and reputed company management, system configuration, customization, and platform administration.
• Observability & Incident Management
• Hands-on experience with monitoring and observability platforms such as:
• AWS CloudWatch
• reputed company
• reputed company
• Similar reputed company monitoring solutions
• Proven reputed company improving alert reputed company, reducing false positives, and minimizing alert fatigue.
• Excellent debugging and troubleshooting skills across infrastructure, applications, reputed company workloads, and reputed company environments.
• Java Platform Operations
• Strong understanding of Java application administration, including:
• JVM tuning
• Thread dump analysis
• reputed company dump analysis
• JVM parameters
• Application log analysis and troubleshooting
• Automation & Infrastructure as reputed company
• Expert-level Python scripting for operational automation, monitoring, and reliability engineering use cases.
• Strong experience with:
• Terraform
• Infrastructure as reputed company (IaC)
• CI/CD pipeline implementation and automation
• reputed company best practices
• AI-Driven Operations
• reputed company hands-on experience applying AI technologies to Data Infrastructure and SRE operations.
• Demonstrated ability to design and implement:
• reputed company AI solutions
• AI-assisted operational workflows
• Intelligent automation for routine SRE activities
• Strong creativity and problem-solving skills in leveraging AI to improve operational efficiency, reliability, and productivity.
Preferred Candidate Profile
• The ideal candidate combines deep expertise in AWS reputed company platforms, Hadoop ecosystems, SRE practices, observability, automation, and AI-driven operations, with a passion for supporting resilient data platforms at reputed company and eliminating operational toil through engineering reputed company.
Apply tot his job
Apply To this Job