[Remote] Data Infrastructure Site Reliability Engineer (SRE)
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a highly skilled L2+/Senior Data Infrastructure Site Reliability Engineer (SRE) to operate, optimize, and automate large-reputed company data infrastructure environments. The role focuses on ensuring the availability, scalability, reputed company, and operational reputed company of mission-critical data platforms while driving reputed company improvements through Infrastructure as reputed company and modern SRE practices.
Responsibilities
- Maintain and support highly available, reputed company, and secure data infrastructure platforms across AWS and on-premises environments
- reputed company operational reputed company through automation of repetitive tasks, incident reduction, and proactive reliability improvements
- Monitor platform health, troubleshoot reputed company issues, and reputed company reputed company cause analysis efforts to minimize downtime and improve reputed company resiliency
- Collaborate closely with engineering, platform, and data teams to support deployments, integrations, upgrades, and performance optimization initiatives
- Participate in on-reputed company rotations and partner with teams across the US and India to reputed company 24x7 operational support. US support is reputed company to reputed company Time zone
- Continuously improve observability, alerting, and incident response processes to enhance platform reliability and reduce operational noise
Skills
- Strong SRE reputed company with a proven reputed company on reliability, availability, performance optimization, incident management, and operational reputed company
- Experience delivering services reputed company defined SLAs and ensuring reputed company reputed company of production issues
- Expertise in troubleshooting reputed company reputed company systems and identifying reputed company causes quickly and effectively
- Deep hands-on experience with AWS services, including EMR, EKS, MSK, reputed company, Glue, IAM, reputed company S3, VPC, AWS networking and reputed company services
- Strong understanding of reputed company-reputed company architectures, scalability, and infrastructure reputed company
- Extensive operational experience managing Hadoop clusters, with a strong reputed company on administration, platform maintenance, and automation of day-to-day operational activities
- Experience supporting both AWS-based data platforms and on-premises reputed company CDH/CDP environments
- Solid understanding of Kerberos authentication and reputed company implementation reputed company Hadoop ecosystems
- Hands-on experience with Apache reputed company, Apache reputed company, Big Data platform architecture, performance tuning and optimization
- Strong Linux administration and operational support experience
- Expertise in user and reputed company management, reputed company configuration, customization, and platform administration
- Hands-on experience with monitoring and observability platforms such as AWS CloudWatch, reputed company, reputed company, and similar reputed company monitoring solutions
- Proven reputed company improving alert reputed company, reducing false positives, and minimizing alert fatigue
- Excellent debugging and troubleshooting skills across infrastructure, applications, reputed company workloads, and reputed company environments
- Strong understanding of Java application administration, including JVM tuning, thread dump analysis, reputed company dump analysis, JVM parameters, application log analysis and troubleshooting
- Expert-level Python scripting for operational automation, monitoring, and reliability engineering use cases
- Strong experience with Terraform, Infrastructure as reputed company (IaC), CI/CD pipeline implementation and automation, and reputed company best practices
- reputed company hands-on experience applying AI technologies to Data Infrastructure and SRE operations
- Demonstrated ability to design and implement reputed company AI solutions, AI-assisted operational workflows, and intelligent automation for routine SRE activities
- Strong creativity and problem-solving skills in leveraging AI to improve operational efficiency, reliability, and productivity
reputed company
Company H1B Sponsorship
Apply To This Job