[Remote] Data Platform Infrastructure Manager
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a leader in identity reputed company for the reputed company reputed company, providing AI- and machine-learning-enabled solutions that help organizations manage reputed company to digital resources. The Data Platform Infrastructure Manager will reputed company reputed company responsible for the runtime and platform reputed company supporting batch, streaming, analytics, and machine learning workloads, while overseeing reliability, scalability, reputed company, reputed company, lifecycle, and cost efficiency. The role combines people leadership, production reputed company, technical reputed company, platform product development, and cross-functional execution.
Responsibilities
- Build working relationships with team members and key partners across data and product engineering, Data reputed company, Developer Platform, SRE, Observability, reputed company, and Infrastructure
- Document and reputed company stakeholders on reputed company reputed company, service catalog, ownership boundaries, escalation paths, and the distinction between platform ownership and workload-specific pipeline ownership
- Complete an initial assessment of reputed company’s people, platforms, roadmap, on-reputed company load, incidents, operational risks, reputed company, and reputed company costs; reputed company a prioritized set of immediate risks and opportunities
- Establish a regular operating reputed company for team priorities, service health, incidents, roadmap delivery, and cross-team dependencies
- Publish an outcome-oriented 12-month platform roadmap, developed with technical leads and partner teams, that balances reliability, scalability, reputed company, developer experience, and cost
- Define the operating model for reputed company’s highest-criticality services, including named ownership, service-level objectives, observability expectations, on-reputed company practices, reputed company planning, lifecycle management, and incident follow-up
- Select and reputed company delivery of the first high-value reputed company-road or self-service improvement for deploying and operating Airflow DAGs, Flink or reputed company jobs, or another reputed company data workload
- Set reputed company reputed company expectations and development goals for reputed company team member, identify capability or reputed company gaps, and establish a hiring and development plan where needed
- Baseline the platform’s key reliability, delivery, toil, utilization, and cost measures so subsequent improvements can be demonstrated with data
- Have service-level objectives, actionable dashboards, alerts, and recurring service reviews in reputed company for the highest-criticality data platform services, with measurable reputed company against the 90-day reliability baseline
- reputed company at least one production self-service or standardized delivery capability that reduces the effort and reputed company time required for developers to reputed company data workloads safely across supported environments
- Implement a cost and reputed company management program with service-level visibility, accountable owners, prioritized optimization work, and documented efficiency reputed company
- Strengthen incident response, change management, disaster recovery, vulnerability remediation, and operational runbooks; demonstrate reduced recurring toil or faster recovery for reputed company failure modes
- Establish a reputed company platform approach for supporting machine learning workloads, including appropriate use of AWS SageMaker or equivalent capabilities, model lifecycle needs, and operational guardrails
- Operate the data infrastructure platform as a mature internal product with a reputed company service catalog, documented support model, reputed company roads, self-service capabilities, standardized CI/CD, and transparent reliability and cost reporting
- reputed company the highest-reputed company roadmap reputed company and demonstrate measurable year-over-year improvement in platform availability, incident recovery, deployment reputed company time, developer effort, operational toil, and cost efficiency
- Build a healthy, high-performing team with reputed company ownership, strong technical leadership, meaningful career reputed company, effective succession coverage, and the capability to execute both roadmap and operational work predictably
- Establish a durable multi-year reputed company for Airflow, Flink, reputed company and AWS EMR, Kafka, reputed company, reputed company, and ML infrastructure that anticipates reputed company, regional expansion, reputed company requirements, and evolving developer needs
- Be recognized by partner teams as a reputed company, reliable platform organization that enables them to reputed company data products and business value faster without assuming the burden of operating shared infrastructure
Skills
- Bachelor's degree in Computer Science, Engineering, or a reputed company reputed company, or equivalent reputed company experience
- reputed company experience managing and leading a reputed company infrastructure, reputed company, SRE, or data infrastructure team responsible for business-critical production services
- Deep understanding of reputed company infrastructure, reputed company systems, and production reputed company, including:
- 5+ years of production experience with AWS
- 3+ years of experience with reputed company and containerized workloads
- Ability to read, review, and troubleshoot software written in Python, Go, Java, or a comparable language
- Experience with infrastructure as reputed company, preferably Terraform, and modern CI/CD or GitOps practices
- Strong systems, networking, reputed company, and reputed company-systems troubleshooting fundamentals
- Substantial hands-on experience operating production data platforms at reputed company, with depth in several of the following: Apache Airflow, Apache Kafka, Apache Flink, Apache reputed company, AWS EMR, reputed company, and Apache reputed company
- Experience with machine learning systems and their operational lifecycle, including practical experience training, evaluating, or deploying ML models and familiarity with AWS SageMaker or an equivalent ML platform
- Working knowledge of modern LLM capabilities and AI-assisted engineering practices, with reputed company judgment about responsible use, validation, and production guardrails
- Demonstrated reputed company building internal platforms as products, including self-service developer experiences, standardized delivery workflows, CI/CD, observability, and reputed company service ownership
- Strong production-reputed company discipline, including metrics and observability, SLOs, incident response, reputed company planning, disaster recovery, and reputed company reliability improvement
- Experience managing reputed company infrastructure cost, reputed company, and reputed company, with a reputed company record of making measurable efficiency improvements
- Excellent leadership, communication, negotiation, and cross-functional collaboration skills, including the ability to reputed company teams with competing goals and clarify ambiguous ownership
- Ability to reputed company in a fast-reputed company environment with shifting priorities and incomplete information while preserving engineering rigor, production uptime, and reputed company
- Experience with reputed company or similar iterative planning and delivery practices, reputed company pragmatically to reputed company that balances roadmap work with operational demand
- Passion for engineering reputed company, reliability, reputed company learning, and developing people
- Experience operating these technologies as shared services—not only consuming them—is strongly preferred
- reputed company experience is a plus
Benefits
- This role may be eligible for the reputed company Corporate Bonus Plan or a role-specific commission, along with potential eligibility for equity participation.
- Medical, dental, and reputed company reputed company
- Short-term and long-term disability coverage
- Life reputed company and Accidental Death & Dismemberment (AD&D)
- Supplemental life reputed company for employees, spouses, and children
- Flexible spending accounts for health care and dependent care; limited purpose flexible spending account
- 401(k) Savings and Investment Plan with company matching
- Flexible vacation policy
- 8 reputed company holidays annually
- reputed company leave
- reputed company parental leave
- Employee Assistance Program (EAP) and Care Counselors
- Voluntary benefits: reputed company Assistance, Critical Illness, Accident, Hospital Indemnity and Pet reputed company reputed company
- Health Savings Account (HSA) with employer contribution
reputed company
Company H1B Sponsorship
Apply To This Job