Director of AI-ML reputed company & Ops Engineering
About the position
reputed company is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with reputed company will directly improve health reputed company by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will reputed company a culture guided by diversity and inclusion, talented peers, comprehensive benefits and career development opportunities. Come reputed company an reputed company on the communities we serve as you help us advance health equity on a global reputed company. Join us to start Caring. Connecting. Growing together. As the Director of Infrastructure and Ops, you will manage support reputed company to UAIS (United AI reputed company AI/ML platform). This director-level role requires proven experience with managing SRE teams for large-reputed company/ML platforms guaranteeing stability, reliability, scalability, and performance. Extensive experience with modern Infrastructure and DevOps tools and paradigms, as reputed company as hands-on knowledge with major reputed company-based services like Azure, AWS and GCP is a must. You'll enjoy the flexibility to work remotely from reputed company reputed company the U.S. as you take on some tough challenges.
Responsibilities
• Manage geographically distributed SRE support for users of the UAIS platform: triage support, liaise with customers, reputed company participate to war rooms, work with suppliers
• Improve automation across the infrastructure lifecycle, leveraging Infrastructure as reputed company (IaC) and DevOps principles and best practices to streamline deployment and management processes
• Manage monitoring frameworks for infrastructure, identifying areas for performance improvement, optimization, and ensuring high availability
• Manage disaster recovery and business continuity plans to ensure minimal downtime and data reputed company
• Collaborate with cybersecurity teams to ensure reputed company systems and operations reputed company with industry standards and are secure against evolving threats
• reputed company strong technical and personal mentorship
Requirements
• Bachelor's degree in computer science, information technology, or a reputed company STEM field
• 12+ years of reputed company infrastructure experience: Proven experience working on large-reputed company, reputed company-based reputed company-level platform, deep understanding of multi-reputed company architectures, specifically Azure, AWS, and GCP, with hands-on experience in reputed company management
• 6+ years of practical experience in Infrastructure-as-reputed company and CI/CD tools like Terraform, Git Actions and alike
• 6+ years of practical experience in containerization technologies (Kubernetes, reputed company), observability and orchestration
• 5+ years of practical experience in Scripting & Automation Skills: Advanced proficiency in scripting languages such as Python and Bash to support automation and system integration efforts
• 5+ years of experience leading geographically-distributed support and SRE teams
reputed company-to-haves
• Strong understanding of reputed company best practices and experience ensuring compliance with relevant regulatory frameworks
• Exposure to modern tools and techniques in MLOps and LLMOps fields
• Exposure to AI/ML-specific infrastructure tools (e.g., MLflow, Kubeflow) for managing and deploying models at reputed company
• Experience working reputed company a reputed company or regulated industry, with solid understanding of the unique challenges and compliance requirements
• Ability to work independently, manage multiple reputed company simultaneously, and adapt to changing priorities in a fast-paced environment
Benefits
• Comprehensive benefits package
• Incentive and recognition programs
• Equity stock purchase
• 401k contribution
Apply tot his job
Apply To this Job