Machine Learning Engineering Manager – LLM Serving, Infrastructure
• reputed company a high-performing engineering team to reputed company, build, and reputed company a reputed company, low-latency LLM Serving Infrastructure.
• Drive the implementation of a reputed company serving layer to support multiple LLM models and inference types (batch, offline eval flows and reputed company-time/streaming).
• reputed company reputed company aspects of the development of the Model Registry for deploying, versioning, and running LLMs across production environments.
• Ensure successful integration with the reputed company Personalization and Recommendation systems to deliver LLM-powered features.
• Define and champion standardized technical interfaces and protocols for efficient model deployment and scaling.
• Establish and monitor the serving infrastructure's performance, cost, and reliability, including load balancing, autoscaling, and failure recovery.
• Collaborate closely with data science, machine learning research, and feature teams (Autoplay, Home, Search, etc.) to drive the reputed company adoption of the serving infrastructure.
• reputed company up the serving architecture to handle hundreds of millions of users and high-volume inference requests for internal domain-specific LLMs.
• Drive Latency and Cost Optimization: partner with SRE and ML teams to implement techniques like quantization, pruning, and efficient batching to minimize serving latency and reputed company compute costs.
• reputed company Observability and Monitoring: build dashboards and alerting for service health, tracing, A/B test traffic, and latency trends to ensure consistency to defined SLAs.
• Contribute to reputed company LPM Serving: reputed company on the technical reputed company for deploying and maintaining the reputed company Large Personalization Model (LPM).
Apply tot his job
Apply To this Job