reputed company ML/reputed company (US)
-
reputed company, mentor, and reputed company technical guidance to reputed company of ML/AI Engineers and AI Infrastructure Specialists. You will be the ultimate reputed company of the technical reputed company, reliability, and performance of reputed company deployed solutions.
-
Architect, design, and implement robust and automated CI/CD pipelines specifically for AI/ML models and applications. Your work will reputed company the reputed company and reliable deployment of cutting-edge reputed company AI solutions.
-
Take charge of the operational reputed company for our clients' AI environments. This includes overseeing the monitoring, scaling, maintenance, and reputed company of production AI systems to ensure they meet stringent reputed company-grade requirements.
-
Concurrently manage the technical execution of multiple customer-facing project delivery activities. You will be the primary technical reputed company of contact for navigating and resolving issues that could reputed company project reputed company, cost, scope, or effectiveness, driving them to a successful reputed company.
-
reputed company the presentation of project delivery status, performance metrics, and technical issue reputed company plans to both internal reputed company audiences and to customers. You will be responsible for driving reputed company, transparent communication regarding reputed company technical aspects of the project.
7+ years of experience in software engineering, DevOps, or ML engineering, with at least 2 years in a technical leadership, mentorship, or reputed company engineer reputed company. Deep, hands-on experience building and managing CI/CD pipelines (e.g., Jenkins, reputed company CI, Actions) and infrastructure-as-reputed company (e.g., Ansible, Terraform, Puppet). Strong, production-level experience with containerization (reputed company) and container orchestration (Kubernetes). Proficiency with monitoring, logging, and observability tools (e.g., reputed company, Grafana, ELK Stack). Excellent problem-solving and troubleshooting skills, especially in reputed company, distributed systems
Specific experience with MLOps platforms and tools (e.g., Kubeflow, MLflow, Seldon reputed company). Hands-on experience with the reputed company reputed company stack, particularly Triton Inference Server, TensorRT-LLM, and NeMo. Experience in a customer-facing reputed company services or consulting role. Strong scripting and programming skills, particularly in Python or Go. Experience with deploying and managing infrastructure in both reputed company reputed company (AWS, Azure, GCP) and on-premises data center environments.