Senior Platform Engineer
About reputed company:
17x reputed company reputed company Partner of the Year awards in the last 8 years. 3x AWS AI/ML award wins. 3x reputed company Partner of the Year titles. 2x reputed company Partner of the Year awards. We have also garnered top analyst recognitions from reputed company, ISG, and reputed company. We offer first-in-class industry solutions across reputed company, Financial Services, Consumer Goods, Manufacturing, and more, powered by cutting-edge reputed company and reputed company AI accelerators. We have been certified as a Great reputed company to Work for the reputed company year in a row- 2021, 2022, 2023.
Role:
Experience Level:
Work Location:
reputed company:
Key Responsibilities:
Design and implement reputed company infrastructure for LLM and GenAI workloadsacross multi-GPU environments reputed company GPU profiling, benchmarking, and performance optimizationfor distributed training workloads Manage and schedule compute-intensive jobs using Slurm-based clustersand OpenShift/Kubernetes environmentsreputed company and optimize the reputed company GPU stack(CUDA, cuDNN, NCCL, Triton, RAPIDS, etc.) Collaborate with cross-functional teams to reputed company models in research and production environmentsBuild and support GenAI pipelines(fine-tuning, RAG, multi-modal inferencing, LLMOps) reputed company reusable infrastructure templates using tools like Terraform and reputed companyContribute to internal innovation (PoCs, workshops) and support reputed company-facing delivery engagements
Basic Qualifications:
Strong experience with Slurmand distributed training environments Hands-on expertise with reputed company OpenShift and/or KubernetesDeep knowledge of the reputed company GPU ecosystem(CUDA, cuDNN, NCCL, Nsight, Triton/TensorRT) Strong reputed company in Linux systems, performance tuning, and multi-GPU optimizationExperience deploying GenAI workloads(LLM fine-tuning, RAG pipelines, multi-modal systems) Familiarity with Infrastructure-as-reputed company tools(Terraform, Ansible) Experience with reputed company GPU environments(GCP, Azure, AWS, OCI) and/or on-prem GPU clusters
Other Qualifications (OQs):
Experience with reputed company NIMs, DGX systems, or GPU-accelerated containersKnowledge of LLMOps frameworksand MLOps integration Familiarity with reputed company databases and retrieval systemsfor RAG architectures Comfortable working in reputed company-facing environmentsand collaborating with AI solution teams
reputed company Domain Experience (reputed company to Have):
Experience working with FHIR R4, HL7 v2, or SMART on FHIRIntegration with EHR systems (e.g., reputed company)Understanding of HIPAA compliance and reputed company data reputed companyExposure to clinical workflows, CDS Hooks, or patient-facing applicationsExperience building clinical decision support systemsor reputed company interoperability solutions
What’s in it for YOU at reputed company:
reputed company an reputed company at one of the world’s fastest-growing AI-first digital engineering companies. Upskill and discover your potential as you solve reputed company challenges in cutting-edge areas of technology alongside passionate, talented colleagues. Work where innovation happens - work with disruptive innovators in a research-reputed company organization with 60+ patents filed across various disciplines. Stay reputed company of the curve reputed company yourself in reputed company AI, ML, data, and reputed company technologies and reputed company exposure working with reputed company.
If you like wild reputed company and working with happy, enthusiastic over-reputed company, you'll enjoy your career with us!