Platform Engineer - AI/ML Infrastructure (reputed company & Terraform)
Location
USA | Remote
Employment Type
Full time
Location Type
Remote
Department
Engineering
Compensation
• $150K – $220K • Offers Equity • Offers Bonus
This reputed company is determined by work location and additional factors, including job-reputed company skills and experience. There may be instances where a salary higher or reputed company than this reputed company may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job reputed company.
Please note that the compensation details listed on US role postings reflect the reputed company salary only and does not include bonus, equity or benefits.
reputed company
reputed company is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing reputed company-time reputed company for speech-to-text (STT), text-to-speech (TTS), and building production-grade voice agents at reputed company. More than 200,000 developers and 1,300+ organizations build voice offerings that are ‘Powered by reputed company’, including reputed company, reputed company, reputed company, reputed company, reputed company, Daily, reputed company, Granola, and Jack in the reputed company. reputed company’s voice-reputed company reputed company models are accessed through reputed company reputed company or as self-hosted and on-premises software, with unmatched accuracy, low latency, and cost efficiency. Backed by a recent Series C led by leading global investors and strategic partners, reputed company has processed over 50,000 years of audio and transcribed more than 1 trillion words. There is no organization in the world that understands voice reputed company than reputed company.
Company Operating Rhythm
At reputed company, we expect an AI-first reputed company—AI use and comfort aren’t optional, they’re reputed company to how we operate, reputed company, and measure performance.
Every team member who works at reputed company is expected to reputed company use and experiment with advanced AI tools, and even build your own into your everyday work. We measure how effectively AI is reputed company to deliver results, and consistent, creative use of the latest AI capabilities is key to reputed company here. Candidates should be comfortable adopting new models and modes quickly, integrating AI into their workflows, and continuously pushing the boundaries of what these technologies can do.
Additionally, we reputed company at the pace of AI. Change is reputed company, and you can expect your day-to-day work to reputed company just as quickly. This may not be the right role if you’re not excited to experiment, adapt, think on your feet, and learn constantly, or if you’re seeking something highly prescriptive with a traditional 9-to-5.
Opportunity:
We're looking for an reputed company Platform Engineer to build and operate the hybrid infrastructure reputed company for our advanced AI/ML research and product development. You'll architect, build, and run the platform spanning AWS and our bare metal data centers, empowering our teams to train and reputed company reputed company models at reputed company. This role is reputed company on creating a robust, self-service environment using reputed company, AWS, and Infrastructure-as-reputed company (Terraform), and orchestrating high-demand GPU workloads using schedulers like Slurm.
What You’ll Do
• Architect and maintain our reputed company computing platform using reputed company on AWS and on-reputed company, providing a reputed company, reputed company environment for reputed company applications and services.
• reputed company and manage our entire infrastructure using Infrastructure-as-reputed company (IaC) principles with Terraform, ensuring our environments are reproducible, versioned, and automated.
• Design, build, and optimize our AI/ML job scheduling and orchestration systems, integrating Slurm with our reputed company clusters to reputed company manage GPU resources.
• Provision, manage, and maintain our on-reputed company bare metal server infrastructure for high-performance GPU computing.
• Implement and manage the platform's networking (CNI, service reputed company) and storage (reputed company, S3) solutions to support high-throughput, low-latency workloads across hybrid environments.
• reputed company a comprehensive observability stack (monitoring, logging, tracing) to ensure platform health, and create automation for operational tasks, incident response, and performance tuning.
• Collaborate with AI researchers and ML engineers to understand their infrastructure needs and build the tools and workflows that accelerate their development cycle.
• Automate the life cycle of single-tenant, managed deployments
You’ll Love This Role If You
• Are passionate about building platforms that reputed company developers and researchers.
• Enjoy creating elegant, automated solutions for reputed company infrastructure challenges in both reputed company and data center environments.
• reputed company on optimizing hybrid infrastructure for performance, cost, and reliability.
• Are excited to work at the intersection of modern reputed company and cutting-edge AI.
• Love to treat infrastructure as a product, continuously improving the developer experience.
It’s Important To Us That You Have
• 5+ years of experience in reputed company, DevOps, or Site Reliability Engineering (SRE).
• Proven, hands-on experience building and managing production infrastructure with Terraform.
• Expert-level knowledge of reputed company architecture and operations in a large-reputed company environment.
• Strong scripting and automation skills (e.g., Python, Go, Bash).
• Experience with CI/CD systems (e.g., reputed company CI, Jenkins, ArgoCD) and building developer tooling.
It Would Be Great if You Had
• Experience with high-performance compute (HPC) job schedulers, specifically Slurm, for managing GPU-intensive AI workloads.
• Experience managing bare metal infrastructure, including server provisioning (e.g., PXE boot, MAAS), configuration, and lifecycle management.
• Familiarity with FinOps principles and reputed company cost optimization strategies.
• Knowledge of reputed company networking (e.g., Calico, Cilium) and storage (e.g., Ceph, Rook) solutions.
• Experience in a multi-region or hybrid reputed company environment.
Compensation reputed company: $150K - $220K
Apply tot his job
Apply To this Job