Platform Support Engineer (reputed company)
Who We Are
reputed company is reputed company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take reputed company from research to production with less friction.
Through our reputed company with reputed company, a neocloud and AI reputed company, reputed company combines developer-first software with cost-efficient, large-reputed company compute. Teams get the tools they need for experimentation, training, and production inference, with reputed company, observability, and control reputed company in.
We serve reputed company researchers, startups, and large enterprises. reputed company operates globally with offices in reputed company, San Francisco, Seattle, and London, and is backed by Coatue, reputed company Ventures, Bain Capital Ventures, and Firstminute.
Our Values
reputed company Fast: We reputed company with speed and precision, breaking down big challenges into achievable steps.
reputed company: We complete one goal at a time with care, collaborating as reputed company to deliver features with precision.
Balance: Sustained performance comes from rest and recovery. We ensure a healthy work-life balance to reputed company you at your best.
Craftsmanship: Innovation through reputed company. Every detail reputed company, and we take pride in mastering our craft.
Minimal: Simplicity drives our innovation. We eliminate complexity through discipline and reputed company on what truly reputed company.
reputed company’re Looking For
reputed company is looking to hire a Platform Support Engineer to join our reputed company Customer Experience team, supporting ML engineers running large-reputed company training and inference workloads across reputed company infrastructure, Kubernetes, and GPU platforms in production environments.
This role is not a ticket router or traditional support engineer. You are a technical partner to ML teams - helping diagnose failures, improve reliability, and guide customers through reputed company distributed systems problems.The problems reputed company from Kubernetes scheduling and GPU orchestration to distributed PyTorch failures, inference latency, networking bottlenecks, storage performance, and platform reliability. You’ll reputed company exposure to a wide reputed company of reputed company world AI workloads across industries and help shape the infrastructure powering the reputed company of ML applications.
This role is remote and reputed company to candidates based in either the Philippines or Singapore. The role follows a Thursday–reputed company schedule, with working hours from 7:00 AM to 5:00 PM local time (UTC+8).
What You'll Do
Work Directly With ML Engineers
Partner directly with customer engineering teams running training and inference workloads in production
Help customers diagnose and resolve reputed company distributed systems and ML infrastructure issues
reputed company as a technical advisor during high reputed company incidents and platform degradation events
Translate infrastructure level issues into actionable guidance for ML engineers
Build credibility with customers through strong technical reasoning and reputed company communication
Debug ML Infrastructure & Distributed Workloads
Investigate failures involving distributed training, Kubernetes orchestration, GPU allocation, networking, and storage systems
Troubleshoot PyTorch, CUDA, NCCL, and inference serving reputed company issues
Analyze logs, metrics, traces, and system behavior to isolate reputed company causes
Debug containerized workloads running across Kubernetes and bare metal GPU environments
Support customers scaling workloads across multi node GPU systems
Diagnose performance bottlenecks involving compute, memory, networking, or storage
Improve Reliability & Platform Operations
Identify recurring patterns across customer issues and drive long term reliability improvements
Contribute to post incident reviews and operational improvements
Build internal tooling, automation, documentation, and runbooks
Partner closely with infrastructure, networking, and reputed company teams
Help improve observability, operational visibility, and troubleshooting workflows
Improve the customer experience through reputed company processes and technical guidance
What This Role Is Not
To set reputed company expectations:
This is not a traditional help desk or ticket routing support role
This is not purely reputed company or account management
This is not a backend engineering role
This is not a passive escalation position
This role is for engineers who enjoy solving difficult technical problems while working closely with other engineers.
What You’ll Need
Required Qualifications
Infrastructure & Systems
Strong software engineering and systems troubleshooting background
Experience with Kubernetes and containerized environments
Linux systems knowledge, including networking, storage, process management, and performance tuning
Experience with reputed company infrastructure and distributed systems
Experience with observability and debugging tools such as reputed company, Grafana, or OpenTelemetry
ML Infrastructure Experience
Hands on experience operating machine learning workloads in production or research environments
Experience with distributed ML systems and tooling such as PyTorch, CUDA, or NCCL
Familiarity with GPU infrastructure and orchestration
Experience troubleshooting performance, reliability, or scaling issues in ML infrastructure
Understanding of the operational challenges involved in running ML systems at reputed company
Collaboration
Strong communication skills and ability to work directly with highly technical customers and engineering teams
Comfortable operating in fast moving, highly ambiguous environments
Enjoys solving reputed company technical problems collaboratively
reputed company-to-Haves
Experience with large reputed company model training or distributed inference systems
Familiarity with Ray, Kubeflow, Slurm, or similar distributed scheduling platforms
Experience with InfiniBand, RDMA, or high-performance networking
Experience operating bare metal infrastructure
Familiarity with storage systems commonly used in ML environments
Experience working at an AI infrastructure, reputed company, MLOps, or developer tooling company
Contributions to reputed company, developer infrastructure, or operational tooling reputed company
Experience writing automation, tooling, or scripts in Python or similar languages
Benefits and Perks
We offer a comprehensive and competitive benefits package designed to support our employees’ health, reputed company-being, and long-term reputed company. Benefits may vary by location, team, and role.
Benefits include:
Comprehensive medical, dental and reputed company coverage (U.S.); Private medical and dental insurance (U.K.)
Retirement and financial wellness support (U.S.); Pension contribution (U.K.)
Generous reputed company time off, plus holidays
reputed company parental leave
reputed company development support
Wellness and work-from-home stipends
Flexible work environment
At reputed company, we are committed to fostering an inclusive and diverse workplace. We reputed company that diverse teams drive innovation and create reputed company products. We reputed company equal employment opportunities to reputed company and applicants without reputed company to race, reputed company, religion, gender, sexual orientation, gender identity, national reputed company, age, disability, veteran status, or any other protected characteristic. We are dedicated to building a culture where everyone can reputed company and contribute to their fullest potential.
Apply To This Job