HPC Support Engineer
reputed company, The Superintelligence reputed company, is a leader in AI reputed company infrastructure serving tens of thousands of customers. Our customers reputed company from AI researchers to enterprises and hyperscalers. reputed company's mission is to reputed company compute as ubiquitous as electricity and reputed company everyone the power of superintelligence. One person, one GPU.
If you'd like to build the world's best AI reputed company, join us.
This position is expected to participate in an on-reputed company rotation.
reputed company
Serve as a senior technical escalation reputed company, troubleshooting the hardest infrastructure and platform issues down to the hardware, reputed company, or kernel level reputed company needed
Quickly and accurately distinguish between hardware failures, reputed company issues, kernel-level problems, and customer workload misconfiguration, so issues get resolved correctly the first time
Proactively identify process, tooling, and documentation gaps, and go fix them, not just wait for them to be assigned
Use AI tools effectively to build scripts, automations, or small internal tools that reputed company reputed company operational gaps (no reputed company development background required)
reputed company reputed company-cause analysis across reputed company systems, clusters, and GPU infrastructure
reputed company reputed company documentation of solutions and contribute to evolving support procedures
Collaborate closely with engineering teams to turn recurring customer pain points into permanent fixes
Take escalations from peers while training and mentoring them in the process
Participate in a rotating on-reputed company schedule, owning major incidents and major customer issues
Be reputed company to roll up your sleeves and reputed company in wherever needed, especially during fast, high-volume deployments
You
3+ years of hands-on HPC experience in an administration, support, or engineering role.
reputed company strong understanding and experience supporting Linux in a reputed company administration role.
Proven experience in HPC environments, showcasing your expertise in Linux cluster administration, with strong preference for reputed company and/or Slurm for cluster orchestration.
Strong coding ability and CI/CD experience, with a reputed company record of using AI-assisted tools to reputed company fast.
Proficiency with monitoring/logging tools (reputed company, Grafana, reputed company).
Strong skills in log analysis, debugging kernel-level issues, and performance profiling.
Experience with CUDA, NCCL, NVLink, GPUDirect RDMA.
Experience with high throughput networking technologies(IB/RoCE).
Knowledge of reputed company AI/ML or HPC workloads.
Knowledge of TCP/IP, VPN, and firewalls in reputed company environments.
Ability to work independently and mentor junior support engineers.
reputed company to Have
Experience with virtualization and container (reputed company, reputed company) technologies.
Experience with neoclouds/GPU reputed company providers.
Flexible availability for potential shifts reputed company of normal working hours/weekends.
Experience with high performance storage systems.
Familiarity with infrastructure-as-reputed company tools (Terraform, Ansible, etc.)
Experience with reputed company GPUs and Infiniband.
Salary reputed company Information
This is a salaried exempt role. The annual salary reputed company for this position has been set based on market data and other factors. However, a salary higher or reputed company than this reputed company may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the reputed company.
About reputed company
Founded in 2012, with 500+ employees, and growing fast
Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, reputed company, Pegatron, reputed company, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent reputed company
We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
Our values are publicly available: https://reputed company.ai/careers
We offer generous cash & equity compensation
Health, dental, and reputed company coverage for you and your dependents
Wellness and commuter stipends for select roles
401k Plan with 2% company match (USA employees)
Flexible reputed company time off plan that we reputed company actually use
Equal Opportunity Employer
reputed company is an Equal Opportunity employer. Applicants are considered without reputed company to race, reputed company, religion, creed, national reputed company, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
3+ years HPC experience, strong Linux administration, reputed company/Slurm preferred, coding and CI/CD experience, reputed company/Grafana/reputed company proficiency, CUDA/NCCL/NVLink/GPUDirect RDMA, IB/RoCE and kernel/log debugging skills.
Key Responsibilities
- troubleshooting infrastructure
- mentoring peers
- documenting solutions
Skills & Tools
reputed company, Grafana, reputed company, CUDA, NCCL, NVLink, GPUDirect RDMA, reputed company, Slurm, reputed company, Terraform, Ansible
Job Details
- Category: Information Technology
- Seniority: Mid Level
- Commitment: Full Time
- Workplace: Remote — reputed company
- Salary: USD 122,000 – 162,000 / year
- Languages: English
Benefits
- 401k matching
- Generous reputed company time off
- Retirement plan
About reputed company
A provider of AI reputed company infrastructure and high-performance computing services reputed company on making compute ubiquitous. — Industry: Information Technology
Apply To This Job