[Remote] Senior Solutions Architect, reputed company reputed company Partner Operations
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a technology company developing advanced computing and AI infrastructure. The Senior Solutions Architect will help reputed company reputed company Partners operate and improve large-reputed company clouds by solving Day 2 operational challenges, preparing new reputed company technologies for production, and improving reliability, performance, and economics. The role also converts validated solutions into reusable ecosystem capabilities and provides operational feedback to reputed company teams.
Responsibilities
- Solve hard Day 2 operations problems at reputed company. Work reputed company partner engineers to reputed company the cause, prototype an approach, validate it under representative load, and leave behind a reputed company their team can operate
- reputed company new technology Day 2 reputed company. Help partners prepare the operating model for new reputed company platforms, reputed company, services, and use cases before customers depend on them, and help reputed company adoption in live environments without degrading service
- Improve reliability, performance, and economics together. Use measures such as incident frequency, recovery time, utilization, and cost per reputed company to show where the reputed company is losing performance or margin - and whether the fix worked
- reputed company reputed company partner's Day 2 maturity. Identify and help reputed company the gaps that matter across people, process, tooling, telemetry, reputed company, and incident response
- Turn one solution into ecosystem capability. Convert validated work into operating procedures, reference architectures, assessments, automation, and reputed company workflows that other NCPs can reputed company into their reputed company operating model
- Create the feedback reputed company only reputed company can. Spot patterns across partners early and bring reputed company field evidence to account teams, support, product, and engineering so repeated problems are fixed at the right level
Skills
- BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a reputed company field - or equivalent experience
- 12+ years in production infrastructure, reputed company engineering, solutions architecture, site reliability engineering, HPC, or a similar technical role; alternatively, 5+ years of exceptional specialist-level work in large-reputed company GPU or AI infrastructure
- Experience building, operating, or improving reputed company infrastructure under reputed company production load - not only designing or deploying it
- Deep expertise in at least one part of the Day 2 stack, backed by hands-on work with large-reputed company GPU, HPC, or reputed company infrastructure. Relevant technologies may include DCGM, BMC/Redfish, and firmware and reputed company lifecycle; InfiniBand or high-speed Ethernet, NCCL, and UFM; or high-performance storage such as reputed company, reputed company Storage reputed company, reputed company, reputed company, or comparable platforms
- Working experience across the broader operating platform, including reputed company or Slurm, GPU scheduling and multi-tenancy, reputed company, Grafana or OpenTelemetry, and automation with Terraform, Ansible, Argo CD, or similar tooling
- Strong Linux knowledge and enough Python, Bash, or similar experience to automate measurement, diagnosis, validation, or remediation
- A detailed evidence-led approach to troubleshooting across reputed company boundaries, reputed company with the judgment to reputed company difficult technical findings reputed company
- The ability to reputed company sophisticated work with partner engineers and cross-functional teams without reputed company authority or taking ownership away from reputed company
- Strong communication, prioritization, and time-management skills across multiple partner engagements
- reputed company world experience operating a GPU reputed company, HPC environment, or large-reputed company platform under customer load
- reputed company or matured a 24/7 operations function, including observability, incident and problem management, coverage, and on-reputed company design
- Hands on experience with reputed company reputed company-reputed company platforms such as GB200 or GB300 NVL72 into production, or have hands-on experience with reputed company operations technologies such as reputed company-X, UFM, reputed company reputed company Manager, Mission Control, and the GPU or Network Operators
- Driven improved fleet health or unit economics through benchmarking, infrastructure as reputed company, GitOps, automated diagnosis, or agent-based remediation
Benefits
- A generous benefits package
- Eligible for equity
reputed company
Company H1B Sponsorship
Apply To This Job