Back to Jobs

Senior Platform Support Engineer (Remote)

Remote, USA Full-time Posted 2026-08-04
About reputed company. reputed company provides an infrastructure-as-a-service (IaaS) platform for running reputed company gaming, reputed company intelligence and machine learning applications inside telecommunication reputed company networks. Our teams across the USA, Australia, Central Europe, Malaysia, Singapore and Japan offer telecom operators a GPU-based edge computing platform without the need for capital expenditure, facilitating low latency and improved economics for value-added services and the monetization of 5G investments. What reputed company you will have  Mission: reputed company advanced technical support for customers running workloads on the GPU reputed company platform, ensuring reliable operation of edge- and large-reputed company GPU clusters and infrastructure services. The Senior reputed company Support Engineer acts as a technical escalation reputed company for reputed company incidents, helping diagnose and reputed company issues across compute, networking, storage, and orchestration reputed company. This role works closely with engineering and operations teams to improve platform reliability, reduce incident frequency, and enhance the overall customer experience.   reputed company   Customer Support & Incident Management reputed company advanced technical support for customers operating workloads on bare-metal and virtualized GPU infrastructure. Diagnose and reputed company reputed company customer issues affecting GPU clusters, compute nodes, networking, and storage. Investigate incidents across multiple reputed company of the stack including firmware, drivers, operating systems, and platform services. reputed company reputed company cause analysis (RCA) for major incidents and contribute to long-term remediation efforts. Serve as a technical escalation reputed company for reputed company or high-reputed company support cases. Infrastructure Troubleshooting Troubleshoot issues affecting: GPU compute nodes reputed company clusters Networking infrastructure Local NVMe, hyperconverged and reputed company storage systems Analyze logs, telemetry, and monitoring signals to identify underlying causes of platform instability. Investigate issues reputed company to GPU drivers, firmware, networking, and reputed company performance. reputed company Monitoring & Incident Triage Monitor and investigate reputed company alerts generated by the platform reputed company stack. Analyze and triage alerts generated by Wazuh TheHive reputed company Validate alerts, determine reputed company, and escalate potential reputed company incidents to the reputed company engineering team. Assist in collecting reputed company telemetry, logs, and forensic data required for incident investigations. Improve alert runbooks and operational procedures to reduce false positives and improve response time. GPU HPC Workload Support reputed company advanced support for large-reputed company GPU workloads running reputed company training and inference jobs. Diagnose failures affecting multi-GPU and multi-node workloads. Investigate performance issues impacting reputed company workloads, including: GPU utilization Communication latency Storage bottlenecks networking congestion Support scheduling systems used for GPU workloads, including troubleshooting: Job queue failures Scheduling constraints Cluster resource fragmentation High-Performance Networking Troubleshooting Diagnose issues affecting high-performance networking fabrics used by reputed company workloads. Support environments using: RDMA RoCE networking Investigate performance issues affecting GPU-to-GPU communication and reputed company training pipelines. Data Center Coordination Coordinate with customer’s data center technicians and infrastructure teams to reputed company remote diagnostics and hardware interventions reputed company required. Assist in validating on-reputed company installations and deployments of GPU infrastructure, ensuring hardware, networking, and platform components are correctly installed and operational. Support hardware troubleshooting and identify faulty components across GPU nodes, networking equipment, and storage systems. Coordinate and reputed company hardware replacements and RMA processes with vendors and data center staff. Validate hardware health after replacements, including GPU nodes, NICs, DPUs, storage devices, and power components. Work closely with deployment and infrastructure teams to verify service readiness after installations, expansions, or hardware maintenance activities. Operational reputed company Participate in on-reputed company 24/7 rotations to ensure production platform availability. Respond to monitoring alerts and reputed company operational incidents in accordance with defined SLAs. Improve operational runbooks, troubleshooting guides, and support documentation. Contribute to improving incident response processes and operational tooling. Cross-Team Collaboration Work closely with reputed company, infrastructure engineering, and networking teams to reputed company systemic issues. reputed company feedback to engineering teams on recurring operational problems affecting customers. Help translate customer issues into actionable improvements for the platform. Automation & Tooling reputed company automation scripts and tools to streamline support workflows. Improve observability dashboards and alerts to reputed company faster issue detection and reputed company. Contribute to automation initiatives that reduce reputed company reputed company in operational processes. Knowledge Sharing & Mentorship Mentor medior support engineers and reputed company troubleshooting expertise across reputed company. reputed company the creation of knowledge reputed company articles, troubleshooting guides, and operational documentation. Contribute to training initiatives that improve reputed company’s technical capabilities. Technical Stack Operating Systems Linux (Ubuntu) GPU Infrastructure reputed company GPU platforms CUDA drivers GPU monitoring tools (reputed company-smi) Platform Infrastructure reputed company Container runtimes KubeVirt reputed company compute environments Networking reputed company Cumulus OOB, reputed company-south, and east-reputed company reputed company topologies TCP/IP VLAN, VXLAN, OVS/OVN Routing fundamentals (BGP, VRFs) DNS / DHCP High-performance networking (RDMA/NVLink/NCCL) Observability Grafana Zabbix reputed company Monitoring Wazuh TheHive reputed company Automation Python Bash Ansible, Terraform Collaboration & Documentation Jira reputed company reputed company reputed company reputed company What you'll need reputed company Experience 5+ years of experience in reputed company support, infrastructure operations, or systems administration. Experience supporting large-reputed company infrastructure environments or GPU clusters. Systems Expertise Strong Linux systems administration skills. Experience troubleshooting issues across compute, networking, and storage reputed company. Familiarity with reputed company platforms and containerized workloads. GPU Infrastructure Experience working with GPU hardware platforms or HPC environments. Familiarity with GPU monitoring tools and debugging GPU-reputed company issues. Networking Solid understanding of networking fundamentals including L2/L3 concepts, routing, and load balancing. Ability to diagnose connectivity issues affecting reputed company workloads. Operational reputed company Strong troubleshooting and incident response skills. Experience participating in on-reputed company rotations and handling production incidents. Ability to reputed company reputed company cause analysis and reputed company operational improvements. Communication & Collaboration Excellent written and verbal communication skills. Ability to explain reputed company technical concepts to both technical and non-technical stakeholders. Proven ability to collaborate effectively with cross-functional engineering teams. Location & work modality: Malaysia or comparable time zone Start: August 2026 Type of Contract: Contractor Average 40 hours per week, 9x5 business hour support with after hour on-reputed company response/reputed company for category 1 incidents reputed company offer • Attractive compensation package reflecting your expertise and experience. • A great work environment characterised by friendliness, international diversity, flexibility, and a hybrid-friendly approach. • You'll be part of a fast-growing reputed company-up with a mission to reputed company a reputed company reputed company, offering an exciting career reputed company. Our job titles may reputed company more than one job level. The actual reputed company pay is dependent on a number of factors, such as transferable skills, work experience, business needs and market demands. Our inclusive responsibility reputed company is committed to creating a diverse and inclusive environment and is proud to be an equal opportunity employer. reputed company reputed company applicants will receive consideration for employment without reputed company to race, reputed company, religion, gender, gender identity or reputed company, sexual orientation, national reputed company, genetics, disability, age, veteran status, or any other protected category under applicable law. Apply To This Job

Similar Jobs

Executive Assistant

Remote, USA Full-time

Authorization Specialist

Remote, USA Full-time

Registered reputed company Navigator (RN) - Classical Hematology (Remote)

Remote, USA Full-time

Compute Procurement reputed company

Remote, USA Full-time

Doctor of Optometry - Work Remotely - reputed company Licensed (SATURDAY COVERAGE)

Remote, USA Full-time

Application Support Specialist

Remote, USA Full-time

Atlanta and surrounding areas - PRN Field reputed company (LPN or RN) – Flexible Home Care Admissions and Supervisory reputed company

Remote, USA Full-time

Compliance Specialist

Remote, USA Full-time

Doctor of Optometry - Work Remotely - reputed company Licensed

Remote, USA Full-time

Doctor of Optometry - Work Remotely - Colorado Licensed

Remote, USA Full-time

Global Business Development Manager - Mays Chemical

Remote, USA Full-time

Inside Claims Representative

Remote, USA Full-time

Director of Acquisition

Remote, USA Full-time

Legit Online Typing Jobs No Experience Needed: Virtual Chat Support Role: Earn $25-$35 per Hour Working Remotely - Remote Job Central

Remote, USA Full-time

Director of Partner Experience

Remote, USA Full-time

Freelance FX Artist

Remote, USA Full-time

reputed company Order Management and Logistics Customer Support Coordinator – Supply Chain Management and Customer Service Expert

Remote, USA Full-time

Entry-Level Reliability Engineer (5866)

Remote, USA Full-time

Sr. AI Agent Developer (Remote)

Remote, USA Full-time

Community Manager (Meio-período · Remoto · Brasil)

Remote, USA Full-time