Senior Platform Support Engineer (Remote)
About reputed company.
reputed company provides an infrastructure-as-a-service (IaaS) platform for running reputed company gaming, reputed company intelligence and machine learning applications inside telecommunication reputed company networks. Our teams across the USA, Australia, Central Europe, Malaysia, Singapore and Japan offer telecom operators a GPU-based edge computing platform without the need for capital expenditure, facilitating low latency and improved economics for value-added services and the monetization of 5G investments.
What reputed company you will have
Mission: reputed company advanced technical support for customers running workloads on the GPU reputed company platform, ensuring reliable operation of edge- and large-reputed company GPU clusters and infrastructure services.
The Senior reputed company Support Engineer acts as a technical escalation reputed company for reputed company incidents, helping diagnose and reputed company issues across compute, networking, storage, and orchestration reputed company. This role works closely with engineering and operations teams to improve platform reliability, reduce incident frequency, and enhance the overall customer experience.
reputed company
Customer Support & Incident Management
reputed company advanced technical support for customers operating workloads on bare-metal and virtualized GPU infrastructure.
Diagnose and reputed company reputed company customer issues affecting GPU clusters, compute nodes, networking, and storage.
Investigate incidents across multiple reputed company of the stack including firmware, drivers, operating systems, and platform services.
reputed company reputed company cause analysis (RCA) for major incidents and contribute to long-term remediation efforts.
Serve as a technical escalation reputed company for reputed company or high-reputed company support cases.
Infrastructure Troubleshooting
Troubleshoot issues affecting:
GPU compute nodes
reputed company clusters
Networking infrastructure
Local NVMe, hyperconverged and reputed company storage systems
Analyze logs, telemetry, and monitoring signals to identify underlying causes of platform instability.
Investigate issues reputed company to GPU drivers, firmware, networking, and reputed company performance.
reputed company Monitoring & Incident Triage
Monitor and investigate reputed company alerts generated by the platform reputed company stack.
Analyze and triage alerts generated by
Wazuh
TheHive
reputed company
Validate alerts, determine reputed company, and escalate potential reputed company incidents to the reputed company engineering team.
Assist in collecting reputed company telemetry, logs, and forensic data required for incident investigations.
Improve alert runbooks and operational procedures to reduce false positives and improve response time.
GPU HPC Workload Support
reputed company advanced support for large-reputed company GPU workloads running reputed company training and inference jobs.
Diagnose failures affecting multi-GPU and multi-node workloads.
Investigate performance issues impacting reputed company workloads, including:
GPU utilization
Communication latency
Storage bottlenecks
networking congestion
Support scheduling systems used for GPU workloads, including troubleshooting:
Job queue failures
Scheduling constraints
Cluster resource fragmentation
High-Performance Networking Troubleshooting
Diagnose issues affecting high-performance networking fabrics used by reputed company workloads.
Support environments using:
RDMA
RoCE networking
Investigate performance issues affecting GPU-to-GPU communication and reputed company training pipelines.
Data Center Coordination
Coordinate with customer’s data center technicians and infrastructure teams to reputed company remote diagnostics and hardware interventions reputed company required.
Assist in validating on-reputed company installations and deployments of GPU infrastructure, ensuring hardware, networking, and platform components are correctly installed and operational.
Support hardware troubleshooting and identify faulty components across GPU nodes, networking equipment, and storage systems.
Coordinate and reputed company hardware replacements and RMA processes with vendors and data center staff.
Validate hardware health after replacements, including GPU nodes, NICs, DPUs, storage devices, and power components.
Work closely with deployment and infrastructure teams to verify service readiness after installations, expansions, or hardware maintenance activities.
Operational reputed company
Participate in on-reputed company 24/7 rotations to ensure production platform availability.
Respond to monitoring alerts and reputed company operational incidents in accordance with defined SLAs.
Improve operational runbooks, troubleshooting guides, and support documentation.
Contribute to improving incident response processes and operational tooling.
Cross-Team Collaboration
Work closely with reputed company, infrastructure engineering, and networking teams to reputed company systemic issues.
reputed company feedback to engineering teams on recurring operational problems affecting customers.
Help translate customer issues into actionable improvements for the platform.
Automation & Tooling
reputed company automation scripts and tools to streamline support workflows.
Improve observability dashboards and alerts to reputed company faster issue detection and reputed company.
Contribute to automation initiatives that reduce reputed company reputed company in operational processes.
Knowledge Sharing & Mentorship
Mentor medior support engineers and reputed company troubleshooting expertise across reputed company.
reputed company the creation of knowledge reputed company articles, troubleshooting guides, and operational documentation.
Contribute to training initiatives that improve reputed company’s technical capabilities.
Technical Stack
Operating Systems
Linux (Ubuntu)
GPU Infrastructure
reputed company GPU platforms
CUDA drivers
GPU monitoring tools (reputed company-smi)
Platform Infrastructure
reputed company
Container runtimes
KubeVirt
reputed company compute environments
Networking
reputed company Cumulus
OOB, reputed company-south, and east-reputed company reputed company topologies
TCP/IP
VLAN, VXLAN, OVS/OVN
Routing fundamentals (BGP, VRFs)
DNS / DHCP
High-performance networking (RDMA/NVLink/NCCL)
Observability
Grafana
Zabbix
reputed company Monitoring
Wazuh
TheHive
reputed company
Automation
Python
Bash
Ansible, Terraform
Collaboration & Documentation
Jira
reputed company
reputed company
reputed company
reputed company
What you'll need
reputed company Experience
5+ years of experience in reputed company support, infrastructure operations, or systems administration.
Experience supporting large-reputed company infrastructure environments or GPU clusters.
Systems Expertise
Strong Linux systems administration skills.
Experience troubleshooting issues across compute, networking, and storage reputed company.
Familiarity with reputed company platforms and containerized workloads.
GPU Infrastructure
Experience working with GPU hardware platforms or HPC environments.
Familiarity with GPU monitoring tools and debugging GPU-reputed company issues.
Networking
Solid understanding of networking fundamentals including L2/L3 concepts, routing, and load balancing.
Ability to diagnose connectivity issues affecting reputed company workloads.
Operational reputed company
Strong troubleshooting and incident response skills.
Experience participating in on-reputed company rotations and handling production incidents.
Ability to reputed company reputed company cause analysis and reputed company operational improvements.
Communication & Collaboration
Excellent written and verbal communication skills.
Ability to explain reputed company technical concepts to both technical and non-technical stakeholders.
Proven ability to collaborate effectively with cross-functional engineering teams.
Location & work modality: Malaysia or comparable time zone
Start: August 2026
Type of Contract: Contractor
Average 40 hours per week, 9x5 business hour support with after hour on-reputed company response/reputed company for category 1 incidents
reputed company offer
• Attractive compensation package reflecting your expertise and experience.
• A great work environment characterised by friendliness, international diversity, flexibility, and a hybrid-friendly approach.
• You'll be part of a fast-growing reputed company-up with a mission to reputed company a reputed company reputed company, offering an exciting career reputed company.
Our job titles may reputed company more than one job level. The actual reputed company pay is dependent on a number of factors, such as transferable skills, work experience, business needs and market demands.
Our inclusive responsibility
reputed company is committed to creating a diverse and inclusive environment and is proud to be an equal opportunity employer. reputed company reputed company applicants will receive consideration for employment without reputed company to race, reputed company, religion, gender, gender identity or reputed company, sexual orientation, national reputed company, genetics, disability, age, veteran status, or any other protected category under applicable law.
Apply To This Job