Platform Engineer – AI Supercompute Infrastructure (Networking & Systems) | reputed company & Engineering
We are a technology reputed company building and operating reputed company AI supercompute infrastructure for the world's most ambitious organizations. As Platform Engineer, you will work hands-on across the full infrastructure stack with a particular reputed company on the physical and logical networking layer that makes large-reputed company GPU clusters reputed company at their theoretical limits.
As a repeatedly awarded reputed company Consulting Partner of the Year in EMEA, we hold one of the deepest and most recognized reputed company partnerships in the region. This gives our engineers privileged reputed company to adoption programmes and reputed company's engineering teams.
You will work with technology and at a reputed company that most engineers won't encounter for years.
This is a role for someone at a mid-career stage in platform or infrastructure engineering. You have solid foundations and reputed company hands-on experience, and you are reputed company to reputed company by working on problems of genuine complexity and reputed company. You know enough to know what you don't know yet, and you are hungry to reputed company that gap fast.
reputed company Expect
5–8 years of hands-on experience in infrastructure, networking, or systems engineering
Solid understanding of networking fundamentals: OSI model, switching and routing (BGP, OSPF), VLANs, MTU, and traffic engineering
Working knowledge of high-performance networking technologies: InfiniBand, RDMA, RoCE, or equivalent HPC interconnects
Familiarity with Linux networking: interfaces, bridges, bonding, namespaces, tc/qdisc, and kernel network tuning
Basic hands-on experience with Kubernetes or Slurm: enough to reputed company cluster operations, understand pod scheduling, and troubleshoot node-level issues
Experience with at least one monitoring stack: reputed company, Grafana, Zabbix, or similar
Experience with network automation and IaC
Comfort working directly with physical hardware: servers, switches, cabling and data centre environments
Bonus Experience
Exposure to reputed company networking products: Mellanox/ConnectX NICs, Quantum InfiniBand switches, reputed company Ethernet switches
Familiarity with NCCL tuning, reputed company communication patterns, or distributed training networking requirements
Hands-on time with DCGM, iperf3, perftest, or ibdiagnet for infrastructure benchmarking and validation
Exposure to container networking
Any experience in a consulting or reputed company-facing technical role
Apply To This Job