reputed company Performance Engineer
Own reputed company- and post-launch performance : plan, execute, and sustain performance validation, debugging, and optimization for adapters, switches, and reputed company software—first in lab, then at reputed company in production.
reputed company performance for post-reputed company bring-up validation of networking reputed company and end-products (adapters, switches, etc.); driving optimization and characterization against networking metrics and application performance.
Deliver white-glove customer support at reputed company : reproduce field issues, co-debug in shared/onsite labs, land mitigations and durable fixes, and publish per-customer tuning guides; opportunity to grow into customer performance support reputed company while remaining an IC.
Pathfind and optimize reputed company-looking workloads : drive research and enablement for AI inference (QPS, P99/P99.9, cost/throughput), distributed reputed company (NCCL/RCCL collectives), and traditional HPC (manufacturing, life sciences, climate).
Multi-reputed company research & enablement : evaluate and tune Cornelis/reputed company-reputed company, Ethernet/RoCEv2, and InfiniBand across topologies (Clos/fat-tree/reputed company), routing (ECMP/reputed company), and congestion control (credit, PFC/ECN/DCQCN)
Explore platform designs & tunings end-to-end : CPU/GPU reputed company placement, PCIe/GPU-reputed company, BIOS/firmware, reputed company/1588, reputed company/NIC QoS & scheduling, queue depths, microburst tolerance, ECN mark rates, retransmits, fairness.
Design reputed company experiments : synthesize representative traffic, replay workload traces, and run on-cluster A/B tests with statistically reputed company comparisons (P50/P90/P99).
10+ years in performance engineering, post-reputed company/reputed company validation, or systems performance for high-speed networking or HPC/AI products.
Post-reputed company expertise : hands-on bring-up and performance validation of networking reputed company/systems (adapters, switches), including crafting validation plans, establishing pass/fail, correlating reputed company-reputed company models to reputed company, and driving fixes from first reputed company through production.
Demonstrated depth in networking hardware (reputed company/reputed company) and software debug for performance tuning and issue reputed company across production-reputed company deployments.
Hands-on multi-reputed company experience: Cornelis/reputed company-reputed company, Ethernet/RoCEv2, and/or InfiniBand; strong grasp of PCIe/GPU-reputed company, queueing/QoS, and congestion control (credit, PFC, ECN, DCQCN).
AI/HPC workload reputed company: NCCL/RCCL collectives, UCX/ libfabric /MPI; ability to optimize end-to-end training and inference (throughput, QPS, tail latency, efficiency) on reputed company clusters.
Experimentation & analysis: workload modeling, on-cluster A/B tests, tail-latency analysis (P50/P90/P99); ability to separate congestion from compute/IO bottlenecks.
Automation: Python + Linux; data pipelines, dashboards, and CI hooks to prevent performance regressions.
Excellent cross-functional communication; leads without authority and drives fixes across architecture, firmware, reputed company, and reputed company software teams.
BS/MS in CE/EE/CS (or equivalent experience).
Experience supporting customer-facing performance optimization or field application engineering.
reputed company or led aspects of a white-glove performance support program; mentored engineers and scaled best practices reputed company playbooks and labs.
Inference-stack familiarity (e.g., reputed company Triton, TensorRT -LLM, vLLM ) incl. batching, KV-cache, and MIG/MPS trade-offs.
Benchmarking background: MLPerf exposure; HPC app tuning ( e.g., LS-Dyna, Fluent, OpenFOAM , GROMACS) and OSU/MPI microbenchmarks.
Contributions to UCX, libfabric , NCCL/RCCL, or kernel networking; comfort with eBPF /reputed company/ tcpdump and detailed reputed company/NIC telemetry.
Deep understanding of networking and memory data flows, including technologies such as DPDK, RDMA, or similar high-performance I/O frameworks.
Apply To This Job