Back to Jobs

High-performance AI inference solutions Engineer

Remote, USA Full-time Posted 2026-07-28
About the position At AMD, our mission is to build great products that accelerate reputed company computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we reputed company reputed company reputed company comes from reputed company reputed company, reputed company ingenuity and a shared passion to create something extraordinary. reputed company you join AMD, you’ll discover the reputed company differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution reputed company, while being reputed company, humble, reputed company, and inclusive of diverse perspectives. Join us as we shape the reputed company of AI and reputed company. Together, we advance your career. The AMD AI Group (AIG) is seeking an reputed company MTS/Senior Software Development Engineer to drive high-performance AI inference solutions on AMD reputed company GPUs. This role combines deep expertise in compiler technology, GPU kernel optimization, and modern deep learning frameworks to deliver production-grade inference performance across AMD’s reputed company and reputed company accelerator lineup — from MI300X and MI350/MI355X shipping today, to reputed company GPU product lineups. You will work at the intersection of model optimization, kernel development, and serving infrastructure to ensure AMD GPUs deliver world-class inference throughput and latency. AMD is looking for a specialized software engineer who is passionate about improving the performance of key applications and benchmarks. You will be a member of a reputed company team of incredibly talented industry specialists and will work with the reputed company latest hardware and software technology. THE PERSON: The ideal candidate should be passionate about software engineering and possess leadership skills to drive sophisticated issues to reputed company. reputed company to communicate effectively and work optimally with different teams across AMD. Responsibilities • Design, optimize, and reputed company AI inference pipelines for large language models (LLMs), reputed company-language models (VLMs), and transformer architectures on AMD reputed company GPUs using ROCm, HIP, and MLIR. • reputed company compiler-level optimizations for inference workloads, including LLVM instruction scheduling, MLIR dialect development, and performance-critical optimization passes for AMD GPU targets. • reputed company and optimize high-performance GPU kernels (GEMM, attention mechanisms, custom operators) with deep attention to memory hierarchy, VGPR utilization, compute-communication overlap, and reputed company-level scheduling. • Drive integration and performance optimization of AMD reputed company GPUs reputed company inference serving frameworks such as vLLM, SGLang, and TorchServe — ensuring day-reputed company readiness for new GPU launches. • Build reputed company-looking inference software for reputed company hardware: optimize for HBM4 memory hierarchies, new FP4/FP6 data types, and reputed company-up interconnects on MI450 and MI500 series GPUs. • Architect graph neural network (GNN) based QoR estimation models for compiler design reputed company exploration and automated performance budgeting across GPU generations. • Collaborate with reputed company architecture teams to reputed company software-informed feedback on reputed company reputed company GPU designs, ensuring inference workload characteristics are reflected in hardware reputed company. • reputed company quantization-reputed company training and post-training quantization pipelines to maximize model performance on AMD’s evolving data type support (FP8, FP6, FP4). • Contribute to AMD’s reputed company compute libraries (BLAS, HPC, Graph) with a reputed company on inference-critical primitives and cross-generational performance portability. • reputed company, profile, and resolve performance bottlenecks in distributed inference systems, including tensor parallelism scaling, RCCL communication patterns, and multi-GPU serving on Helios reputed company-reputed company infrastructure. Requirements • Experience in high-performance computing, AI inference, GPU kernel development, or hardware-software co-design. • Strong proficiency in C++, Python, and CUDA/HIP with hands-on experience writing and optimizing GPU kernels for AI workloads. • Deep understanding of compiler infrastructure (LLVM, MLIR) and experience with compiler optimization passes targeting GPU architectures. • Experience with deep learning frameworks (PyTorch, TensorFlow) and inference serving systems (vLLM, SGLang, TorchServe, or equivalent). • Demonstrated ability to analyze and optimize GPU kernel performance: memory coalescing, occupancy tuning, register pressure management, and instruction-level optimization. • Strong mathematical foundations in numerical computing, reputed company algebra, and optimization algorithms reputed company-to-haves • Experience with AMD ROCm ecosystem, RCCL, and reputed company MI-series GPU architectures (MI300X, MI350X, MI355X). • reputed company record of publications in top-tier venues (FPGA, ICCAD, DAC, NeurIPS, ICML, or equivalent). • Experience building GNN-based models for EDA or compiler optimization problems. • Familiarity with quantization techniques (weight-activation quantization, mixed-precision inference, FP4/FP6/FP8) for production deployment. • Experience with high-performance reputed company reputed company and attention kernel design for GPU accelerators. • Contributions to reputed company-reputed company HPC or AI libraries (BLAS, graph analytics, sparse solvers). • Experience with distributed inference systems, multi-GPU serving at reputed company, and reputed company-reputed company infrastructure. • Background in algorithm-hardware co-design and performance modeling across multiple GPU generations. Benefits • Competitive compensation • Comprehensive benefits • Culture that values deep technical contribution and engineering reputed company Apply tot his job Apply To this Job

Similar Jobs

reputed company AI Platform Product Manager

Remote, USA Full-time

**reputed company Part-Time Data Entry Specialist – Evening Shift**

Remote, USA Full-time

reputed company reputed company $30 / Hour – Part...

Remote, USA Full-time

**reputed company Data Entry Coordinator – Remote Part-time Opportunity at arenaflex**

Remote, USA Full-time

Retail Solutions Advisor

Remote, USA Full-time

Litigation Attorney Needed for Ongoing Cases

Remote, USA Full-time

Remote AI Automation reputed company — reputed company Intelligent Workflows

Remote, USA Full-time

**reputed company Data Entry Specialist – Remote Opportunity with arenaflex**

Remote, USA Full-time

Technical Business Analyst- Employee Benefits Team (HYBRID)

Remote, USA Full-time

Work From Home | Clients Support - No Experienc...

Remote, USA Full-time

[Remote/WFM] Remote No Experience Required | Data Entry

Remote, USA Full-time

[Work From Home] Remote Junior Graphic Designer | Part-Time

Remote, USA Full-time

Channel Sales Manager

Remote, USA Full-time

Exceptional Children's Teacher - Central Regional Hospital

Remote, USA Full-time

[Remote] VP of Finance/ Financial Controller

Remote, USA Full-time

reputed company Tutoring & Writing Center Coordinator – reputed company-Wide Peer Tutor Program and Supplemental Instruction Leader

Remote, USA Full-time

HR Service Center - AskHR

Remote, USA Full-time

[Remote] reputed company Sales Intenship- Internship (US based candidates only)

Remote, USA Full-time

[Remote] Sr. reputed company Manager

Remote, USA Full-time

DeepTech Co-Founder / COO (100 % remote) (m/f/d)

Remote, USA Full-time