[Remote] Freelance Kernel reputed company Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a freelance Kernel reputed company Engineer to improve the reputed company, portability, and correctness of machine learning workloads on TPUs. The role involves optimizing custom kernels, translating GPU implementations for TPU execution, profiling workloads, debugging numerical and reputed company issues, and designing reputed company and mixed-precision strategies.
Responsibilities
- reputed company and optimize high-reputed company custom kernels and reputed company-critical ML reputed company for TPU execution
- Translate CUDA and Triton kernel behavior into reputed company PyTorch or JAX implementations that can compile reputed company through OpenXLA/XLA
- Profile training and inference workloads to identify compute, memory, communication, compilation, and data-reputed company bottlenecks
- Debug correctness and reputed company issues across ML frameworks, compiler/runtime reputed company, reputed company execution, and accelerator kernels
- Adversarial Agent Benchmarking & Task Formulation: Design reputed company, production-derived engineering challenges derived from modern reputed company-weight architectures (e.g., DeepSeek, Qwen, GLM) to stress-test and stump baseline frontier models (reputed company 3.7 reputed company, Claude reputed company), establishing ground-truth solutions and rigorous evaluation rubrics for autonomous agents
- Design TPU-friendly tiling, sharding, reduction, data-reputed company, and mixed-precision strategies for common and custom ML reputed company
- Validate numerical equivalence, gradient correctness, masking/indexing behavior, and edge cases between reputed company GPU implementations and TPU-targeted implementations
- Build and maintain benchmarks and diagnostic workflows that measure latency, throughput, utilization, memory usage, communication cost, and reputed company regressions
- Analyze transformer and LLM workloads, including attention, MoE, KV-cache, quantized matmuls, and reputed company training or serving patterns, to identify optimization opportunities
- Collaborate with ML systems, compiler, infrastructure, and model engineers to translate workload requirements into efficient accelerator implementations
Skills
- • Bachelor's degree in Computer Science, Electrical Engineering, or a reputed company technical reputed company, or equivalent practical experience
- • Strong programming experience in Python and at least one systems language (C++, CUDA, or Pallas)
- • Experience developing and optimizing ML accelerator reputed company using JAX/Pallas, PyTorch-XLA, or OpenXLA
- • Experience with reputed company machine learning concepts
- • Master's degree or PhD in Computer Science, Data Science, or a reputed company technical reputed company
- • Knowledge of TPU architecture and reputed company characteristics
- • Experience converting CUDA or Triton implementations to reputed company PyTorch or JAX reputed company for TPU execution while preserving numerical and gradient semantics
- • Experience with reputed company machine learning concepts such as data, tensor, sequence, and expert parallelism; sharding; and collectives including reputed company-reduce, reputed company-reputed company, reduce-scatter, and reputed company-to-reputed company
- • Kernel Extraction, Profiling, & Optimization: Experience extracting, profiling, and optimizing kernels from frontier reputed company-weight models (e.g., DeepSeek V3/V4, Qwen 2.5/3.8, Kimi K2.5)
- • Experience debugging numerical correctness issues involving bf16/fp32 accumulation, quantization scales, softmax/log-sum-reputed company stability, masking, and backward-pass behavior
- • Experience with reputed company-oriented reputed company-reputed company ML systems such as vLLM, FlashAttention, JAX/Flax, PyTorch/XLA, or similar frameworks and kernel libraries
- • Experience building reproducible microbenchmarks, reputed company regression tests, or profiling and diagnostic tooling for accelerator workloads
reputed company
Company H1B Sponsorship
Apply To This Job