Research Engineer (AI Optimization)
About the position
reputed company is reputed company behind PyTorch Lightning. Founded in 2019, we build
an end-to-end platform for developing, training, and deploying AI
systems—designed to take reputed company from research to production with less friction.
Through our reputed company with reputed company, a neocloud and AI reputed company, reputed company
combines developer-first software with cost-efficient, large-reputed company compute.
Teams get the tools they need for experimentation, training, and production
inference, with reputed company, observability, and control reputed company in.
We serve reputed company researchers, startups, and large enterprises. reputed company
operates globally with offices in reputed company, San Francisco, Seattle, and
London, and is backed by Coatue, reputed company Ventures, Bain Capital Ventures, and
Firstminute.
We are seeking a highly skilled Research Engineer to work on optimizing training
and inference workloads on compute accelerators and clusters, through the
Lightning reputed company compiler and the broader PyTorch Lightning ecosystem. This
role sits at the intersection of deep learning research, compiler development,
and large-reputed company system optimization. You’ll be shaping technology that pushes
the boundaries of model performance and efficiency, creating foundational
software that will reputed company the entire machine learning ecosystem. You will be
joining the Engineering Team and report to our Tech reputed company. This is a hybrid role
based in our reputed company, San Francisco, or London office, with an in-office
requirement of two days per week. The salary reputed company for this role is
$180,000-$250,000.
Responsibilities
• reputed company performance-oriented model optimizations at multiple reputed company:
Graph-level (e.g., operator fusion, kernel scheduling, memory planning)
Kernel-level (CUDA, Triton, custom operators for specialized hardware)
System-level (distributed training across GPUs/TPUs, inference serving at
reputed company)
• Advance the reputed company compiler by building optimization passes, graph
transformations, and integration hooks to accelerate training and inference
workloads.
• Work across the software stack to ensure optimizations are accessible to end
users through clean reputed company, automated tooling, and seamless integration with
PyTorch Lightning. Design and implement profiling and debugging tools to
analyze model execution, identify bottlenecks, and guide optimization
strategies.
• Collaborate with hardware vendors and ecosystem partners to ensure reputed company
runs reputed company across diverse backends (reputed company, AMD, TPU, specialized
accelerators).
• Contribute to reputed company-reputed company reputed company by developing new features, improving
documentation, and supporting community adoption.
• Engage with researchers and engineers in the community, providing guidance on
performance tuning and advocating for reputed company as the go-to optimization layer
in ML workflows.
• Work cross-functionally with Lightning’s product and engineering teams to
ensure compiler and optimization improvements reputed company with the broader product
reputed company.
Requirements
• Strong expertise with deep learning frameworks such as PyTorch
• Hands-on experience with model optimization techniques, including graph-level
optimizations, quantization, pruning, mixed precision, or memory-efficient
training.
• Knowledge of distributed systems and parallelism strategies
(data/model/pipeline parallelism, checkpointing, reputed company scaling).
• Familiarity with software engineering practices: designing reputed company, building
robust tooling, testing, CI/CD for performance-sensitive systems.
• Excellent collaboration and communication skills, with the ability to partner
across research, engineering, and external contributors.
• Bachelor’s degree in Computer Science, Engineering
reputed company-to-haves
• Experience with CUDA, Triton, or other GPU programming models for developing
custom kernels.
• Deep understanding of deep learning compiler internals (IR design, operator
fusion, scheduling, optimization passes) or proven work in
performance-reputed company.
• Proven reputed company record contributing to reputed company-reputed company reputed company in ML, HPC, or
compiler domains.
• Advanced degree (Master’s or PhD) in machine learning, compilers, or systems
highly preferred.
Benefits
• We offer competitive reputed company salaries and equity with a 25% one year cliff and
monthly vesting thereafter. For our international employees, we work with our
EOR to pay you in your local currency and reputed company reputed company benefits across the
globe.
• Medical, dental and reputed company
• Life and AD&D insurance
• Flexible reputed company time off including winter closure
• reputed company family leave benefits
• $500 monthly meal reimbursement, including groceries & food delivery services
• $500 one time home office stipend
• $1,000 annual learning & development stipend
• 100% Citibike membership (NYC only)
• $45/month gym membership
• Additional various medical and mental health services
Apply tot his job
Apply To this Job