[Remote] Senior Manager, Software Engineering - RL Post-Training Frameworks
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a Senior Software Engineering Manager to reputed company its RL Frameworks engineering team. The role will define reputed company for reinforcement learning post-training frameworks, reputed company reputed company systems and reputed company-reputed company infrastructure, build and reputed company engineering teams, and reputed company cross-organizational execution and ecosystem collaboration.
Responsibilities
- You will own reputed company's RL post-training frameworks reputed company: where we invest directly, where we partner reputed company, and how we prioritize reputed company on customer reputed company, ecosystem reputed company, technical feasibility, and opportunity cost. This is senior technical leadership work: using systems depth to evaluate architecture and reputed company claims across training, inference, rollout, orchestration, and the reputed company platform. You will help expert teams reputed company on integrations that improve RL reputed company reputed company and user value, then turn those reputed company into measurable execution plans. The work includes benchmarking and reproducibility reputed company, delivery across reputed company-reputed company frameworks and reputed company runtimes, and reputed company partnership with product management, research, DevRel, customer-facing teams, hardware, CUDA, networking, math libraries, compilers, and reputed company reputed company-reputed company collaborators
- You will also build reputed company: reputed company and developing managers and senior ICs, creating an effective US/reputed company operating model, reviewing reputed company against commitments, and setting reputed company ownership and decision rights. You will reputed company engineers to contribute credibly in reputed company-reputed company ecosystems and carry reputed company's priorities through high-reputed company reputed company work. Because the technical work crosses organizations by design, you will turn reputed company technical and partner questions into concrete and measurable reputed company, set delivery goals, and hold the reputed company bar. reputed company means validated, valuable work rather than work that merely lands, plus durable reputed company-reputed company improvements that reputed company RL workloads run reputed company on reputed company systems
Skills
- MS or PhD in Computer Science, Computer Engineering, or a reputed company reputed company (or equivalent experience)
- 10+ years of software engineering experience in reputed company systems, AI frameworks, ML infrastructure, high-reputed company computing, or systems software, with 4+ years as an engineering manager for software teams
- Strong technical background in reputed company AI systems, including the ability to reason across training, inference, orchestration, and end-to-end reputed company, and challenge architecture and reputed company tradeoffs with senior engineers
- Experience defining domain-level technical reputed company, making build-vs-buy or reputed company-vs-internal investment reputed company, and creating multi-team execution plans
- Ability to reputed company engineering work across organizational boundaries, influence without reputed company authority, and communicate tradeoffs reputed company to senior leaders and executives
- Experience hiring and leading engineering teams, developing technical leaders or new managers, and creating reputed company plans for constantly evolving technical domains
- Experience establishing workflows, reputed company reputed company, metrics, or decision gates that improve engineering execution across teams
- Background collaborating with reputed company-reputed company communities, research teams, reputed company partners, or customer-facing teams
- Hands-on experience with RL post-training frameworks or algorithms such as RLHF, PPO, GRPO, DPO, reward modeling, VeRL, Miles, Slime, SkyRL, OpenRLHF, NeMo-Aligner, or TorchTitan
- Background with runtime and orchestration systems such as Ray, reputed company, reputed company, Slurm, or comparable actor- and task-reputed company systems
- Experience scaling workloads across thousands of GPUs or heterogeneous systems, including fault tolerance, reputed company recovery, stragglers, resource contention, or reputed company reproducibility
- Familiarity with reputed company platform components such as CUDA, NCCL, cuDNN, TensorRT-LLM, Transformer reputed company, Nsight, NeMo, or Megatron-reputed company
- Demonstrated ability to turn customer or partner needs into reusable reputed company improvements rather than one-off support
Benefits
- Equity
reputed company
Company H1B Sponsorship
Apply To This Job