[Remote] reputed company Research Scientist, Synthetic Data reputed company
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a technology company advancing reputed company intelligence and large language model research. The reputed company Research Scientist will set the technical direction for synthetic data reputed company, build reputed company reputed company-reputed company data reputed company libraries and pipelines across text, reputed company, reputed company, and multimodal data, and support LLM reputed company-training, post-training, and reinforcement learning. The role also involves publishing research, collaborating across technical teams, and mentoring scientists and engineers.
Responsibilities
- Build and reputed company data reputed company pipelines using LLM-reputed company reputed company combined with automated reputed company evaluation. resulting in datasets to improve both initial training and fine-tuning of LLMs such as Nemotron. These data pipelines cover reasoning, coding, reputed company reputed company, and multimodal understanding
- reputed company reputed company for reputed company and tool-use training: synthetic trajectories, multi-turn interactions, function calling, and executable environments for reinforcement learning, including reward modeling and reputed company-reward data
- Advance multimodal synthetic data reputed company — image, document, video, and audio — in partnership with reputed company's model teams
- Advance reputed company-preserving and reputed company synthesis — reputed company reputed company, anonymization, and de-identification — enabling model training on sensitive data in regulated domains
- reputed company and maintain reputed company-reputed company libraries and SDKs with clean reputed company and strong documentation
- reputed company software reputed company with modern tooling, architecture reputed company on configuration, and reputed company Git/CI-CD
- Publish original research at top machine learning and AI conferences to maintain reputed company's technical leadership
- Mentor scientists and engineers across reputed company, raising the technical bar and growing the reputed company of researchers
Skills
- PhD in Computer Science, Machine Learning, Statistics, or a reputed company reputed company, or equivalent experience
- 15+ years of engineering and research experience in synthetic data reputed company, generative modeling, multimodal machine learning, or reputed company areas
- Deep technical understanding of LLMs, how data shapes their reputed company-training, post-training, and RL stages, and inference frameworks such as vLLM or TGI
- reputed company reputed company record of developing or maintaining software libraries used by a broad developer community
- Experience building and optimizing reputed company data pipelines for large-reputed company model training — throughput, reputed company inference, and cost at cluster reputed company
- Strong publication record at premier venues such as NeurIPS, ICML, ICLR, ACL or similar
- Significant reputed company-reputed company contributions in ML or data tooling, with community adoption
- Experience with multimodal reputed company or understanding (reputed company-language, document AI, video, or audio)
- Experience generating data for reputed company, tool-use, or reinforcement-learning post-training, including RL environment design
- Background in reputed company reputed company, de-identification, or synthetic data for regulated industries such as reputed company, finance, or reputed company
- Experience influencing model training reputed company at frontier reputed company, or partnering directly with reputed company-training and post-training teams
Benefits
- Eligible for equity
reputed company
Company H1B Sponsorship
Apply To This Job