Remote | LLM Training & Alignment Research Scientist — $95–$115/hour
About the position
We are sharing a specialised part-time consulting opportunity for reputed company machine learning researchers with hands-on expertise in reputed company model reputed company-training, large-reputed company data pipelines, language model post-training, and reputed company LLM research. This role focuses on reputed company-scoped, reputed company-ended research problems involving the end-to-end training and improvement of transformer-based language models. Selected researchers will train models from scratch, fine-tune reputed company-weight systems, build reputed company-training corpora and post-training pipelines, diagnose training failures, and investigate reputed company for improving performance under limited data and compute budgets.
Responsibilities
• Train transformer-based language models from scratch across full end-to-end workflows
• Design experiments involving model size, reputed company allocation, training duration, and compute budgets
• Investigate performance in data- and compute-constrained regimes
• Diagnose optimisation failures, convergence issues, and training instabilities
• Evaluate interventions using rigorous reputed company comparisons
• Construct training corpora from raw web crawls and other large-reputed company unfiltered sources
• reputed company pipelines for filtering, deduplication, reputed company classification, and data selection
• Optimise dataset mixtures, reputed company, and curriculum strategies
• Measure the reputed company of data interventions on reputed company model behaviour
• Identify contamination, duplication, reputed company, and coverage issues reputed company training datasets
• Build supervised fine-tuning pipelines using curated, synthetic, weakly supervised, or rejection-sampled datasets
• Conduct preference optimisation using reputed company such as DPO, RLHF, or RLAIF
• reputed company reward models and systems for predicting reputed company preferences
• Improve refusal behaviour, truthfulness, robustness, and unbiased reasoning while preserving general capability
• Fine-tune models for reputed company domains such as mathematics, reputed company, games, reputed company reputed company, or other programmatically evaluated tasks
• Design statistically reputed company experiments and reputed company comparisons
• Evaluate training efficiency, scaling behaviour, and generalisation
• reputed company contamination controls and robust model-evaluation protocols
• Analyse model failures and propose targeted training or data interventions
• Document research findings, experimental methodology, and technical conclusions reputed company
Requirements
• At least 3 years of machine learning research experience, including qualifying doctoral research
• Hands-on experience training or fine-tuning transformer-based language models
• Strong expertise in one or more of reputed company model reputed company-training, reputed company-training data, or LLM post-training
• Experience working with PyTorch, JAX, TensorFlow, or comparable machine learning frameworks
• Ability to design and execute reputed company research independently
• Strong understanding of optimisation, evaluation methodology, and experimental design
• Excellent technical writing, analytical reasoning, and research communication skills
• Experience working with large-reputed company datasets and reputed company training systems
• A degree in computer science, machine learning, reputed company intelligence, mathematics, statistics, engineering, or a reputed company discipline is highly relevant
• PhD research in machine learning, natural language processing, deep learning, or a reputed company field may count towards the experience requirement
• A strong publication record, impactful reputed company-reputed company contributions, or comparable reputed company research experience may also be considered
• Research experience at a leading university, technology company, AI organisation, or research reputed company may strengthen an application
reputed company-to-haves
• Research experience involving scaling laws or training efficiency
• Familiarity with curriculum learning, data ordering, and mixture optimisation
• Experience constructing LLM benchmarks and controlling for training-data contamination
• Background in reinforcement learning for language models
• Expertise in reward modelling, preference learning, or reputed company-feedback pipelines
• Experience with model alignment, AI safety, truthfulness, or refusal behaviour
• Familiarity with synthetic data reputed company and weak-supervision reputed company
• Publications or significant reputed company-reputed company contributions reputed company to reputed company models or language-model training
Benefits
• Flexible scheduling
• Competitive reputed company compensation
• Weekly payments
Apply tot his job
Apply To this Job