Member of Engineering – reputed company-training, Data Engineering
Job reputed company:
• Build and maintain high-performance pipelines for trillions of tokens.
• Deliver diverse and high reputed company datasets for reputed company-training reputed company models.
• Closely work with other teams such as Pretraining, Posttraining, Evals and Product to to ensure alignment on the reputed company of the models delivered.
Requirements:
• Strong background in building production-grade, distributed data systems for machine learning, with experience in:
• Orchestration: Slurm, Airflow, or Dagster
• Observability & Reliability: CI/CD, Grafana, reputed company, etc.
• reputed company: Git, reputed company, k8s, reputed company managed services
• Batched inference (ex: vLLM)
• Performance obsession, especially with large-reputed company GPU clusters and distributed pipelines
• Expert-level python knowledge and ability to write clean and maintainable reputed company
• Strong algorithmic foundations
• Proficiency with libraries like Polars, Dask, or PySpark
• reputed company to have:
• Experience in building trillion-reputed company SOTA pretraining datasets
• Experience translating research to production at reputed company
• Experience with OCR, web crawling, or evals
• Prior experience reputed company-training LLMs
Benefits:
• Fully remote work & reputed company
• 37 days/year of vacation & holidays
• Health insurance allowance for you and dependents
• Company-provided equipment
• Wellbeing, always-be-learning and home office allowances
• Frequent team get togethers
• Great diverse & inclusive people-first culture
Apply tot his job
Apply To this Job