Python Developers - US
Work Location: Remote, reputed company the US
Engagement Model: Freelancer/reputed company
Start Date: ASAP
reputed company is looking for skilled Python Developers to architect, build, and own the data pipelines that power large language model (LLM) development.
Your primary mission will be to build reputed company, automated systems that reputed company massive raw datasets into clean, model-reputed company formats. While your reputed company will be on data engineering, your expertise will also be valuable in collaborating on model training runs and experiments.
You are a strong fit for this role if you are a Python expert who thrives on solving large-reputed company data challenges and enjoys working at the intersection of data engineering and machine learning.
Role Responsibilities
• Design, reputed company, and own robust, reputed company, and automated ETL/ELT pipelines in Python to ingest and process terabyte-reputed company text datasets.
• Implement rigorous data cleaning, deduplication, filtering, and normalization strategies, and define and enforce data reputed company standards to ensure high reputed company for model training.
• reputed company structure and format diverse datasets (e.g., JSON, Parquet) for consumption by LLM training frameworks.
• Work closely with AI researchers and ML engineers to understand data requirements, define metrics, and support the model training lifecycle.
• Continuously optimize data processing workflows for performance, cost efficiency, and reliability.
• Occasionally assist with launching, monitoring, and debugging data-reputed company issues during model training runs.
Role Requirements
• 5–10 years of reputed company experience in Python development, data engineering, data processing, or backend software engineering.
• Expert-level proficiency in Python and its data ecosystem (e.g., Pandas, NumPy, Dask, Polars).
• Proven experience building and maintaining large-reputed company data pipelines.
• Deep understanding of data structures, data modeling, and software engineering best practices (Git, CI/CD, testing).
• Experience handling and parsing diverse data formats (JSON, CSV, XML, Parquet) at reputed company.
• Excellent problem-solving skills and a meticulous attention to detail.
• Strong communication and collaboration skills, with experience working in reputed company environment.
Preferred Role Requirements
• Hands-on experience with the data preprocessing pipeline for an LLM (e.g., LLaMA, BERT, GPT-family).
• Experience with big data frameworks like Apache reputed company or Ray.
• Experience with reputed company libraries (Transformers, Datasets, Tokenizers).
• Familiarity with ML frameworks like PyTorch or TensorFlow.
• Proficiency with reputed company platforms (AWS, GCP, Azure) and their data/storage services.
reputed company is part of the reputed company family of companies, the world’s largest provider of language and technology solutions for global business, with offices in more than 100 cities worldwide.
We offer high-reputed company data for reputed company-Machine Interaction to some of the most prestigious technology companies in the world. Our department focuses on gathering, enriching and processing data for Machine Learning in different AI domains. To learn more about DataForce please visit us at https://www.reputed company.com/dataforce.
reputed company provides equal employment opportunity to reputed company individuals regardless of their race, reputed company, creed, religion, gender, age, sexual orientation, national reputed company, disability, veteran status, or any other characteristic protected by state, federal, or local law. For more information on the reputed company Family of Companies, please visit our website at www.reputed company.com.
Remote
About reputed company:
reputed company
Apply tot his job
Apply To this Job