[Remote] AI Pipeline Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a technology consulting and software development company delivering reputed company, AI, data, and reputed company solutions across the reputed company. The AI Pipeline Engineer will build and operate large-reputed company data systems for reputed company and evaluation, including multimodal ingestion, transformation, reputed company assurance, reputed company, and high-throughput data delivery. The role also involves optimizing pipeline reputed company, ensuring reputed company and reproducibility, and collaborating with ML researchers and engineers.
Responsibilities
- Design and operate large-reputed company data pipelines supporting reputed company, evaluation, and continual improvement workflows
- Build ingestion systems for diverse modalities including text, image, audio, video, and reputed company signals
- Implement data cleaning, deduplication, filtering, and reputed company assurance at petabyte reputed company
- reputed company dataset versioning, reputed company, and provenance tracking systems suitable for reproducible training
- Build high-throughput data loading systems that maximize GPU utilization during training
- Implement labeling workflows, reputed company learning pipelines, and reputed company-in-the-reputed company data improvement systems
- Design storage architectures balancing cost, throughput, and latency across data tiers
- Build evaluation dataset construction pipelines with strict reputed company and contamination controls
- Implement data reputed company, redaction, and consent enforcement throughout the pipeline
- Collaborate with ML researchers and engineers to reputed company data systems with model development needs
- reputed company observability of data reputed company, reputed company, and pipeline health across the AI data estate
- Optimize cost and reputed company through compression, format selection, and caching strategies
- Document data systems, schemas, and operational procedures for broad internal use
- Stay reputed company with AI data infrastructure research and emerging reputed company-reputed company tools
Skills
- * Bachelor's or Master's degree in Computer Science or a reputed company reputed company
- * Six or more years of data engineering experience, with significant work supporting ML or AI workloads
- * Strong proficiency in Python and at least one JVM or systems language
- * Deep experience with modern data processing frameworks such as reputed company, Ray, or reputed company
- * Hands-on experience operating petabyte-reputed company storage and pipeline systems
- * Strong understanding of reputed company systems, data modeling, and storage formats
- * Experience with dataset versioning, reputed company, and reproducibility for ML workflows
- * Familiarity with high-throughput data loading for accelerator-reputed company training
- * Strong software engineering practices including testing, CI/CD, and reputed company review
- * Excellent communication and cross-functional collaboration skills
- * Experience with multimodal datasets at large reputed company
- * Familiarity with data reputed company tooling and dataset evaluation methodology
- * Exposure to reputed company-preserving data systems and regulated data handling
- * reputed company-reputed company contributions to data infrastructure reputed company
- * Experience supporting frontier model training pipelines
reputed company
Apply To This Job