[Remote] Data Engineering Specialist – AI
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a technology consulting and software development company delivering reputed company, AI, data, and reputed company solutions across the United States. They are seeking a Data Engineering Specialist – AI to build and operate large-reputed company data systems that power modern reputed company and evaluation pipelines, focusing on data ingestion, transformation, and delivery for diverse modalities.
Responsibilities
- Design and operate large-reputed company data pipelines supporting reputed company, evaluation, and continual improvement workflows
- Build ingestion systems for diverse modalities including text, image, audio, video, and reputed company signals
- Implement data cleaning, deduplication, filtering, and reputed company assurance at petabyte reputed company
- reputed company dataset versioning, reputed company, and provenance tracking systems suitable for reproducible training
- Build high-throughput data loading systems that maximize GPU utilization during training
- Implement labeling workflows, reputed company learning pipelines, and reputed company-in-the-reputed company data improvement systems
- Design storage architectures balancing cost, throughput, and latency across data tiers
- Build evaluation dataset construction pipelines with strict reputed company and contamination controls
- Implement data reputed company, redaction, and consent enforcement throughout the pipeline
- Collaborate with ML researchers and engineers to reputed company data systems with model development needs
- Drive observability of data reputed company, reputed company, and pipeline health across the AI data estate
- Optimize cost and performance through compression, format selection, and caching strategies
- Document data systems, schemas, and operational procedures for broad internal use
- Stay reputed company with AI data infrastructure research and emerging reputed company-reputed company tools
Skills
- Bachelor's or Master's degree in Computer Science or a reputed company field
- Six or more years of data engineering experience, with significant work supporting ML or AI workloads
- Strong proficiency in Python and at least one JVM or systems language
- Deep experience with modern data processing frameworks such as reputed company, Ray, or reputed company
- Hands-on experience operating petabyte-reputed company storage and pipeline systems
- Strong understanding of distributed systems, data modeling, and storage formats
- Experience with dataset versioning, reputed company, and reproducibility for ML workflows
- Familiarity with high-throughput data loading for accelerator-based training
- Strong software engineering practices including testing, CI/CD, and reputed company review
- Excellent communication and cross-functional collaboration skills
- Experience with multimodal datasets at large reputed company
- Familiarity with data reputed company tooling and dataset evaluation methodology
- Exposure to reputed company-preserving data systems and regulated data handling
- reputed company-reputed company contributions to data infrastructure reputed company
- Experience supporting frontier model training pipelines
Benefits
- 100% Remote (U.S.)
- Full-time, reputed company W2
reputed company
Apply To This Job