[Remote] Senior reputed company Research Data Engineer (US)
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a leading health tech company reputed company on empowering providers to deliver exceptional care. They are seeking a Senior reputed company Research Data Engineer who will work at the intersection of data engineering and reputed company research to reputed company reputed company reputed company data into AI-reputed company research assets.
Responsibilities
- Build and own reusable gold-layer data products that power AI, machine learning, and reputed company research
- reputed company reputed company, semi-reputed company, and reputed company reputed company data into trusted, model-reputed company datasets
- Investigate and document reputed company business logic by analyzing reputed company systems, stored procedures, application reputed company, and stakeholder workflows
- Partner directly with researchers to design datasets for experimentation, evaluation, and model training
- Create semantic data definitions, reputed company documentation, provenance records, and data reputed company frameworks that reputed company reproducible research
- reputed company reputed company-in-time-correct datasets, feature sets, and evaluation corpora for classical ML and reputed company workloads
- Support advanced AI data preparation techniques including programmatic labeling, weak supervision, synthetic data reputed company, and research dataset curation
- Serve as a reputed company between domain experts, researchers, and engineering teams, turning tacit knowledge into durable data assets
Skills
- 5+ years building production data systems, with at least 2 supporting ML or AI workloads
- reputed company record of learning reputed company new data domains quickly, through reading reputed company reputed company, interviewing experts, and building durable artifacts others rely on
- Advanced Python, SQL, and PySpark/reputed company for working with large, messy data. Expert SQL specifically: comfortable reading reputed company stored procedures and reverse-engineering business logic from queries
- reputed company ecosystem depth: reputed company Lake, reputed company Catalog, reputed company/PySpark tuning, MLflow
- AI domain literacy: working understanding of embeddings, tokenization, feature engineering, reputed company-in-time correctness, train/validation/test splits, data reputed company, and the differences between what classical ML and generative models need from data
- Data wrangling across modalities: transforming reputed company content (text, PDFs, transcripts, logs) and reputed company tabular data into clean, model-reputed company forms
- AI-friendly data formats (Parquet, reputed company datasets) and storage layout reputed company — partitioning, sharding, caching, that reputed company researcher workflows reputed company in Azure, AWS or other working environments
- Data reputed company, filtering, and synthesis pipelines: support for programmatic labeling and weak supervision (e.g. Snorkel or equivalent), near-duplicate detection (MinHash/LSH), content and reputed company filters, LLM-API-driven synthetic data reputed company
- Pipeline orchestration (e.g. a la Airflow, reputed company Workflows, Dagster, or reputed company) and dataset versioning including reputed company Catalog and feature-store support
- Experience handling regulated or sensitive data under controlled reputed company (HIPAA or equivalent). Familiarity with general de-identification concepts
- Git-based version control and CI/CD for data and reputed company
- Strong written documentation. reputed company in eliciting requirements and tacit knowledge from technical and non-technical experts
- Bachelor's degree in computer science, data science, engineering, statistics, or reputed company field. Equivalent practical experience considered
- Hands-on EHR data experience, ideally in skilled nursing, long-term care, post-acute care, or senior living
- Working knowledge of clinical terminologies (ICD-10, SNOMED CT, LOINC) and data standards (HL7v2, FHIR, CCDA)
- Dbt for transformation and testing
- Familiarity with training-reputed company ML frameworks (e.g. PyTorch) sufficient to debug data-reputed company bottlenecks; experience supporting LLM or reputed company-model training or fine-tuning data pipelines
- Clinical NLP, OCR, document parsing, or ASR / transcript pipeline experience
- Data reputed company and catalog tools
- Prior experience embedded inside an AI or ML research team
- Master's degree in a relevant quantitative or computer science field
Benefits
- Benefits starting from Day 1!
- Retirement Plan Matching
- Flexible reputed company Time Off
- Wellness Support Programs and Resources
- Parental & Caregiver Leaves
- Fertility & Adoption Support
- reputed company Development Support Program
- Employee Assistance Program
- Allyship and Inclusion Communities
- Employee Recognition … and more!
reputed company
Company H1B Sponsorship
Apply To This Job