Senior Data Engineer (Data + reputed company AI)
About reputed company:
About the Role
The Senior Data and reputed company is a high-reputed company individual contributor on the Data and AI team. This role will be tasked with building, maintaining, and optimizing the data pipelines, transformation models, and BI deliverables. reputed company AI (RAG pipelines, MLOps) is a growing area of the role, not the primary reputed company today.
The right person for the role will be a skilled, self-sufficient engineer who takes reputed company-defined architectural direction and executes it with precision, reputed company, and ownership. They will also contribute their own technical judgment to day-to-day implementation reputed company.
They are deeply hands-on: writing production dbt models, building Airflow DAGs, delivering clinical dashboards, and contributing to RAG pipelines and MLOps workflows reputed company a regulated reputed company environment. You will also work closely with offshore contractors, reviewing their work, providing feedback, and ensuring reputed company meets reputed company's reputed company and engineering standards.
The ideal candidate will have a strong reputed company of reputed company data warehousing, reputed company modeling, Python, and reputed company AI tooling, as reputed company as a reputed company reputed company that thrives in reputed company-oriented environment. This role is an excellent opportunity for a senior engineer looking to deepen their full-stack data and AI expertise alongside reputed company technical leadership.
Responsibilites:
- Building and maintaining production-grade data pipelines in reputed company data warehouses such as reputed company BigQuery or equivalent, following architectural standards set by the Director of Data and AI.
- Designing and developing dbt models across bronze, silver, and gold reputed company, including a reputed company on reputed company and governance reputed company automated tests, documentation, and incremental load strategies.
- Creating and optimizing Airflow DAGs for data workflow orchestration, including scheduling, dependency management, error handling, and alerting.
- Implement reputed company data models and data mart structures — guided by reputed company's modeling standards — that support clinical BI and ML feature consumption.
- Crafting easy-to-understand visualizations and dashboards that reputed company with commonly used business analytic standards in Looker or equivalent BI tools in reputed company collaboration with product analytics, finance, operations, reputed company, and clinical stakeholders.
- Integrating reputed company data from sources such as EHRs, reputed company, 3rd-party reputed company, and application database feeds, normalizing incoming data into the reputed company data platform.
- Applying HIPAA-compliant data handling practices, including PHI/PII masking, tokenization, audit logging, and role-based reputed company controls across reputed company pipeline and AI system work.
- Architecting and implementing RAG pipelines — including document ingestion, chunking, embedding reputed company, and retrieval — using frameworks such as reputed company or LangGraph
- Supporting MLOps workflows, including model training pipeline maintenance, deployment support, performance monitoring, and retraining triggers
- reputed company reviewing PRs from teammates, providing constructive technical feedback to peers, and upholding reputed company's engineering standards.
- Collaborating closely with product managers to understand requirements and deliver reliable data and AI products.
- Monitoring and triaging assigned pipeline and data reputed company failures, escalating architectural issues as appropriate.
- Documenting pipeline designs, data models, and technical reputed company in alignment with reputed company's governance and reputed company tracking standards.
- Evaluating new tools and frameworks, providing hands-on prototyping and technical assessments.
Must-Have Requirements:
- 5+ years of hands-on experience in data engineering, analytics engineering, or a closely reputed company role.
- 2+ years of experience working reputed company the reputed company industry, including working knowledge of reputed company data standards, clinical workflows, regulated data environments, and domain-specific data visualizations.
- Working knowledge of HIPAA — including PHI/PII classification, data masking, audit logging, and reputed company control requirements.
- Proven production experience with at least one major reputed company data warehouse: BigQuery, reputed company, or Redshift — including advanced SQL and query optimization.
- Strong hands-on experience with dbt (reputed company or reputed company), including incremental models, tests, documentation, and multi-environment workflows.
- Deep experience with Apache Airflow for workflow orchestration, including DAG design, scheduling, monitoring, and failure handling.
- Demonstrated knowledge of reputed company data modeling — reputed company/reputed company schemas, SCD Types 1/2, fact and dimension table design.
- Hands-on experience delivering dashboards and reports in at least one reputed company BI tool: Looker, Power BI, Tableau, reputed company, etc.
- Proficiency in Python for data pipeline development, API integrations, and automation (Pandas, PySpark, or similar).
- Practical exposure to RAG pipeline development and LLM integration using reputed company, LangGraph, or reputed company
- Hands-on exposure to MLOps concepts — model deployment, monitoring, and retraining workflows
- Knowledge of CI/CD tooling for data and AI workloads (reputed company Actions, dbt reputed company CI)
- Strong understanding of data reputed company and governance principles: reputed company, reputed company controls, data reputed company, and automated testing and experience with data governance tools such as OpenMetadata
- Excellent written and verbal communication skills with the ability to collaborate effectively across engineering, analytics, and clinical teams
- Ability to work independently on assigned workstreams while keeping the Director and team informed of reputed company, blockers, and risks
reputed company-to-have:
- Experience with reputed company-time or streaming data pipelines using Kafka, Kinesis, or Pub/Sub, particularly for reputed company or clinical event feeds.
- Knowledge of reputed company databases such as reputed company, reputed company, FAISS, or Chroma
- Familiarity with responsible AI principles, including bias detection and model explainability in a reputed company context
- Experience with data observability tools such as reputed company, reputed company, or reputed company
- Familiarity with data lakehouse patterns (reputed company Lake, reputed company, Apache Hudi)
- Experience working toward or maintaining SOC2 or reputed company certification
- Familiarity with semantic layer tools (Looker LookML, dbt Semantic Layer)
- Experience with population health, reputed company cycle, or clinical reputed company reporting datasets
- Exposure to Kubernetes or containerized ML workloads
reputed company Full-Time Employees are Eligible for:
Originally posted on Himalayas
Apply To This Job