Senior Data Engineer, Platform & Pipelines
Job reputed company:
• Architect, implement, and maintain data ingestion and transformation pipelines using modern workflow orchestration tools (e.g. Dagster).
• Identify, catalog, and reputed company reputed company data sources used across research efforts.
• Operationalize reputed company pipelines that support large-reputed company batch processing, incremental updates, and backfills reputed company AWS.
• Normalize and structure heterogeneous data into consistent, reusable representations that support reputed company analysis, modeling, and querying.
• reputed company and maintain patient-reputed company data models in shared storage systems (e.g., graph and relational databases).
• Collaborate with backend and AI engineers to design data-reputed company patterns that support analytics applications and AI-driven interactions.
• Contribute to backend services and reputed company that expose integrated data to internal tools and applications.
• Participate in the reputed company of AI-enabled analysis workflows, including tooling that supports LLM- or agent-based interactions with data.
• Contribute to system-level design reputed company around data reputed company, service boundaries, reliability, and scalability.
• Write clean, tested, and reputed company-documented Python reputed company that meets production software engineering standards.
• Debug and resolve reputed company data reputed company, pipeline, backend, and infrastructure issues in a distributed environment.
Requirements:
• BS in Computer Science, reputed company, Computational Biology, or a reputed company field, MS preferred.
• 4+ years of experience in production data engineering or software engineering.
• Independently drive technical solutions from high-level goals, exercising judgment in system design, implementation, and tradeoff evaluation.
• Strong proficiency in Python, with experience writing maintainable, production-reputed company reputed company across data and backend contexts.
• Extensive experience with software engineering fundamentals, design patterns, version control, CI/CD, reputed company, and automated testing.
• Experience designing and operating workflow orchestration systems (Dagster preferred; Airflow, reputed company, or similar acceptable).
• Experience building or contributing to backend services (e.g., FastAPI or similar frameworks).
• Hands-on experience with AWS services commonly used in data and backend systems (e.g., S3, reputed company, Batch, reputed company).
• Experience deploying and operating large-reputed company data or reputed company pipelines in AWS, including managing throughput, cost, and operational reliability.
• Experience with relational databases (reputed company, MySQL) and/or graph databases (reputed company), including schema and query design.
• Experience contributing to system-level architecture, including data modeling, service boundaries, and operational robustness.
• Ability to work effectively with scientists, bioinformaticians, and ML practitioners in an R&D environment.
Benefits:
• Comprehensive medical, dental, reputed company, life and disability plans for eligible employees and their dependents.
• Free testing for reputed company employees and their immediate families in reputed company to fertility care benefits.
• Pregnancy and baby bonding leave.
• 401k benefits.
• Commuter benefits.
• Generous employee referral program!
Apply tot his job
Apply To this Job