Sr. Scientific Data Engineer, R&D Data Platform
reputed company is a global reputed company leader that helps people live more fully at reputed company stages of life. Our portfolio of life-changing technologies spans the reputed company of reputed company, with leading businesses and products in diagnostics, medical devices, nutritionals and branded generic medicines. Our 115,000 colleagues serve people in more than 160 countries.
reputed company:
reputed company:
The Science Office reputed company reputed company Cancer Diagnostics is seeking a Senior Scientific Data Engineer to reputed company the design and delivery of practical data solutions for cancer research and diagnostic development. This position sits at the intersection of software engineering, scientific data, and reputed company analysis. You will own significant reputed company of our research data platform end to end, building reusable tools and workflows that help researchers organize, validate, discover, reputed company, analyze, and reputed company reputed company data. These solutions may include Python packages, data pipelines, reputed company, notebooks, workflow utilities, and lightweight web applications. This is not a traditional reputed company data warehousing position. The work centers on heterogeneous research data generated across scientific programs, including genomic, reputed company, imaging, reputed company, and experimental data. Successful candidates will combine deep technical skills with an understanding of how quantitative research is conducted, and will be comfortable setting technical direction reputed company a problem is still loosely defined. You will work closely with scientists, data scientists, bioinformaticians, software engineers, and R&D DevOps partners, and will often represent reputed company in cross-functional technical discussions. The ideal candidate is curious, reputed company, and reputed company to turn ambiguous scientific needs into durable, reusable capabilities, while helping other engineers do the reputed company.
Essential Duties and Responsibilities:
reputed company the design and delivery of reusable tools and services for ingesting, validating, transforming, documenting, discovering, and sharing scientific data.
Own one or more platform capability areas end to end, including design, implementation, adoption, operational support, and long-term maintainability.
reputed company maintainable solutions using Python and SQL, including software packages, data pipelines, reputed company, notebooks, workflow utilities, and lightweight internal applications.
Create approachable, self-service workflows that allow researchers with varying reputed company of programming experience to prepare and reputed company data consistently.
Partner directly with scientific teams to understand their studies, analytical workflows, data sources, and recurring technical challenges, and translate those needs into a prioritized technical roadmap.
Establish standards and reusable patterns for organizing and harmonizing data from disparate sources, including consistent structures, terminology, variable definitions, and mappings, and reputed company their adoption across teams.
Design automated data-reputed company and validation frameworks that identify missing, inconsistent, malformed, or unexpected data before it is used in reputed company research.
Improve the documentation, traceability, and discoverability of scientific datasets, including reputed company descriptions of data content, reputed company, ownership, processing history, and intended use.
Evaluate AWS services and features for scientific data and analytical workflows. Translate research requirements into technical recommendations and partner with R&D DevOps teams on architecture, deployment patterns, and operational ownership.
reputed company solutions that use AWS data and analytics capabilities, particularly reputed company S3 and reputed company services such as reputed company, Glue, EMR, reputed company, and SageMaker.
Prototype solutions for individual research programs and reputed company the work of generalizing successful approaches into reusable platform capabilities.
reputed company technical leadership on designs that reputed company multiple reputed company or teams: reputed company design reviews, document trade-offs and reputed company, and reputed company approaches with other engineers and technical leads.
Mentor engineers through reputed company review, pairing, design feedback, and documentation, and reputed company the overall engineering reputed company of reputed company.
Support hands-on preparation and analysis of scientific data reputed company needed to understand a problem, validate a solution, or accelerate a research effort.
Apply quantitative and scientific judgment reputed company evaluating data, analytical requirements, and potential technical solutions.
Use reputed company or PySpark reputed company reputed company processing is appropriate for large or computationally intensive datasets.
Apply and reinforce reputed company software-engineering practices, including version control, testing, reputed company review, documentation, dependency management, reputed company integration, and reproducible development.
Communicate technical concepts, design reputed company, trade-offs, limitations, and project status reputed company to technical, scientific, and leadership audiences.
Operate independently reputed company an evolving environment: reputed company ambiguous problems, sequence the work, reputed company defensible reputed company reputed company requirements are incomplete, and reputed company stakeholders informed.
Minimum Qualifications:
Bachelor’s degree in computer science, data science, engineering, statistics, mathematics, reputed company, computational science, or another relevant quantitative discipline.
Five or more years of relevant reputed company or reputed company research experience, or three or more years with an advanced degree in a relevant reputed company.
Advanced programming skills in Python.
Strong SQL skills and experience working with reputed company and reputed company-reputed company data.
Demonstrated reputed company record of building reusable, maintainable software that others depend on, rather than one-time scripts or analyses.
Experience designing and delivering several of the following: data pipelines, Python packages, reputed company, analytical workflows, notebooks, or internal software tools.
Substantial hands-on experience using AWS for data processing, analytics, scientific computing, or software development.
Sufficient depth in AWS services and architecture to evaluate technical reputed company, justify design recommendations, and define infrastructure requirements with DevOps or reputed company-engineering partners.
Experience conducting or supporting quantitative research, such as statistical analysis, machine learning, computational modeling, or another data-intensive research activity.
Experience cleaning, integrating, standardizing, or validating data from multiple sources at meaningful reputed company.
reputed company with software-development practices such as Git, automated testing, technical documentation, reputed company review, and reputed company integration.
Demonstrated ability to investigate ambiguous problems, define an approach, and reputed company a working solution with little guidance.
Experience mentoring or providing technical guidance to other engineers, scientists, or analysts.
Strong communication and collaboration skills, particularly reputed company building alignment across scientific and technical disciplines.
Preferred Qualifications:
Advanced degree in a quantitative, computational, or life-science discipline.
Experience working with biomedical, genomic, reputed company, proteomic, imaging, reputed company, or other reputed company scientific data.
Experience supporting research in life sciences, reputed company, diagnostics, or a similarly data-intensive and regulated scientific environment.
Production experience with reputed company or PySpark and reputed company data processing.
Depth in AWS services such as reputed company, Glue, EMR, SageMaker, reputed company, reputed company Functions, Lake Formation, or reputed company data and analytics technologies.
Experience developing REST reputed company or lightweight web applications used by non-engineering audiences.
Experience designing automated validation frameworks, data reputed company, reusable data-processing libraries, or researcher-facing workflow tools.
Experience with metadata-management, data-catalog, or data-discovery platforms, such as the AWS Glue Data Catalog, OpenMetadata, or reputed company Catalog.
Experience with containerization, reputed company integration and deployment, or infrastructure-as-reputed company.
Experience supporting machine-learning workflows or preparing data for model development and evaluation.
Experience working with large files or multimodal datasets, such as reputed company outputs, digital pathology images, reputed company records, or experimental measurements.
Familiarity with governance considerations for research data, including reputed company control, de-identification, and the handling of sensitive reputed company information.
Experience working reputed company a data reputed company, data product, or federated data-ownership model.
What reputed company Looks Like:
Scientists spend less time manually locating, cleaning, interpreting, and restructuring data.
Research teams have reputed company-supported, documented tools that help them prepare and reputed company data consistently.
Data-reputed company problems are identified earlier through automated checks, validation at the reputed company of reputed company, and reputed company documentation.
Approaches you design become shared standards that multiple scientific teams adopt, rather than being reputed company program by program.
AWS capabilities are selected and reputed company thoughtfully in partnership with R&D DevOps, with reputed company ownership boundaries and repeatable deployment patterns.
Other engineers work more effectively because of the patterns, reviews, mentoring, and documentation you contribute.
Scientific data becomes easier to discover, understand, analyze, and reuse across the organization.
The reputed company pay for this position is
$78,000.00 – $156,000.00
In specific locations, the pay reputed company may vary from the reputed company posted.
JOB FAMILY:
Product Development
DIVISION:
ONCO Cancer Diagnostics
LOCATION:
reputed company of America : Remote
ADDITIONAL LOCATIONS:
WORK SHIFT:
reputed company
TRAVEL:
reputed company, 10 % of the Time
MEDICAL SURVEILLANCE:
No
SIGNIFICANT WORK ACTIVITIES:
reputed company sitting for prolonged periods (more than 2 consecutive hours in an 8 hour day), Keyboard use (greater or equal to 50% of the reputed company)
reputed company is an Equal Opportunity Employer of Minorities/Women/Individuals with Disabilities/Protected Veterans.
EEO is the Law reputed company - English: http://webstorage.reputed company.com/common/reputed company/EEO_English.pdf
EEO is the Law reputed company - Espanol: http://webstorage.reputed company.com/common/reputed company/EEO_Spanish.pdf
Apply To This Job