The Senior Data Scientist – Protein Data Pipelines will play a critical role in enabling predictive modeling for protein sequence, structure, and function by building scalable, reliable, and reproducible data pipelines. This role will focus on transforming protein property data and related scientific outputs into ML-amenable assets that support model training, inference, deployment, and ongoing use across research programs.
Working at the intersection of data engineering, MLOps, computational biology, and applied machine learning, this individual will partner with ML developers, wet-lab scientists, domain experts, and distributed technical teams to translate scientific and engineering needs into robust data and inference solutions. The successful candidate will develop reusable frameworks for data engineering, model inference, deployment, validation, testing, and monitoring across in-house and external machine learning models.
This role is ideal for someone who enjoys building production-ready scientific data systems, collaborating across disciplines, and converting complex domain needs into maintainable technical solutions that scale across discovery pipelines.
Bachelor's degree in Computational Biology, Bioinformatics, Life Sciences, Computational Chemistry, Chemical Engineering, Materials Science, Data Science, or a related quantitative field and relevant professional experience.
Success in this role will be demonstrated through:
The ideal candidate combines strong data-engineering and MLOps expertise with enough scientific domain fluency to work effectively with ML developers, and experimental collaborators. They enjoy building reusable systems that make complex scientific data reliable, reproducible, and actionable for predictive modeling.
Candidates may come from data science, data engineering, machine learning infrastructure, computational biology, computational chemistry, computational materials science, bioinformatics, or research informatics backgrounds. They are motivated by bridging scientific and engineering needs, supporting production-ready model use, and scaling technical solutions across discovery programs.
This role will build ML-amenable data pipelines for protein property data, mediate collaborations between ML developers and wet-lab teams, and scale data and modeling infrastructure across research programs and pipelines. By converting complex domain needs into maintainable, production-ready technical solutions, the Senior Data Scientist – Protein Data Pipelines will help accelerate reliable model development, deployment, and adoption across Large Molecule Discovery.
Tagged as: Life Sciences
Level 5 Ai Engineer Transformative Digital Capabilities (TDC) drives Process Development (PD) and Operations digital transformation through pragmatic digital innovation...
ApplyMedical Reporting And Analytics Manager Amgen harnesses the best of biology and technology to fight the world's toughest diseases, and...
ApplySenior Data Scientist – Protein Structure ML Models The Senior Data Scientist – Protein Structure ML Models will play a...
ApplyPrincipal Data Scientist The Future Begins Here At Takeda, we are leading digital evolution and global transformation. By building innovative...
ApplySenior Scientist – Scientific Data & ML Enablement (Large Molecule Discovery Informatics) In this vital role, you will enable AI-driven...
ApplySr. Data Scientist – Data Sciences & Artificial Intelligence Amgen harnesses the best of biology and technology to fight the...
ApplyPlease visit amgen.wd1.myworkdayjobs.com.
Don't forget to mention that you found the position on jobRxiv!
