Scientist II / Senior ML Scientist, Data-Efficient Learning for Drug Discovery

Lilasciences
Location
Cambridge, MA USA; London, UK; San Francisco, CA USA
Job Type
Full-time
Posted
September 2, 2026
Views
2

Job Description

Your Impact at LILA

Lila Sciences is seeking a Machine Learning Scientist, Data-Efficient Learning for Drug Discovery to build models and learning strategies for settings where data is scarce, expensive, and intentionally generated. This role is focused on training useful models from low-quantity but high-quality datasets ranging from as few as tens to low thousands of examples, often in tightly focused areas of chemical space, and deciding what data should be acquired next.

This is an applied scientific ML role in a frontier research area. The work is not a matter of applying standard models out of the box. You will use and develop approaches across active learning, meta-learning, fine-tuning, uncertainty estimation, experimental design, and multimodal modeling to help Lila build closed-loop systems that learn efficiently from targeted data acquisition.

This role connects model training with scientific decision-making: data acquisition plans should be useful to computational chemists evaluating compound priorities, computational biophysicists deciding when simulation is warranted, and cofolding modelers deciding which protein-ligand data would improve structure-aware models.

What You'll Be Building

  • Build ML models that perform well in low-data regimes for drug discovery and molecular optimization.
  • Design data acquisition strategies that identify which compounds, assays, DEL selections, simulations, structural predictions, or experiments should be run next to maximize learning.
  • Develop active learning, meta-learning, fine-tuning, transfer learning, and uncertainty-aware modeling approaches for focused chemical spaces.
  • Train models on low-quantity, high-quality datasets generated by Lila's experimental, computational, and agentic discovery systems.
  • Build multimodal models that can integrate DEL data, simulation outputs, assay data, protein and structural information, chemical features, literature or text-derived signals, images, and experimental metadata.
  • Partner with experimental, computational, and drug discovery teams to ensure data acquisition plans are scientifically meaningful and operationally feasible.
  • Evaluate models through learning curves, prospective validation, retrospective benchmarks, uncertainty calibration, and decision-focused metrics.
  • Develop closed-loop learning workflows that continuously update models as new data arrives from experiments, simulations, and automated systems.
  • Translate model predictions and uncertainty into practical recommendations for compound selection, assay selection, batch design, or next experiments.
  • Work with platform and agent teams to expose model-driven recommendations as tools for scientists and AI agents.

What You'll Need to Succeed

  • PhD or equivalent experience in machine learning, computational chemistry, computational biology, statistics, computer science, bioengineering, or a related field.
  • Strong experience training ML models in low-data regimes.
  • Experience with active learning, Bayesian optimization, experimental design, meta-learning, fine-tuning, transfer learning, uncertainty estimation, or related data-efficient learning methods.
  • Experience building ML models for scientific, molecular, biological, chemical, pharmacological, biochemical, or other high-dimensional experimental datasets.
  • Experience with multimodal learning or methods that combine heterogeneous data sources.
  • Ability to reason about data acquisition strategy, not only model fitting.
  • Strong scientific judgment and ability to connect model behavior to experimental decisions.
  • Practical experience with PyTorch, JAX, scikit-learn, or equivalent ML tools.
  • Ability to collaborate across ML, data, computational science, experimental, and drug discovery teams.

Bonus Points For

Your Impact at LILA

Lila Sciences is seeking a Machine Learning Scientist, Data-Efficient Learning for Drug Discovery to build models and learning strategies for settings where data is scarce, expensive, and intentionally generated. This role is focused on training useful models from low-quantity but high-quality datasets ranging from as few as tens to low thousands of examples, often in tightly focused areas of chemical space, and deciding what data should be acquired next.

This is an applied scientific ML role in a frontier research area. The work is not a matter of applying standard models out of the box. You will use and develop approaches across active learning, meta-learning, fine-tuning, uncertainty estimation, experimental design, and multimodal modeling to help Lila build closed-loop systems that learn efficiently from targeted data acquisition.

This role connects model training with scientific decision-making: data acquisition plans should be useful to computational chemists evaluating compound priorities, computational biophysicists deciding when simulation is warranted, and cofolding modelers deciding which protein-ligand data would improve structure-aware models.

What You'll Be Building

  • Build ML models that perform well in low-data regimes for drug discovery and molecular optimization.
  • Design data acquisition strategies that identify which compounds, assays, DEL selections, simulations, structural predictions, or experiments should be run next to maximize learning.
  • Develop active learning, meta-learning, fine-tuning, transfer learning, and uncertainty-aware modeling approaches for focused chemical spaces.
  • Train models on low-quantity, high-quality datasets generated by Lila's experimental, computational, and agentic discovery systems.
  • Build multimodal models that can integrate DEL data, simulation outputs, assay data, protein and structural information, chemical features, literature or text-derived signals, images, and experimental metadata.
  • Partner with experimental, computational, and drug discovery teams to ensure data acquisition plans are scientifically meaningful and operationally feasible.
  • Evaluate models through learning curves, prospective validation, retrospective benchmarks, uncertainty calibration, and decision-focused metrics.
  • Develop closed-loop learning workflows that continuously update models as new data arrives from experiments, simulations, and automated systems.
  • Translate model predictions and uncertainty into practical recommendations for compound selection, assay selection, batch design, or next experiments.
  • Work with platform and agent teams to expose model-driven recommendations as tools for scientists and AI agents.

What You'll Need to Succeed

  • PhD or equivalent experience in machine learning, computational chemistry, computational biology, statistics, computer science, bioengineering, or a related field.
  • Strong experience training ML models in low-data regimes.
  • Experience with active learning, Bayesian optimization, experimental design, meta-learning, fine-tuning, transfer learning, uncertainty estimation, or related data-efficient learning methods.
  • Experience building ML models for scientific, molecular, biological, chemical, pharmacological, biochemical, or other high-dimensional experimental datasets.
  • Experience with multimodal learning or methods that combine heterogeneous data sources.
  • Ability to reason about data acquisition strategy, not only model fitting.
  • Strong scientific judgment and ability to connect model behavior to experimental decisions.
  • Practical experience with PyTorch, JAX, scikit-learn, or equivalent ML tools.
  • Ability to collaborate across ML, data, computational science, experimental, and drug discovery teams.

Bonus Points For

  • Drug discovery experience, especially in molecular optimization, screening, or design-make-test-learn workflows.
  • General understanding of pharmacology, biochemistry, or mechanisms of molecular activity.
  • Experience with DEL, high-throughput screening, medicinal chemistry, assay data, simulation-derived features, protein or structure-based features, text or literature features, or scientific images.
  • Experience with closed-loop experimentation, autonomous labs, or agent-driven scientific workflows.
  • Experience with generative molecular design, candidate prioritization, or batch selection workflows.
  • Familiarity with causal inference, optimal experimental design, decision theory, or Bayesian methods.
  • Comfort working with frontier ML techniques where standard out-of-the-box approaches are insufficient.

Compensation

We offer competitive base compensation with bonus potential and generous early-stage equity. Your final offer will reflect your background, expertise, and expected impact.

U.S. Benefits.Full-time U.S. employees receive a comprehensive benefits program including medical, dental, and vision coverage; employer-paid life and disability insurance; flexible time off with generous company wide holidays; paid parental leave; an educational assistance program; commuter benefits, including bike share memberships for office based employees; and a company subsidized lunch program.

International Benefits.Full-time employees outside the U.S. receive a comprehensive benefits program tailored to their region. USD salary ranges apply only to U.S.-based positions; international salaries are set to local market.

Expected Base Salary Range

$228,000—$358,000 USD

About LILA

Lila Sciences is building Scientific Superintelligence™ to solve humankind's greatest challenges. We believe science is the most inspiring frontier for AI. Rather than hard-coding expert knowledge into tools, LILA builds systems that can learn for themselves.

LILA combines advanced AI models with proprietary AI Science Factory™ instruments into an operating system for science that executes the entire scientific method autonomously, accelerating discovery at unprecedented speed, scale, and impact across medicine, materials, and energy. Learn more at www.lila.ai.

Guided by our core values of truth, trust, curiosity, grit, and velocity, we move with startup speed while tackling problems of historic importance. If this sounds like an environment you'd love to work in, even if you don't meet every qualification listed above, we encourage you to apply.

We’re All In

Lila Sciences is committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status.

Information you provide during your application process will be handled in accordance with ourCandidate Privacy Policy.

A Note to Agencies

Lila Sciences does not accept unsolicited resumes from any source other than candidates. The submission of unsolicited resumes by recruitment or staffing agencies to Lila Sciences or its employees is strictly prohibited unless contacted directly by Lila Science’s internal Talent Acquisition team. Any resume submitted by an agency in the absence of a signed agreement will automatically become the property of Lila Sciences, and Lila Sciences will not owe any referral or other fees with respect thereto.

Frequently Asked Questions

Where is the job located, and is it remote/hybrid/on-site?
The position is located in Cambridge, MA USA; London, UK; or San Francisco, CA USA. The job posting does not specify a remote, hybrid, or on-site work-mode policy.
What are the key responsibilities of this role?
You will build ML models for low-data regimes, design data acquisition strategies, and develop active learning, meta-learning, and uncertainty-aware modeling approaches. You will also train models on low-quantity, high-quality datasets, build multimodal models, partner with experimental and drug discovery teams, and translate model predictions into practical recommendations.
What qualifications and experience do I need to apply?
You need a PhD or equivalent experience in machine learning, computational chemistry/biology, statistics, computer science, bioengineering, or a related field. You must have experience training ML models in low-data regimes, using active learning or Bayesian optimization, working with scientific/molecular datasets, and using PyTorch, JAX, or scikit-learn.
What is the salary range for this position?
The expected base salary range for U.S.-based positions is $228,000 to $358,000 USD. International salaries are set to local market rates.
What benefits does the company offer?
U.S. benefits include medical, dental, and vision coverage; life and disability insurance; flexible time off and company holidays; paid parental leave; educational assistance; commuter benefits (including bike share memberships); and subsidized lunches. International employees receive a comprehensive benefits program tailored to their region.

Ready to Apply?

Apply for this Position

You'll be redirected to the company's application page

Share this job:

Job Information

Source: greenhouse
AI Relevance: 92/100 (Highly relevant)
Remote Type: onsite
Allowed Locations: Cambridge, MA USA; London, UK; San Francisco, CA USA
Skills & Tags:
Physical Sciences AI

Get Similar Jobs by Email

Weekly digest of Lilasciences and similar companies. Free.

Related Jobs

Apply for this Position

Get weekly job alerts