Senior Data Engineer, Bioinformatics, Cheminformatics, Materials

Lilasciences
Location
San Francisco, CA USA
Job Type
Full-time
Posted
September 2, 2026
Views
1

Job Description

Your Impact at LILA

Lila’s mission is to accelerate scientific discovery with AI, and that depends on trustworthy scientific data. As a Data Engineer, you’ll build ETL pipelines and data models for Lila’s scientific data platform, working at the intersection of data engineering, computational biology, chemistry, and materials science.

You’ll partner with AI researchers and experimentalists to turn raw lab instrument outputs into validated, analysis-ready datasets. The core challenge is data modeling: transforming messy, per-instrument measurements into clean, well-typed data that is efficient to query, reliable to use, and ready for downstream analysis.

You’ll also build domain-specific analysis functions and reusable data pipelines that help scientists and AI researchers move faster without re-deriving bespoke solutions.

What You'll Be Building

• Design pipelines that turn raw lab output into analysis-ready scientific data. • Model heterogeneous data from bio, chemistry, and materials instruments. • Build validation checks, schema-evolution gates, and data quality workflows. • Develop reusable analysis functions for scientific and AI research workflows. • Improve automation and observability across instrument-to-result data flows. • Build canonical datasets that scientists and AI researchers can trust. • Use AI coding tools to accelerate pipeline development and team velocity.

What You'll Need to Succeed

• 2–6 years of experience in data engineering, bioinformatics, cheminformatics, or computational science. •
Strong Python skills, including typed, tested, production-quality code.
• Strong SQL skills, especially with Postgres or similar relational databases.
• Experience building ETL pipelines, data models, and reusable data transformations.
• Data science foundation, including statistics and pandas, NumPy, or similar tools.
• Experience translating noisy scientific measurements into accurate, validated datasets.
• Workflow orchestration experience, ideally Flyte, Airflow, Prefect, Dagster, or Nextflow.
• Active use of AI coding tools in day-to-day engineering work.

Bonus Points For

• Experience with columnar or lakehouse stacks such as Parquet, Iceberg, DuckDB, Polars, or Ibis.
• Familiarity with event-driven pipelines such as NATS or Kafka. • Exposure to lab instrument data formats, LIMS, or ELN systems.
• Familiarity with life sciences assays, sequencing, imaging, or flow cytometry.
• Familiarity with materials or chemistry methods such as XRD, XRF, SEM, TGA, or DSC.
• Experience with curve fitting, peak detection, or unit and dimensional analysis.

Compensation

We offer competitive base compensation with bonus potential and generous early-stage equity. Your final offer will reflect your background, expertise, and expected impact.

U.S. Benefits.Full-time U.S. employees receive a comprehensive benefits program including medical, dental, and vision coverage; employer-paid life and disability insurance; flexible time off with generous company wide holidays; paid parental leave; an educational assistance program; commuter benefits, including bike share memberships for office based employees; and a company subsidized lunch program.

International Benefits.Full-time employees outside the U.S. receive a comprehensive benefits program tailored to their region. USD salary ranges apply only to U.S.-based positions; international salaries are set to local market.

Expected Base Salary Range

$144,000—$240,000 USD

About LILA

Lila Sciences is building Scientific Superintelligence™ to solve humankind's greatest challenges. We believe science is the most inspiring frontier for AI. Rather than hard-coding expert knowledge into tools, LILA builds systems that can learn for themselves.

LILA combines advanced AI models with proprietary AI Science Factory™ instruments into an operating system for science that executes the entire scientific method autonomously, accelerating discovery at unprecedented speed, scale, and impact across medicine, materials, and energy. Learn more at www.lila.ai.

Your Impact at LILA

Lila’s mission is to accelerate scientific discovery with AI, and that depends on trustworthy scientific data. As a Data Engineer, you’ll build ETL pipelines and data models for Lila’s scientific data platform, working at the intersection of data engineering, computational biology, chemistry, and materials science.

You’ll partner with AI researchers and experimentalists to turn raw lab instrument outputs into validated, analysis-ready datasets. The core challenge is data modeling: transforming messy, per-instrument measurements into clean, well-typed data that is efficient to query, reliable to use, and ready for downstream analysis.

You’ll also build domain-specific analysis functions and reusable data pipelines that help scientists and AI researchers move faster without re-deriving bespoke solutions.

What You'll Be Building

• Design pipelines that turn raw lab output into analysis-ready scientific data. • Model heterogeneous data from bio, chemistry, and materials instruments. • Build validation checks, schema-evolution gates, and data quality workflows. • Develop reusable analysis functions for scientific and AI research workflows. • Improve automation and observability across instrument-to-result data flows. • Build canonical datasets that scientists and AI researchers can trust. • Use AI coding tools to accelerate pipeline development and team velocity.

What You'll Need to Succeed

• 2–6 years of experience in data engineering, bioinformatics, cheminformatics, or computational science. •
Strong Python skills, including typed, tested, production-quality code.
• Strong SQL skills, especially with Postgres or similar relational databases.
• Experience building ETL pipelines, data models, and reusable data transformations.
• Data science foundation, including statistics and pandas, NumPy, or similar tools.
• Experience translating noisy scientific measurements into accurate, validated datasets.
• Workflow orchestration experience, ideally Flyte, Airflow, Prefect, Dagster, or Nextflow.
• Active use of AI coding tools in day-to-day engineering work.

Bonus Points For

• Experience with columnar or lakehouse stacks such as Parquet, Iceberg, DuckDB, Polars, or Ibis.
• Familiarity with event-driven pipelines such as NATS or Kafka. • Exposure to lab instrument data formats, LIMS, or ELN systems.
• Familiarity with life sciences assays, sequencing, imaging, or flow cytometry.
• Familiarity with materials or chemistry methods such as XRD, XRF, SEM, TGA, or DSC.
• Experience with curve fitting, peak detection, or unit and dimensional analysis.

Compensation

We offer competitive base compensation with bonus potential and generous early-stage equity. Your final offer will reflect your background, expertise, and expected impact.

U.S. Benefits.Full-time U.S. employees receive a comprehensive benefits program including medical, dental, and vision coverage; employer-paid life and disability insurance; flexible time off with generous company wide holidays; paid parental leave; an educational assistance program; commuter benefits, including bike share memberships for office based employees; and a company subsidized lunch program.

International Benefits.Full-time employees outside the U.S. receive a comprehensive benefits program tailored to their region. USD salary ranges apply only to U.S.-based positions; international salaries are set to local market.

Expected Base Salary Range

$144,000—$240,000 USD

About LILA

Lila Sciences is building Scientific Superintelligence™ to solve humankind's greatest challenges. We believe science is the most inspiring frontier for AI. Rather than hard-coding expert knowledge into tools, LILA builds systems that can learn for themselves.

LILA combines advanced AI models with proprietary AI Science Factory™ instruments into an operating system for science that executes the entire scientific method autonomously, accelerating discovery at unprecedented speed, scale, and impact across medicine, materials, and energy. Learn more at www.lila.ai.

Guided by our core values of truth, trust, curiosity, grit, and velocity, we move with startup speed while tackling problems of historic importance. If this sounds like an environment you'd love to work in, even if you don't meet every qualification listed above, we encourage you to apply.

We’re All In

Lila Sciences is committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status.

Information you provide during your application process will be handled in accordance with ourCandidate Privacy Policy.

A Note to Agencies

Lila Sciences does not accept unsolicited resumes from any source other than candidates. The submission of unsolicited resumes by recruitment or staffing agencies to Lila Sciences or its employees is strictly prohibited unless contacted directly by Lila Science’s internal Talent Acquisition team. Any resume submitted by an agency in the absence of a signed agreement will automatically become the property of Lila Sciences, and Lila Sciences will not owe any referral or other fees with respect thereto.

Frequently Asked Questions

Where is the job located, and is it remote/hybrid/on-site?
The position is located in San Francisco, CA USA. The job posting does not specify a remote, hybrid, or on-site work-mode policy, though it mentions commuter benefits and a subsidized lunch program for office-based employees.
What are the required qualifications and experience level for this role?
You need 2–6 years of experience in data engineering, bioinformatics, cheminformatics, or computational science. Requirements include strong Python and SQL (especially Postgres) skills, experience building ETL pipelines and data models, a data science foundation (statistics, pandas, NumPy), workflow orchestration experience, and active use of AI coding tools.
What are the key responsibilities of the Senior Data Engineer?
You will design ETL pipelines and model heterogeneous data from biological, chemical, and materials instruments. Responsibilities include building validation checks, schema-evolution gates, and reusable analysis functions, improving automation and observability across data flows, and partnering with AI researchers and experimentalists to create trusted, analysis-ready datasets.
What is the salary range for this position?
The expected base salary range for U.S.-based positions is $144,000 to $240,000 USD. The company also offers competitive base compensation with bonus potential and generous early-stage equity.
What benefits does the company offer?
Full-time U.S. employees receive medical, dental, and vision coverage, employer-paid life and disability insurance, flexible time off, paid parental leave, educational assistance, commuter benefits (including bike share memberships), and a subsidized lunch program. International employees receive a comprehensive benefits program tailored to their region.

Ready to Apply?

Apply for this Position

You'll be redirected to the company's application page

Share this job:

Job Information

Source: greenhouse
AI Relevance: 85/100 (Highly relevant)
Remote Type: onsite
Allowed Locations: San Francisco, CA USA
Skills & Tags:
Software
Explore related roles:

Get Similar Jobs by Email

Weekly digest of Lilasciences and similar companies. Free.

Related Jobs

Apply for this Position

Get weekly job alerts