Machine Learning Research Engineer

Profluent
Location
Emeryville, California, United States; Hybrid (2-3 days on-site)
Job Type
Full-time
Posted
July 23, 2026
Views
7

Job Description

Profluent is an AI-first protein design company. Founded in 2022, we develop deep generative models to design and validate novel, functional proteins to revolutionize biomedicine. Based in Emeryville, CA, we are backed by leading investors including Altimeter Capital, Bezos Expeditions, Spark Capital, Insight Partners, Air Street Capital, AIX Ventures, and Convergent Ventures, and have raised over $150M to date.

We're looking for an experiencedMachine Learning Engineerto build and improve the models and ML systems that drive our protein design efforts. In this role, you'll deploy and optimize large-scale generative models for protein design, and develop the surrounding infrastructure and tooling that enable our ML and protein design scientists to work faster and more confidently. As an early member of a small, fast-moving engineering team, you'll have significant ownership over our ML stack and the opportunity to shape how our platform evolves.

Responsibilities

  • Build robust, reproducible and user-friendly pipelines for automated model fine-tuning, alignment and evaluation
  • Design and implement  modular, easy-to-maintain, multi-model pipelines for protein design.
  • Develop highly scalable ETL pipelines to process petabyte-scale protein data for model pretraining
  • Optimize model training and inference code to maximize throughput and resource utilization when deployed at scale
  • Develop software and infrastructure that enable the ML team to work quickly and frictionlessly in distributed and multi-cloud environments
  • Partner with ML and protein design scientists to prototype research ideas and bring them into production

Who You Are

  • You're comfortable taking ownership and working independently in a fast-moving environment
  • You're an execution-oriented engineer who maintains high standards, and focuses on the highest-impact work
  • You're comfortable owning the full stack of your work, from training code to the infrastructure it runs on
  • You care deeply about model quality, efficiency, and reliability
  • You're willing to step beyond your core responsibilities when the team needs it

Representative Projects

  • Building hyperparameter search frameworks for SFT and Alignment workflows
  • Increasing protein language model throughput during long context generation
  • Updating existing model architectures to work and run efficiently on new GPU hardware
  • Implementing a protein design pipeline that integrates prompt retrieval, sequence generation, attribute prediction, and structure prediction
  • Establishing an ETL pipeline for sampling and tokenizing training datasets from an internal database of billions of sequences
  • Developing a benchmarking and evaluation system for newly trained sequence generation models
  • Contributing to the development of an internal service that provides transparent multi-node job submission for ML scientists

Qualifications

  • BS or MS in Computer Science, Machine Learning, or a related field
  • 3+ years of hands-on experience building and training ML models in PyTorch
  • Strong Python and software engineering fundamentals, including testing, code quality, and version control
  • Experience profiling, benchmarking, and optimizing ML model training and inference
  • Experience implementing or optimizing transformer-based architectures
  • Familiarity with cloud infrastructure and containerization (GCP, AWS, Azure, Kubernetes, Docker)
  • Strong fundamentals in ML, statistics, and/or linear algebra

Preferences

  • Familiarity with protein language models or computational biology
  • Experience with GPU-level optimization (CUDA, Triton)
  • Experience with distributed training (DDP, FSDP, multi-node GPU clusters)
  • Experience with databases and data processing pipelines
  • Experience orchestrating multi-step ML workflows
  • Experience building backend systems that serve ML models in production
  • Contributions to open source ML projects or published research

What We Offer

Profluent is an AI-first protein design company. Founded in 2022, we develop deep generative models to design and validate novel, functional proteins to revolutionize biomedicine. Based in Emeryville, CA, we are backed by leading investors including Altimeter Capital, Bezos Expeditions, Spark Capital, Insight Partners, Air Street Capital, AIX Ventures, and Convergent Ventures, and have raised over $150M to date.

We're looking for an experiencedMachine Learning Engineerto build and improve the models and ML systems that drive our protein design efforts. In this role, you'll deploy and optimize large-scale generative models for protein design, and develop the surrounding infrastructure and tooling that enable our ML and protein design scientists to work faster and more confidently. As an early member of a small, fast-moving engineering team, you'll have significant ownership over our ML stack and the opportunity to shape how our platform evolves.

Responsibilities

  • Build robust, reproducible and user-friendly pipelines for automated model fine-tuning, alignment and evaluation
  • Design and implement  modular, easy-to-maintain, multi-model pipelines for protein design.
  • Develop highly scalable ETL pipelines to process petabyte-scale protein data for model pretraining
  • Optimize model training and inference code to maximize throughput and resource utilization when deployed at scale
  • Develop software and infrastructure that enable the ML team to work quickly and frictionlessly in distributed and multi-cloud environments
  • Partner with ML and protein design scientists to prototype research ideas and bring them into production

Who You Are

  • You're comfortable taking ownership and working independently in a fast-moving environment
  • You're an execution-oriented engineer who maintains high standards, and focuses on the highest-impact work
  • You're comfortable owning the full stack of your work, from training code to the infrastructure it runs on
  • You care deeply about model quality, efficiency, and reliability
  • You're willing to step beyond your core responsibilities when the team needs it

Representative Projects

  • Building hyperparameter search frameworks for SFT and Alignment workflows
  • Increasing protein language model throughput during long context generation
  • Updating existing model architectures to work and run efficiently on new GPU hardware
  • Implementing a protein design pipeline that integrates prompt retrieval, sequence generation, attribute prediction, and structure prediction
  • Establishing an ETL pipeline for sampling and tokenizing training datasets from an internal database of billions of sequences
  • Developing a benchmarking and evaluation system for newly trained sequence generation models
  • Contributing to the development of an internal service that provides transparent multi-node job submission for ML scientists

Qualifications

  • BS or MS in Computer Science, Machine Learning, or a related field
  • 3+ years of hands-on experience building and training ML models in PyTorch
  • Strong Python and software engineering fundamentals, including testing, code quality, and version control
  • Experience profiling, benchmarking, and optimizing ML model training and inference
  • Experience implementing or optimizing transformer-based architectures
  • Familiarity with cloud infrastructure and containerization (GCP, AWS, Azure, Kubernetes, Docker)
  • Strong fundamentals in ML, statistics, and/or linear algebra

Preferences

  • Familiarity with protein language models or computational biology
  • Experience with GPU-level optimization (CUDA, Triton)
  • Experience with distributed training (DDP, FSDP, multi-node GPU clusters)
  • Experience with databases and data processing pipelines
  • Experience orchestrating multi-step ML workflows
  • Experience building backend systems that serve ML models in production
  • Contributions to open source ML projects or published research

What We Offer

  • High-growth opportunity with meaningful impact on the future of protein design
  • Competitive compensation package with equity participation
  • 401(k) with a strong employer match
  • Comprehensive benefits including health/dental/vision insurance
  • Generous PTO policy and commitment to work-life balance
  • Professional development opportunities in a cutting-edge field at the intersection of AI and biology

Profluent Bio, Inc is an equal opportunity employer promoting diversity and inclusion in the workspace. We do not discriminate on the basis of race, color, religion, marital status, age, national origin, ancestry, physical or mental disability, medical conditions, veteran status, sexual orientation, gender (including gender identity and gender expression), sex (which includes pregnancy, childbirth, and breastfeeding), genetic information, taking or requesting statutorily protected leave, or any other basis protected by law.

Employment Eligibility Verification

Legal authorization to work in the United States is required. In compliance with federal law, all persons hired must verify their identity and work eligibility and complete the required employment verification form upon hire.

Hiring Salary Range

$200,000—$330,000 USD

Frequently Asked Questions

Where is the job located, and is it remote/hybrid/on-site?
The position is located in Emeryville, California, United States, and operates on a hybrid model requiring 2-3 days on-site.
What are the required qualifications and experience level for this role?
You need a BS or MS in Computer Science, Machine Learning, or a related field, and 3+ years of hands-on experience building and training ML models in PyTorch. Strong Python, software engineering, and ML/statistics fundamentals are required, alongside experience optimizing transformer-based architectures and profiling ML model training.
What are the key responsibilities of the Machine Learning Research Engineer?
You will build pipelines for automated model fine-tuning, alignment, and evaluation, and design multi-model pipelines for protein design. You will also develop scalable ETL pipelines for petabyte-scale data, optimize training and inference code, build distributed/multi-cloud infrastructure, and partner with scientists to prototype and productionize research ideas.
What is the salary range for this position?
The hiring salary range for this role is $200,000 to $330,000 USD.
Does Profluent offer visa sponsorship for this role?
No. Legal authorization to work in the United States is required, and all hires must verify their identity and work eligibility upon hire.
What benefits and compensation perks are provided?
Profluent offers a competitive compensation package with equity participation, a 401(k) with a strong employer match, comprehensive health, dental, and vision insurance, a generous PTO policy, and professional development opportunities.

Ready to Apply?

Apply for this Position

You'll be redirected to the company's application page

Share this job:

Job Information

Source: greenhouse
AI Relevance: 92/100 (Highly relevant)
Remote Type: hybrid
Allowed Locations: Emeryville, California, United States; Hybrid (2-3 days on-site)
Skills & Tags:
Machine Learning

Get Similar Jobs by Email

Weekly digest of Profluent and similar companies. Free.

Related Jobs

Apply for this Position

Get weekly job alerts