machine.learning.bio

We build AI that reads, writes, and searches the language of biology.

The Dallago lab at Duke University.

What we do

We ask what machine-learned representations can reveal about biology — and where they fail — then use those answers to build more reliable tools for discovery and protein engineering.

Our work spans scalable biological inference, rigorous evaluation of predictive models, and generative design of proteins and molecular interactions. We pair large-scale computation with careful benchmarking and experimental collaboration to identify what models can reliably tell us about biology.

Research

Representation learning for biology

Teaching models to read proteins and genomes as languages.

Generative and programmable design

Writing new proteins, atom by atom, and testing them at the bench.

Search and structure at scale

Making biological search and structure prediction fast enough for whole proteomes.

Benchmarks and honest evaluation

Measuring whether these models actually work where it matters.

Selected work

Preprint 2026 · NVIDIA + Duke

AlphaFold Database expands to proteome-scale quaternary structures

Most proteins do not work alone, but the world's structure database showed them alone. We predicted 31 million candidate complexes across 4,777 proteomes, established confidence criteria against experimentally determined structures, and released 1.81 million high-confidence predictions. 31.3% of them extend beyond anything detectable in the PDB.

ICLR 2026 Oral · NVIDIA

Proteina-Complexa

Binder design has been split between conditional generation and hallucination. We argue that is a false dichotomy, and built one fully atomistic model that does both — then showed it keeps improving when you let it think longer at inference time.

Nature Methods 2025 · NVIDIA

GPU-accelerated homology search with MMseqs2

The bottleneck in structure prediction was never only the network — it was the search in front of it. Redesigning the core alignment algorithms for GPUs reached roughly 100 trillion cell updates per second, and kept it working on low-wattage cards so the speed-up is not reserved for people with a cluster.

ICML 2026 Oral · NVIDIA + Duke

FLIP2

Seven new fitness datasets and splits built to mirror real protein-engineering campaigns rather than random train/test partitions. Across them, simpler models often matched or beat fine-tuned protein language models — a result worth publishing precisely because it is inconvenient.

All publications

News

An updated preprint of the AlphaFold Database quaternary expansion is posted. The manuscript is under review.
FLIP2 is accepted as an Oral at ICML 2026 — one of 168 orals.
Proteina-Complexa is posted, and accepted as an Oral at ICLR 2026.

News archive

Two roles, two doors

Chris is a Senior Research Scientist and Applied Research Team Lead in Digital Biology at NVIDIA, and a Visiting Assistant Professor at Duke University, where he leads this lab. Work described here spans both, and every publication is tagged with the affiliation it was carried out under.

The Duke lab is independent of NVIDIA. Joining the Duke group does not imply employment at, or access to the resources of, NVIDIA. For NVIDIA roles, see the NVIDIA careers page; for the Duke lab, see Join.

Join us

We supervise Duke students, and collaborate with people elsewhere through jointly-supervised theses and fellowships. The constraints are stated plainly on the Join page, up front, before anything else.