Representation learning for biology
Teaching models to read proteins and genomes as languages.
The Dallago lab at Duke University.
We ask what machine-learned representations can reveal about biology — and where they fail — then use those answers to build more reliable tools for discovery and protein engineering.
Our work spans scalable biological inference, rigorous evaluation of predictive models, and generative design of proteins and molecular interactions. We pair large-scale computation with careful benchmarking and experimental collaboration to identify what models can reliably tell us about biology.
Teaching models to read proteins and genomes as languages.
Writing new proteins, atom by atom, and testing them at the bench.
Making biological search and structure prediction fast enough for whole proteomes.
Measuring whether these models actually work where it matters.
Preprint 2026 · NVIDIA + Duke
Most proteins do not work alone, but the world's structure database showed them alone. We predicted 31 million candidate complexes across 4,777 proteomes, established confidence criteria against experimentally determined structures, and released 1.81 million high-confidence predictions. 31.3% of them extend beyond anything detectable in the PDB.
ICLR 2026 Oral · NVIDIA
Binder design has been split between conditional generation and hallucination. We argue that is a false dichotomy, and built one fully atomistic model that does both — then showed it keeps improving when you let it think longer at inference time.
Nature Methods 2025 · NVIDIA
The bottleneck in structure prediction was never only the network — it was the search in front of it. Redesigning the core alignment algorithms for GPUs reached roughly 100 trillion cell updates per second, and kept it working on low-wattage cards so the speed-up is not reserved for people with a cluster.
ICML 2026 Oral · NVIDIA + Duke
Seven new fitness datasets and splits built to mirror real protein-engineering campaigns rather than random train/test partitions. Across them, simpler models often matched or beat fine-tuned protein language models — a result worth publishing precisely because it is inconvenient.
Chris is a Senior Research Scientist and Applied Research Team Lead in Digital Biology at NVIDIA, and a Visiting Assistant Professor at Duke University, where he leads this lab. Work described here spans both, and every publication is tagged with the affiliation it was carried out under.
The Duke lab is independent of NVIDIA. Joining the Duke group does not imply employment at, or access to the resources of, NVIDIA. For NVIDIA roles, see the NVIDIA careers page; for the Duke lab, see Join.
We supervise Duke students, and collaborate with people elsewhere through jointly-supervised theses and fellowships. The constraints are stated plainly on the Join page, up front, before anything else.