PottsMPNN AI framework from MIT frees protein engineering from the constraints of natural sequences
Listen to this article
Read by Anchor
For many years, computational protein design has been governed by a conventional metric that measures model success by how well it reproduces sequences selected by natural evolution over millions of years. However, a new study published in the Proceedings of the National Academy of Sciences (PNAS) by researchers at the Massachusetts Institute of Technology proposes an entirely different approach using a machine learning framework called PottsMPNN, designed to raise success rates for de novo proteins that bear no resemblance to sequences found in nature.
Protein design typically follows a two-step process: it begins with specifying the three-dimensional geometric structure required to perform a specific biological function, such as binding to a pathogen molecule inside a cell, after which a machine learning model is tasked with generating amino acid sequences capable of folding into that target shape. The challenge lies in the fact that radically different amino acid sequences can fold into the same geometry, while a single sequence can adopt different conformations depending on flexibility or functional triggers, making the confinement of artificial intelligence to mimicking natural sequences an obstacle to creating entirely new structures.
The PottsMPNN framework directs generative algorithms toward understanding the physical energy landscape of a structure rather than merely replicating inherited biological sequences.
Professor Amy Keating, head of the Department of Biology at MIT and senior author of the study, explained that evaluating model success through natural sequence recovery is no longer the optimal metric for protein engineering. Developed by researcher Foster Birnbaum, the new framework integrates the physical principles governing protein structural stability, providing the model with a more precise understanding of the relationship between the identity of each amino acid and overall protein stability, while enabling it to accurately predict the impact of mutations on that stability.
To achieve this, the team used geometric noise injection, introducing subtle, intentional variations into the structure during training to reduce the model's tendency to overfit natural sequences and expand the structural diversity it can handle. The model also used pairwise distributions that accurately calculate the physical interactions across all 20 possible amino acid options at each pair of positions, alongside training on sets of evolutionarily related sequences to teach it how distinct sequences can adopt identical structures.
This engineering shift opens immediate practical opportunities for research and development teams, biotechnology laboratories, and pharmaceutical industries in the Gulf, Egypt, and the Levant, shifting workflows from relying on pre-existing biological libraries to custom computational design from scratch. This approach lowers wet lab experimental costs and accelerates the development of targeted therapies and molecular inhibitors for intractable diseases without waiting for the discovery of natural counterparts, strengthening local researchers' independence in developing therapeutics and biosensors tailored to regional healthcare needs.
The researchers note that while the most widely used model in the field since 2022 had remained unmatched for a long time, PottsMPNN demonstrates that reducing reliance on natural sequences improves structural compatibility and energetic prediction. This paves the way for fine-tuning and customizing the model for precise therapeutic tasks, as well as creating entirely new functional proteins unknown to nature.