Editorial: Advances in protein structure, function, and design

Accessibility to ever improving computing infrastructure has led to a paradigm shift towards data-driven modeling in all areas of science and arts. Eponymously, data-driven modeling relies on 1) well curated, domain-knowledge-driven datasets, and 2) appropriate utilization of said data (i.e., avoid overfitting, under sampling, etc.). The domain of protein biology has historically been on the lookout for a reliable method to discern the 3D-shape (structure) of a protein given its amino acid sequence. Precise knowledge of a protein’s structure enables us to first, explain how it works as a tiny molecular machine, and then devise rules to modify existing proteins or design new ones for therapeutic and engineering applications spanning–healthcare, green chemistry, energy, and novel functional materials. One way to accurately determine the 3D-shape of a protein is via experiments (spectroscopy–Nuclear Magnetic Resonance, or crystallography–X-Ray diffraction) to catalog the 3D-Cartesian coordinates of each atom that are present in the protein. Such per-atom information is stored in a PDB (Protein Data Bank) format. Set up in 1976, the PDB (Berman et al., 2000) is a publicly accessible dataset of ~198 k protein structures and has aggressively expanded at the rate of ~11 k new entries per year since 2013. While this is quite a substantial dataset, this barely scratches the surface and constitutes only a meagre ~.09% of the total set of 230M known protein sequences reported till date (UniParc dataset (Bairoch et al., 2005)). This has prompted the emergence of a gamut of data driven deep-learning techniques to reason over known sequence-structure pairs (from PDB) and create neural operators which can then predict the structure from any new protein sequence. Two emerging, yet different schools of thought that fuel these deep-learning pipelines for structure prediction are: 1) family sequence alignment-based (FSA), and 2) single sequencebased (SS). FSA methods such as AlphaFold2 (Jumper et al., 2021) and RosettaFold (Baek et al., 2021) group all sequences a protein family (say, all amylases across species) and corresponding structures into constellations of similarity. During prediction, each input sequence is first sent through a sequence alignment pipeline to find which constellation it belongs to and use structures from the same constellation as templates to thread a possible predicted structure. Such methods, while powerful, cannot account for significant structural changes from point mutations unless such a mutant is a part of the training set (in which case it simply memorizes it). Interestingly, designed proteins with tailored function and disease-causing protein sequence variants fold into very different structures. Structural changes in these proteins are elusive to FSA structure predictors like AlphaFold2 and RosettaFold. On the other hand, SS methods (like RGN2 (Chowdhury et al., 2022)) use natural language processing to encode sequences to highdimensional vectors and map such encodings to atomic coordinates of one (C α) or more atoms OPEN ACCESS

Paper

Full text

PDF

Editorial: Advances in protein structure, function, and design

Semantic Scholar · Medicine · 2023

Abstract

Accessibility to ever improving computing infrastructure has led to a paradigm shift towards data-driven modeling in all areas of science and arts. Eponymously, data-driven modeling relies on 1) well curated, domain-knowledge-driven datasets, and 2) appropriate utilization of said data (i.e., avoid overfitting, under sampling, etc.). The domain of protein biology has historically been on the lookout for a reliable method to discern the 3D-shape (structure) of a protein given its amino acid sequence. Precise knowledge of a protein’s structure enables us to first, explain how it works as a tiny molecular machine, and then devise rules to modify existing proteins or design new ones for therapeutic and engineering applications spanning–healthcare, green chemistry, energy, and novel functional materials. One way to accurately determine the 3D-shape of a protein is via experiments (spectroscopy–Nuclear Magnetic Resonance, or crystallography–X-Ray diffraction) to catalog the 3D-Cartesian coordinates of each atom that are present in the protein. Such per-atom information is stored in a PDB (Protein Data Bank) format. Set up in 1976, the PDB (Berman et al., 2000) is a publicly accessible dataset of ~198 k protein structures and has aggressively expanded at the rate of ~11 k new entries per year since 2013. While this is quite a substantial dataset, this barely scratches the surface and constitutes only a meagre ~.09% of the total set of 230M known protein sequences reported till date (UniParc dataset (Bairoch et al., 2005)). This has prompted the emergence of a gamut of data driven deep-learning techniques to reason over known sequence-structure pairs (from PDB) and create neural operators which can then predict the structure from any new protein sequence. Two emerging, yet different schools of thought that fuel these deep-learning pipelines for structure prediction are: 1) family sequence alignment-based (FSA), and 2) single sequencebased (SS). FSA methods such as AlphaFold2 (Jumper et al., 2021) and RosettaFold (Baek et al., 2021) group all sequences a protein family (say, all amylases across species) and corresponding structures into constellations of similarity. During prediction, each input sequence is first sent through a sequence alignment pipeline to find which constellation it belongs to and use structures from the same constellation as templates to thread a possible predicted structure. Such methods, while powerful, cannot account for significant structural changes from point mutations unless such a mutant is a part of the training set (in which case it simply memorizes it). Interestingly, designed proteins with tailored function and disease-causing protein sequence variants fold into very different structures. Structural changes in these proteins are elusive to FSA structure predictors like AlphaFold2 and RosettaFold. On the other hand, SS methods (like RGN2 (Chowdhury et al., 2022)) use natural language processing to encode sequences to highdimensional vectors and map such encodings to atomic coordinates of one (C α) or more atoms OPEN ACCESS

Similar papers

© 2026 NYSGPT2525 LLC