See discussions, stats, and author profiles for this publication at: [Link]
net/publication/384237748
Structural Bioinformatics and Protein Structure Prediction
Chapter · September 2024
DOI: 10.1007/978-981-97-7123-3_8
CITATIONS READS
0 108
2 authors:
Kavita Patel Ashutosh Mani
Motilal Nehru National Institute of Technology Motilal Nehru National Institute of Technology
7 PUBLICATIONS 10 CITATIONS 96 PUBLICATIONS 683 CITATIONS
SEE PROFILE SEE PROFILE
All content following this page was uploaded by Kavita Patel on 18 June 2025.
The user has requested enhancement of the downloaded file.
eProofing 15/08/24, 7:57 PM
Query Details Back to Main Page
1. As keywords are mandatory for this chapter, please provide 3–6 keywords.
Template based modelling, Template free modelling, Homology modelling, Ab-initio modelling, Threading
2. Reference [88] is given in list but not cited in text. Please cite in text or delete from list.
deleted the reference.
Structural Bioinformatics and Protein Structure Prediction
Kavita Patel Affiliationids : Aff1
Ashutosh Mani ✉
Email : amani@[Link]
Affiliationids : Aff1, Correspondingaffiliationid : Aff1
Aff1 Department of Biotechnology, Motilal Nehru National Institute of Technology, Prayagraj, Uttar Pradesh, 211004, India
Abstract
Structural bioinformatics is a rapidly growing field and is essential in understanding the three-dimensional structure of
biological macromolecules like proteins. This chapter thoroughly introduces structural bioinformatics and how it is used
to predict protein structures. The chapter begins with an introduction to the fundamentals of protein structure
determination, covering experimental methods like nuclear magnetic resonance (NMR) spectroscopy, X-ray
crystallography, and cryo-electron microscopy (cryo-EM). It then explores the difficulties of using experimental methods
and the necessity of computational approaches in protein structure prediction, like homology modeling, ab initio model,
and threading. The focus is on utilizing databases, machine learning techniques, and bioinformatics tools in concert to
improve prediction accuracy. We look at new developments, including the prediction of protein–protein interactions and
the effects of genetic variants on the structure and function of proteins. We also address future directions, emphasizing
the multidisciplinary nature of the discipline and its implications for drug discovery and personalized medicine. These
include the development of novel computational tools and the integration of multi-omics data.
Keywords
∎∎∎
1. Introduction
The dynamic and interdisciplinary area of structural bioinformatics uses concepts from computational biology,
bioinformatics, and biology to analyze and predict the three-dimensional (3D) structures of biological macromolecules, with
a primary focus on proteins [ 1 ]. The basic building blocks of life, proteins are involved in almost every biological function,
including molecular recognition, gene regulation, signal transduction, enzyme catalysis, and signal transduction. Clarifying
the roles, relationships, and methods of action of proteins requires an understanding of their three-dimensional structures [
2 ] AQ1 .
The combination of computational approaches for modeling and predicting protein structures with experimental methods for
determining protein structures gave rise to the discipline of structural bioinformatics. Atomic resolution protein structure
visualization and characterization have been made possible by the development of experimental methods, including cryo-
electron microscopy (cryo-EM) [ 3 ], nuclear magnetic resonance (NMR) spectroscopy [ 4 ], and X-ray crystallography [ 5 ].
[Link] jCXRi4lLmjSpYsSElbKJa0Q== Page 1 of 15
eProofing 15/08/24, 7:57 PM
Proteins are crystallized in X-ray crystallography, and the electron density and three-dimensional structure of the protein are
inferred by examining the diffraction patterns of X-rays going through the crystals. In order to ascertain a protein’s structure
in solution, NMR spectroscopy analyses the interactions between atomic nuclei in the protein. Protein structures are
visualized using electron microscopy in cryo-EM, which determines high-resolution structures of massive macromolecular
complexes by freezing the proteins in vitreous ice.
However, protein structure identification by experimental approaches is frequently labor-intensive, time-consuming, and not
always possible for all proteins, particularly membrane-bound or multi-domain proteins. Furthermore, due to protein flexibility
and conformational changes, experimental methods may not always provide high-resolution structures or may yield
inconclusive results. Because of this, computational tools [ 6, 7, 8, 9 ] are becoming essential for predicting protein structures
in situations when experimental approaches are impractical or yield conflicting results.
There are two main approaches for the structure prediction of the proteins, first one is template based modelling and second
one is template free modelling (Fig. 1 ). Template based modelling methods constructs model by assessing and identifying
the structural frameworks of existing proteins from PDB [ 10 ] which is called homologous templates while template free
modelling methods predict protein without using global template structures. Template based modelling methods accuracy
depends upon the quality of alignments and sequence identities between target protein and template, which is also
dependent on the evolutionary distances between template and query. For proteins with sequence identities >30–50%, the
models have ~85% of the core regions to the native structure, however, the modelling accuracy also decreases as the
sequence identities drops <30% due to lack of significant templates [ 11, 12 ].
Fig. 1
Decision-making diagram for the process of protein structure prediction
Template free modelling methods have been used traditionally to model those proteins which have no homologous templates
identified from the PDB. Template free modelling method based on the knowledge based and physics based energy
functions and use extensive sampling methods to model protein structure, referred as de novo modelling or ab initio
modelling approaches [ 13, 14 ]. This method has not been used historically as its accuracy is low as compare to template
based modelling but recently, the gap between these two approaches has been filled with the use of deep learning to build
protein structure models [ 15 ].
The field of structure prediction has gained significant advancements recently due to the availability of high-throughput
experimental data, advancement in computational algorithms and methods, and integration of resources and bioinformatics
tools. Structural bioinformatics advances biological research and the development of new therapeutics by clarifying protein
structures and providing insights into their relationships, roles, and activities in both health and disease.
[Link] jCXRi4lLmjSpYsSElbKJa0Q== Page 2 of 15
eProofing 15/08/24, 7:57 PM
2. Template Based Modelling
Template based modelling or comparative modelling is a computational approach used for the protein structure prediction
from the known structures of homologous protein, which is referred as templates used to predict the target protein structure
[ 16 ]. The principle underly in template based modelling is that it assume that proteins which have similar sequences share
similar functions and structures. Template based modelling method (Table 1 ) aims to transfer the structural information from
the templates to target protein by aligning the sequences of target protein with known protein structure sequences, which
leads to 3D structure prediction.
Table 1
Template based modelling methods
Name Method Description References
A integrate interface for: 3D
modelling/tertiary structure prediction,
Automated Web Server
IntFOLD quality assessment of 3D model, prediction [ 17 ]
([Link]
of intrinsic disorder, domain prediction,
prediction of protein ligand binding residues
Advanced remote template detection and
Automated Web Server, Automated updated Fold library, Genome
Phyre and build 3D models, ligand binding site [ 18, 19, 20
search ([Link]
Phyre2 detection, analysis of effect of SNPs (amino ]
id=index)
acid variants)
Template search, template- target
Automated Web Server
ESyPred3D alignment, 3d-model building, model [ 21, 22 ]
([Link]
evaluation
A integrated method for: Total folding
energy calculations, build model, repair Command line tool and downloadable program
FoldX [ 23 ]
PDB, analyse complex, Print Networks, and ([Link]
stability of protein
Protein structure prediction, analysis of
molecular dynamics trajectories, plateform
Python platform for structural bioinformatics
Biskit for integrated programs and algorithms [ 24 ]
([Link]
DSSP, T-Coffee, Fold-X, TM-Align and
MODELLER
[ 25, 26, 27
RaptorX Template based protein structure modelling Automated Web Server ([Link]
]
Template search, target template alignment, Software package based on Python and Fortran
MODELLER [ 28 ]
comparative modelling, model evaluation ([Link]
Detection of templates, alignment, 3D Interactive Web Server
HHpred [ 29 ]
modelling ([Link]
Protein modelling, and analysis, ligand
ROSETTA Software program ([Link] [ 30 ]
docking
Template detection, alignment, ligand and
Yasara oligomers modelling, model fragment (Software program [Link] [ 31 ]
hybridizations. simulation
MOE
(Molecular Template search, model build, visualization,
Software program ([Link] [ 32 ]
Operating simulation
Environment)
SWISS- Template search, user defined target-
Automated Web Server ([Link] [ 33 ]
MODEL template alignment, 3D structure prediction
BHAGEERATH- Combination of homology methods and ab Automated Web Server ([Link]
[ 34, 35 ]
H initio folding [Link]/bhageerath/bhageerath_h.jsp)
2.1. Homology Modelling
Homology modelling is used for 3D structure predictions of proteins based on its amino acid sequences and the known
structure of homologous proteins. This method is based on the hypothesis for the proteins which shares high sequence
homology, also will be structurally similar. So, when the sequence similarity is high ≥30% between target and templates, this
method is used. This makes it remarkably easier to find the high similarity templates by aligning them. The identification of
structurally conserved regions among the structures corresponding to templates is a crucial step in homology modelling, it
determine by calculating the C-α distance matrix of each structure and small sections of this matrix get compared to find
the small peptide with low RMSD (root mean square deviations) of related structures [ 36, 37 ].
[Link] jCXRi4lLmjSpYsSElbKJa0Q== Page 3 of 15
eProofing 15/08/24, 7:57 PM
Simple sequence-sequence alignment algorithm like Smith–Waterman algorithm [ 38 ] (for local alignment) and
Needleman–Wunsch algorithm [ 39 ] (for global alignment) is used to find homologous sequence or conserved regions but
its relatively slow dynamic programming methods, BLAST [ 40 ] software is used for rapid sequence-sequence alignment.
The target protein 3D structure is built by transferring the atom coordinates from the corresponding regions of template
structures. The target sequence is superimposed onto the template structure during this procedure, maintaining the
structural features that are conserved.
Template-target alignment shows that some blocks in the target match the template’s structurally conserved regions. Gaps
appear between the aligned model and the template sequence during template-target alignment. It is necessary to identify
and incorporate the structural fold of the gap residues, or loop, between the two conserved core regions. This require
modifications of backbone, orientation or conformational changes are absent from the normal secondary structural
elements. As a result, it is safe to make all of the insertions and deletions inside the alignment form loops, turns, and helixes.
Two approaches mainly used for loop modelling, knowledge based and energy based approach. In knowledge based
approach, it look for a loop region’s structure with endpoints that resemble known structures, and then we sandwich the
loop structure’s coordinates between two cores. Mostly molecular modelling programs like Swiss model, MODELLER [ 41 ],
insight support knowledge based approach for loop modelling. In energy based approach, molecular dynamics or Monte
Carlo simulation techniques [ 42 ] are used to produce the loop form most accurately, and the energy function is utilised to
evaluate the loop quality. It is possible to alter the energy function to produce a loop structure that fits the core more snugly
[ 43, 44 ].
High-precision side-chain rotamer prediction, mainly depends on the rotamers and their packaging, requires the correct
backbone [ 45 ]. The primary method for handling a tangle of this type involves repeatedly predicting rotamers, followed by
backbone shifts and new backbone rotamers, until the process converges. By using this technique, the number of rotamer
predictions and energy minimization stages are decreased. These approaches should be used for the whole protein
structure since they are applicable to both loop modelling and model optimization [ 46, 47 ].
Two approaches (1) Quantum force fields [ 48 ] (2) Self- Parameterizing force fields [ 49 ], are used for the better
optimization of a model which can be achieved by calculating force field. Recent developments in computational biology
have made it possible to apply quantum chemical techniques to provide a more precise interpretation of the charge
distribution across the protein molecule. Initialise the force field with certain parameters, adjust it, reduce model energy, and
store the updated force field if model quality improved; otherwise, revert to the original force field value. By using this
method, the force field during energy minimization can be more accurately directed in the desired direction.
Molecular dynamics simulation can be used to optimise protein models. It generates the true folding dynamics of the
protein by sampling the trajectory of the protein's motions over a 10-fs time span [ 50 ]. As a result, it is anticipated that as
the simulation runs, the model will resemble real structure [ 51 ].
2.2. Threading
Threading also known as fold recognition method, is a computational method used in tertiary structure prediction of
proteins by identifying the most similar fold or structural arrangement for target protein sequence. This method predicts the
3D structure when the similarity between target and template sequence is very low which is ≤ 30% [ 52 ]. Understanding
the link between sequences, structure, and function is still incomplete. Understanding the link between sequences,
structure, and function is still incomplete. This method begins with the generation of structural template or protein folds [
53 ] by identifying the distantly homologous templates [ 54 ].
The hypothesis behind finding the 3D structure by aligning the sequences by PSI-BLAST [ 55 ] which is extension of BLAST
methodology. PSI BLAST perform multiple sequence alignment of the sequences which it get by BLAST alignment. The
alignment is subsequently transformed into a position-specific score matrix (PSSM), which records the amino acid trends at
every MSA position. An method similar to BLAST is used to iteratively search through a sequence database a predefined
number of times using the PSSM rather than the query sequence. The profile, or PSSM, is changed to reflect the sequences
found in the previous round following each step. In order to find more distantly related proteins, PSI-BLAST’s methodology
involves iteratively searching a database utilising profiles, which describe more about the sequence space compatible with
a given protein fold. A template structure is expressed as a descriptor string that characterises the structural environment in
[Link] jCXRi4lLmjSpYsSElbKJa0Q== Page 4 of 15
eProofing 15/08/24, 7:57 PM
the 3D profile approach. Three primary environment classifications exist: (1) the portion of the lateral chains covered by
polar atoms, (2) the area of the lateral chains buried by other protein atoms, and (3) secondary local structure. The
environmental class of each residue in the folded protein structure is represented by an ID string in this instance, which
describes a 3D protein structure. Here, the profile of a template structure is generated using the 3D-1D score table. This
technique is also referred to as sequence-structure since it represents both the target and the template as strings. The
template structure is represented as a string of environmental classes, while the target protein is represented as a string of
amino acids. Using a dynamic programming approach, the fit score between the template and target environment classes is
determined.
Profile Hidden Markov Models (HMMs) can also be used to represent sequence profiles in addition to PSSMs. Utilising
profile HMMs has the benefit of applying substitution probabilities and position-specific gap penalties, which more
accurately reflect the underlying sequence distribution [ 56, 57 ].
TASSER is used to generate 3D model of distantly related homology protein target, which extracts generally the contiguous
fragments from the aligned regions of threading. And for the unaligned regions, it uses a lattice-based template-free
modelling approach. Aside from using constraints from the aligned region, TASSER also uses different knowledge-
based energy functions, which is important for protein folding (e.g., secondary structure formation, hydrogen bonding,
side-chain contact formation) Monte Carlo Simulations and produce large number of reduced models. Based on structural
similarity, low-energy decoys are grouped, and the biggest cluster centroid is chosen for further full-atom refining. The
lowest energy structure is then chosen and put through full-atom refining.
3. Template Free Modelling
This computational approach used in protein structure prediction which does not rely on the known structures of
homologous proteins as templates. As an alternative, TFM uses only the target protein's amino acid sequence to predict
protein structures de novo. This method is especially useful for predicting novel folds or in the absence of closely related
homologous proteins with known structures. There are following different techniques (Table 2 ) used in this approach.
Table 2
Template free modelling methods
Name Method Description References
Orientation-guided folding/deep learning distance,
D-QUARK Automated web server ([Link] [ 58 ]
method for ab-initio protein structure predictions
Hydrogen bond network-guided
D-I-
folding/orientation/template and deep learning Automated web server ([Link] [ 59 ]
TASSER
distance/
Automated web server
Robetta Server
Rosetta Orientation-guided folding/Deep learning distance/ ([Link] [ 30 ]
trRosetta Server:
([Link]
End-to-end deep learning-based protein model Automated web server
AlphaFold2 [ 60 ]
prediction ([Link]
Based on fragment assembly to resolve 3D
FragFold [Link] [ 53 ]
structure prediction
Downloadable program
PSICOV Distance based deep learning model prediction [ 61 ]
([Link]
GREMLIN Distance based deep learning model prediction Automated web server ([Link] [ 62, 63 ]
Open source software
CCMpred Based on Markov model contact based prediction [ 64, 65 ]
([Link]
Protein prediction on contact map prediction using Online server and standalone package
NeBcon [ 66 ]
neural netwrok ([Link]
Protein prediction on contact map prediction using Online server
ResPRE [ 67 ]
coupling precision matrix ([Link]
Automated web server
DMPfold Deep learning-based prediction [ 68 ]
([Link]
RaptorX-
Distance-based deep learning model prediction ([Link] [ 25, 27 ]
Contact
[Link] jCXRi4lLmjSpYsSElbKJa0Q== Page 5 of 15
eProofing 15/08/24, 7:57 PM
TripletRes Distance-based deep learning model prediction Automated web server [ 69 ]
([Link]
3.1. Ab Initio Methods
This method creates a protein model based only on sequence data because structural folds or their counterparts are not
readily available. The physicochemical concept pertaining to the nature of proteins can be comprehended through the ab
initio method. Compared to other approaches of structure prediction, ab initio modelling has a low accuracy [ 70 ]. Protein
structure can be constructed by figuring out the configurational space of the atoms in amino acids if the target sequence is
not having any structural similarities with structures present in the database. This approach makes use of an understanding
of several physics, chemistry, and mathematics concepts. The computation is simplified by using reduced protein
representations. In certain models, a residue can represent only two locations like backbone and side chain [ 71 ]. Some
employ several sites, such as a side link and heavy backbone atoms. Hydrophobic interactions are recognised to be the
primary factor in protein folding, and an empirical energy function may be used to calculate these interactions. Three
criteria need to be determined in order to make the ab initio prediction: (1) proteins representation reduction; (2)
interaction's potential energy function; and (3) a method to explore the conformational space.
Simulated annealing is used to search the configuration space of fragment structures. A step is performed to substitute the
present configuration with the torsion angles of a neighbour chosen at random at a randomly chosen position. Motions that
bring two atoms closer together by 2.5 Å are eliminated, while other motions are assessed. Since most structure prediction
algorithms now in use rely on information gained from experimentally predicted structures, they are not very helpful in
delving into the fundamentals of protein folding. Template-free approaches take into account both the fundamentals of
protein folding and their practical application. The development of template-free approaches may more accurately
represent the prediction of the technical and theoretical level of protein structure than template-based methods since they
are based on information from existing structures. ROSETTA A template-free method developed by the David Baker Lab
that assembles a full structure based on fragments of 3–9 residues from PDB is one of the effective ab initio modelling
approaches [ 72, 73 ]. The fragments are chosen according to the degree of similarity between the predicted and known
secondary structure, much like in template-based approaches. The annealing search strategy used in the Monte Carlo
approach simulates the assembly process. The fragments utilised in QUARK range in size from one to twenty residues, and
Monte Carlo replica-exchange simulation is employed to simulate the assembly process while following the guidance of a
knowledge-based atom-level force field. Fragment assembly is also the foundation of many other techniques, such as
Scrape, PROFESY, FRAGFOLD, etc. The primary difference between these methods and the template-based methods is that
the first do not rely on any global structural blueprint, while the latter do not take advantage of structural similarities or
homology between the target and the proteins that the fragments originate from. For template-free approaches, it is better
at modelling the target of new folds. However, modelling proteins with a length of more than 150 residues remains a
significant difficulty for template-free approaches due to the high computational need and low force field accuracy.
Recently, contact map prediction using a co-evolution method has demonstrated success in breaking through this length
restriction of ab initio structure folding.
3.2. Protein Structure Prediction by Physics Based Energy Minimization
Physics based energy minimization is a computational approach to predict the 3D structure of the proteins by using the
principle of physics and chemistry. This approach describes the interactions between atoms in a protein molecule using
mathematical models called force fields. The protein’s native or physiologically active conformation may be found by
minimising the potential energy of the protein structure, which is the most stable conformation. The choice of a suitable
force field is the first stage in the energy minimization method of protein structure prediction. Protein molecules’ atomic
interactions’ strength and geometry are described by characteristics found in force fields. In order to replicate experimental
observables, these parameters are obtained through the use of empirical fitting, quantum mechanical calculations, and
experimental data. A force field is a mathematical model that uses atom-to-atom interactions to represent the potential
energy of a protein structure. CHARMM [ 74, 75, 76 ], AMBER [ 77, 78, 79 ], and GROMOS [ 80 ] are common force fields used
in protein structure prediction. The process of creating the protein’s first structure usually involves the use of techniques
like threading, homology modelling, or fragment-based assembly, this serves as a starting point for energy minimization.
Using the force field that has been chosen, the potential energy of the protein structure is computed. Bond stretching,
dihedral angle rotation, angle bending, electrostatic interactions and van der Waals interactions are some of the energy
factors that affect the potential energy [ 9 ]. A multidimensional surface with each dimension representing a degree of
[Link] jCXRi4lLmjSpYsSElbKJa0Q== Page 6 of 15
eProofing 15/08/24, 7:57 PM
freedom in the protein conformation can be used to represent the potential energy of a protein structure. The goal of energy
minimization is to locate the global minimum, or the most stable protein conformation, on this potential energy surface.
Molecular dynamics simulations, conjugate gradient, and steepest descent are common optimization algorithms for energy
minimization. By sampling various conformations and calculating their potential energies, energy minimization investigates
the protein’s conformational space. Until a termination condition is satisfied, such as hitting a predetermined energy
threshold or having the optimisation algorithm converge, the energy minimization process keeps going. The final protein
structure is acquired when optimisation is finished [ 81 ].
3.3. Protein Structure Prediction by Fragment Assembly Method
Fragment Assembly (FA) techniques for protein structure prediction are computational approaches that assemble small
peptide fragments into a whole 3D protein structure [ 82 ]. Fragment Assembly techniques create structures de novo, based
only on the amino acid sequence of the target protein, in contrast to Template-Based Modelling (TBM), which uses
templates of known structures of homologous proteins. FA techniques allow the assembly of these fragments into a global
protein structure by taking use of the fact that local structural motifs, such secondary structure elements, are frequently
conserved among proteins.
In Fragment Assembly, a library of short peptide fragments generated from known protein structures is created as the initial
step. These segments usually consist of three to nine amino acids and exhibit local structural motifs such as turns, β-
strands, and α-helices. Fragment libraries can be produced using computational techniques like ab initio structure
prediction or using experimental structures found in databases like the Protein Data Bank (PDB). Proteins often possess
local structural motifs like α-helices, β-strands, and turns, which is why fragment assembly methods take use of this
characteristic. Global protein structures can be produced by the assembly of small peptide fragments that represent these
patterns. Based on projected secondary structure and sequence similarity, appropriate fragments are chosen from the
fragment library for a specified target protein sequence. For assembly, only the fragments that align with the target
protein’s predicted secondary structure and sequence are retained.
Fragment Assembly methods search through the enormous conformational space of protein structures using combinatorial
search strategies. Assembly algorithms find the best possible arrangement of fragments that takes energy and structural
constraints into account by methodically merging and assessing the compatibility of each fragment. The chosen fragments
are combined repeatedly by the assembly algorithm to provide a pool of potential structures. Based on their projected
sequence and secondary structure, fragments are overlapped and aligned. The assemblies that develop are evaluated for
geometric compatibility and energy scores.
The fragment assembly-generated candidate structures are refined to maximise their geometry and eliminate steric
conflicts [ 14, 53 ]. Energy minimization, loop modelling, and molecular dynamics simulations are examples of refinement
approaches which is used to enhance the accuracy and quality of the predicted protein structures. Using geometric and
energy parameters, Fragment Assembly techniques optimise the arrangement of fragments. Protein structure is optimised
by aligning and overlapping fragments to reduce steric conflicts and enhance favourable interactions, such hydrogen bonds
and van der Waals contacts.
3.4. Protein Structure Prediction by Rapid Gradient Descent Based Folding
Methods
One effective template free modelling method is fragment assembly, but the downside is that it depends on the total length
of the protein, the simulations might take hours or even days to complete. Consequently, it is preferable to create
techniques that can quickly produce structures. Gradient descent-based folding techniques can be used to accomplish
this.
A drawback of these methods is that they could be more likely to trap in local minima than to identify global minimum
conformation of the energy distribution. This is especially true in cases like protein folding when the energy landscape is
complicated. Deep learning has recently been used to reliably anticipate pairwise spatial constraints [ 83 ], such as inter-
residue distances, which can smooth the energy landscape and enable gradient-based algorithms [ 84 ] to fold protein
structures appropriately. The initial AlphaFold iteration in CASP13 used a gradient descent-based folding technique to attain
state-of-the-art performance. Moreover, the most recent version of the Rosetta modelling programme, trRosetta, rapidly
folds protein structures using an L-BFGS gradient descent technique and shows that appropriate deep learning-based
[Link] jCXRi4lLmjSpYsSElbKJa0Q== Page 7 of 15
eProofing 15/08/24, 7:57 PM
constraints [ 83 ] may enable high-accuracy predictions even with quick simulations.
3.5. Protein Structure Prediction by Deep Learning
Deep Learning (DL) protein structure prediction is an advanced technique that uses artificial intelligence methods to predict
the 3D structure of proteins based only on their amino acid sequences. Protein structure prediction challenges are suitable
for deep learning models, especially neural networks, which have shown impressive capacities in identifying complex
patterns and correlations in biological data. Preparing training data is the initial stage in Deep Learning’s prediction of
protein structure. A dataset of protein sequences matched to their experimentally established structures must be
assembled; these may be found in sources such as the Protein Data Bank (PDB). The dataset is divide into training, testing,
and validation sets. In order to capture local and global sequence properties pertinent to protein structure, Deep Learning
models develop hierarchical representations of protein sequences [ 85 ]. Deep Learning models automatically learn to
extract meaningful representations by training on large data sets [ 86 ], which helps with accurate structure prediction.
Numerical representations of protein sequences that capture significant sequence properties are called feature vectors.
Embedding methods, in which amino acids are mapped to continuous vector spaces, and one-hot encoding, in which each
amino acid is represented as a binary vector, are common encoding approaches. Deep Learning models can capture the
complex correlations between sequence and structure in proteins because they are very good at recognising patterns in
large, complicated datasets. Deep Learning models may generalise to accurately anticipate the structures of previously
unknown proteins by learning from a variety of examples in the training data. By predicting their structures and deducing
their activities from structural data, deep learning algorithms help with the functional annotation of proteins.
Recurrent neural networks (RNNs), transformer models, convolutional neural networks (CNNs), and other concepts made
up of several layers of linked neurons are common components of Deep Learning models for protein structure prediction [
87 ]. Processing the encoded protein sequences and predicting their matching three-dimensional structures is the
architecture’s intended function. By predicting the structures of designed proteins with certain activities or features, deep
learning-based protein structure prediction aids in the design and engineering of proteins. Stochastic gradient descent
(SGD) or Adam are two optimisation methods that are used to train the Deep Learning model on the provided dataset. By
modifying the neural network's parameters (weights and biases) in response to the difference between the predicted and
real structures, the model gains the ability to map input protein sequences to the relevant 3D structures during training.
Without the need for manual feature creation or intermediary representations, end-to-end learning is made possible by
deep learning models, which allow raw input data (protein sequences) to be directly mapped to output predictions (3D
structures). Complex sequence-structure interactions may be seamlessly integrated with this end-to-end technique, which
also speeds the prediction process. Utilising an independent validation dataset, the trained Deep Learning model’s
prediction accuracy of protein structures is assessed. The global distance test total score (GDT-TS), certain geometric and
energy-based parameters, and the root-mean-square deviation (RMSD) from experimental constructions are examples of
evaluation metrics. Model evaluation aids in determining the projected structures’ accuracy and reliability. The trained Deep
Learning model may be used to forecast the three-dimensional structures of novel protein sequences once it has been
verified. The protein's spatial organisation in three dimensions is represented by the model, which takes as input the amino
acid sequence of a protein and outputs projected coordinates for its atoms.
4. Model Evaluation and Validation
It is necessary to assess the final projected model to ensure that its structural characteristics are in line with the
physicochemical laws. This entails examining irregularities in bond lengths, close to contacts, φ–ψ angles, and other related
factors. Implicitly accounting for these stereochemical features is another technique to assess the quality of modelled
protein. Using statistical profiles of interaction energy and spatial features compiled from structures identified through
experimentation, this technique identifies errors. By comparing the statistical parameters with the developed model, it is able
to ascertain whether sections of a sequence appear to be folded normally and which do not. If structural differences are
discovered, the area is considered to have errors and requires more refinement. A collection of programmes called SAVES
server allows you to verify the model’s correctness by simply uploading the expected structure. One such programme that
can verify general physical properties including bond angles, bond lengths, chirality, and bond angles is called Procheck.
The model’s parameters are compared to those obtained from precisely defined high-resolution structures. The programme
highlights the areas that require more inspection or refinement if it finds any odd traits. Another thorough protein analysis site
that verifies a protein model’s chemical accuracy is WHATIF. It performs a variety of tasks, such as planarity checking, proline
[Link] jCXRi4lLmjSpYsSElbKJa0Q== Page 8 of 15
eProofing 15/08/24, 7:57 PM
puckering, collisions with anomalous bond angles, symmetry axes (close contacts), and bond lengths. Additionally, it permits
the creation of Ramachandran plots to evaluate the model’s quality.
A web server called Atomic Non-Local Environment Assessment (ANOLEA) employs the statistical evaluation method. It
computes the energy of atomic interactions inside a protein chain and contrasts the resultant interaction energy values with
data gathered from an X-ray structural database of proteins. It may be an indication that the appropriate area has not been
accurately modelled if the energy terms of some regions differ notably from those of the typical crystal structures. Typically,
5.0 is the threshold for unfavourable residues. Regions containing mistakes are defined as residues with scores more than
5.0.
Another server that use the statistical method is Verify3D. It makes use of a pre-calculated database made up of 18
environmental profiles derived from high-resolution protein structures that are based on solvent exposure and secondary
structures. The solvent exposure and secondary structure propensity of each residue are computed to evaluate the quality of
a modelled protein. The solvent exposure and secondary structure propensity of each residue are computed to evaluate the
quality of a protein model. A residue is given either a high or low score depending on whether its properties fit into one of the
profiles. The outcome is a two-dimensional graph that shows how well each protein structural residue folds. Normally, the
threshold value is zero. A residue is said to have an unfavourable environment if its score is less than zero.
Using various verification programmes may yield different evaluation findings. Verify3D determines that the protein’s C-
terminus has low-quality residues, despite ANOLEA declaring the model’s full-length protein chain to be favourable. It’s a
good idea to employ a variety of verification methods and determine which ones are in agreement with one another because
there isn’t a single approach that is obviously better than the others. It’s also critical to remember that the evaluation tests
carried out by these programmes only assess the stereochemical correctness; they do not assess the model’s accuracy,
which may or may not have any bearing on biology. Based on the Ramachandran plot of amino acid residues, several
methods forecast the correctness of the anticipated model. It is a 2D scatter plot that displays the torsion angles of every
residue of an amino acid in a protein. Based on established protein structures, the map indicates both permitted and
prohibited areas of the angles. The quality assessment of a novel protein model is aided by this plot.
5. Applications
A vast quantity of data is being integrated more easily thanks to the availability of protein 3D (3-dimensional structure)
structures and other structural analysis tools. This information may be valuable to investigate other avenues for
strengthening our understanding of protein function and structure in the future. A protein’s 3D structure offers greater
information about its binding site and other functionally significant areas, information that may be used in the development of
new drugs. To begin the process of creating drugs based on structure, a target structure must be available. This structure
also directs the modifications that lead molecules undergo. Better knowledge about the residues involved in the interaction
can be obtained from a protein’s 3D complex structure with a ligand. The mechanism of a drug’s pharmacological activity,
binding affinity, and lead modification is explained by the interaction of a receptor with ligand (small molecule). A protein's
natural structure, which is necessary for the protein to function normally, is destroyed when an amino acid is mutated, which
is another way that computational modelling explains why this occurs. Through the representation of structural alterations in
the mutant target protein that result in a loss of appropriate drug binding or interaction, it can also help explain the
mechanism of drug resistance. To maintain stability, a variety of forces, including Van der Waals, hydrophobic interaction,
electrostatic interaction, and hydrogen bonding interaction between protein–ligand complexes. The process of modelling the
intermolecular linkages within the protein–ligand complex is complicated since there are several degrees of freedom and little
data regarding the influence of water on binding.
6. Conclusion
Molecular modelling has emerged as a major and fundamental method for the experts working in the drug creation field.
Molecular modelling helps to understand the relevant physicochemical characteristics of proteins by revealing their three-
dimensional structures. The protein modelling effectively combines theoretical scientific concepts, computational science
techniques, and experimental data to reveal a macromolecule’s structural and biological characteristics. The most effective
technology for prediction of protein structure depend on the kind of problem that has to be solved. The fundamental
strategies and the most current developments in protein structure prediction techniques have been covered in this chapter. A
necessary emphasis has also been placed on the kinds of mistakes that might occur and compound while working with
protein models. The accuracy of predicted protein structure is a crucial step as structure-based drug design depends on
[Link] jCXRi4lLmjSpYsSElbKJa0Q== Page 9 of 15
eProofing 15/08/24, 7:57 PM
protein structure AQ2 . Since X-ray crystallography and NMR spectroscopy are time-consuming methods that don’t seem to
be viable for determining the structure of each particular protein, the creation of a highly accurate and exact modelling tools
need to be the evolved.
Acknowledgements
The authors acknowledge the Department of Biotechnology, MNNIT, Allahabad, for supporting this study smoothly. All
authors have edited, conceptualised, and helped write the manuscript.
Conflict of Interest
“The authors declare that they have no conflicts of interest in this work.”
References
1. Kumar, A., & Chordia, N. (2017). Role of Bioinformatics in Biotechnology. Resaerch and Reviews in BioSciences, 12(1). Ret
rieved May 6, 2024, from [Link] .
2. Khan, R. H., Siddiqi, M. K., & Salahuddin, P. (2017). Protein structure and function.
3. Cheng, Y. (2015). Single-particle cryo-EM at crystallographic resolution. Cell, 161(3), 450–457. [Link]
ell.2015.03.049
4. Cavanagh, J. (1996). Protein NMR spectroscopy: Principles and practice. Academic Press.
5. Methods of Biochemical Analysis. Retrieved May 6, 2024, from [Link] , [Link]
0.1002/9780470110584#page=15 .
6. A method to identify protein sequences that fold into a known three-dimensional structure. Science. Retrieved May 6, 20
24, from [Link] , [Link] .
7. Folding of polypeptide chains in proteins: A proposed mechanism for folding. PNAS. Retrieved May 6, 2024, from https://
[Link]/doi/abs/, [Link] .
8. Levitt, M., & Warshel, A. (1975). Computer simulation of protein folding. Nature, 253(5494), 694–698. [Link]
038/253694a0
9. McCammon, J. A., Gelin, B. R., & Karplus, M. (1977). Dynamics of folded proteins. Nature, 267(5612), 585–590. [Link]
[Link]/10.1038/267585a0
10. The protein structure prediction problem could be solved using the current PDB library. PNAS. Retrieved May 6, 2024, fr
om [Link] , [Link] .
11. Kryshtafovych, A., Monastyrskyy, B., Fidelis, K., Moult, J., Schwede, T., & Tramontano, A. (2018). Evaluation of the templ
ate-based modeling in CASP12. Proteins: Structure, Function, and Bioinformatics, 86(S1), 321–334. [Link]
02/prot.25425
12. Protein structure prediction and structural genomics. Science. Retrieved May 6, 2024, from [Link]
i/abs/ , [Link] .
13. Simons, K. T., Kooperberg, C., Huang, E., & Baker, D. (1997). Assembly of protein tertiary structures from fragments with
similar local sequences using simulated annealing and bayesian scoring functions1. Journal of Molecular Biology, 268(
1), 209–225. [Link]
[Link] jCXRi4lLmjSpYsSElbKJa0Q== Page 10 of 15
eProofing 15/08/24, 7:57 PM
14. Xu, D., & Zhang, Y. (2012). Ab initio protein structure assembly using continuous structure fragments and optimized kno
wledge-based force field. Proteins: Structure, Function, and Bioinformatics, 80(7), 1715–1735. [Link]
ot.24065
15. “It will change everything”: DeepMind’s AI makes gigantic leap in solving protein structures-document-gale academic O
neFile. Retrieved May 6, 2024, from [Link]
t=r&linkaccess=abs&issn=00280836&p=AONE&sw=w&userGroupName=anon%7E51452b6e&aty=open-web-entry .
16. Browne, W. J., North, A. C. T., Phillips, D. C., Brew, K., Vanaman, T. C., & Hill, R. L. (1969). A possible three-dimensional s
tructure of bovine α-lactalbumin based on that of hen’s egg-white lysozyme. Journal of Molecular Biology, 42(1), 65–8
6. [Link]
17. McGuffin, L. J., Adiyaman, R., Maghrabi, A. H. A., et al. (2019). IntFOLD: An integrated web resource for high performanc
e protein structure and function prediction. Nucleic Acids Research, 47(W1), W408–W413. [Link]
z322
18. Kelley, L. A., Mezulis, S., Yates, C. M., Wass, M. N., & Sternberg, M. J. (2015). The Phyre2 web portal for protein modellin
g, prediction and analysis. Nature Protocols, 10(6), 845–858. [Link]
19. PHYRE Protein Fold Recognition Server. Retrieved May 1, 2024, from [Link]
=help/interpret_intensive .
20. The Phyre2 web portal for protein modelling, prediction and analysis-PMC. Retrieved May 1, 2024, from [Link]
[Link]/pmc/articles/PMC5298202/ .
21. ESyPred3D submitting form. Retrieved May 1, 2024, from [Link]
d/ .
22. Lambert, C., Léonard, N., De Bolle, X., & Depiereux, E. (2002). ESyPred3D: Prediction of proteins 3D structures. Bioinfor
matics, 18(9), 1250–1256. [Link]
23. Buß, O., Rudat, J., & Ochsenreither, K. (2018). FoldX as protein engineering tool: Better than random based approache
s? Computational and Structural Biotechnology Journal, 16, 25–33. [Link]
24. Welcome to Biskit!—Biskit: Python for structural bioinformatics. Retrieved April 17, 2024, from [Link] .
25. Källberg, M., Wang, H., Wang, S., et al. (2012). Template-based protein structure modeling using the RaptorX web serve
r. Nature Protocols, 7(8), 1511–1522. [Link]
26. RaptorX. Retrieved April 17, 2024, from [Link] .
27. RaptorX-complex contact: A protein complex contact map prediction server. Retrieved May 6, 2024, from [Link]
[Link]/ComplexContact/ .
28. About MODELLER. Retrieved May 6, 2024, from [Link] .
29. Söding, J., Biegert, A., Lupas, A. N. (2005). The HHpred interactive server for protein homology detection and structure
prediction. Nucleic Acids Research, 33(Web Server issue), W244–W248. [Link] .
30. The Rosetta Software. RosettaCommons. Retrieved May 1, 2024, from [Link] .
[Link] jCXRi4lLmjSpYsSElbKJa0Q== Page 11 of 15
eProofing 15/08/24, 7:57 PM
31. Land, H., & Humble, M. S. (2018). YASARA: A tool to obtain structural guidance in biocatalytic investigations. Methods in
Molecular Biology Clifton NJ, 1685, 43–67. [Link]
32. Wang, Y., Cheng, X., Shan, Q., et al. (2014). Simultaneous editing of three homoeoalleles in hexaploid bread wheat conf
ers heritable resistance to powdery mildew. Nature Biotechnology, 32(9), 947–951. [Link]
33. Schwede, T., Kopp, J., Guex, N., & Peitsch, M. C. (2003). SWISS-MODEL: An automated protein homology-modeling ser
ver. Nucleic Acids Research, 31(13), 3381–3385.
34. Bhageerath-H. Retrieved May 6, 2024, from [Link] .
35. Bhageerath-H: A homology/ab initio hybrid server for predicting tertiary structures of monomeric soluble proteins-PMC.
Retrieved May 6, 2024, from [Link] .
36. Krieger, E., Nabuurs, S. B., & Vriend, G. (2003). Homology modeling. Methods of Biochemical Analysis, 44, 509–523. ht
tps://[Link]/10.1002/0471721204.ch25
37. Rodriguez, R., Chinea, G., Lopez, N., Pons, T., & Vriend, G. (1998). Homology modeling, model and software evaluation:
Three related resources. Bioinformatics (Oxford, England), 14(6), 523–528. [Link]
23
38. Smith, T. F., & Waterman, M. S. (1981). Identification of common molecular subsequences. Journal of Molecular Biology,
147(1), 195–197. [Link]
39. Needleman, S. B., & Wunsch, C. D. (1970). A general method applicable to the search for similarities in the amino acid s
equence of two proteins. Journal of Molecular Biology, 48(3), 443–453. [Link]
4
40. Altschul, S. F., Gish, W., Miller, W., Myers, E. W., & Lipman, D. J. (1990). Basic local alignment search tool. Journal of Mol
ecular Biology, 215(3), 403–410. [Link]
41. Sali, A., & Blundell, T. L. (1993). Comparative protein modelling by satisfaction of spatial restraints. Journal of Molecular
Biology, 234(3), 779–815. [Link]
42. Fiser, A., Do, R. K., & Sali, A. (2000). Modeling of loops in protein structures. Protein Science Publication Protein Societ
y, 9(9), 1753–1773. [Link]
43. Sánchez, R., & Sali, A. (1997). Advances in comparative protein-structure modelling. Current Opinion in Structural Biolo
gy, 7(2), 206–214. [Link]
44. Tappura, K. (2001). Influence of rotational energy barriers to the conformational search of protein loops in molecular dy
namics and ranking the conformations. Proteins, 44(3), 167–179. [Link]
45. Scouras, A. D., & Daggett, V. (2011). The Dynameomics rotamer library: Amino acid side chain conformations and dyna
mics from comprehensive molecular dynamics simulations in water. Protein Science Publication Protein Society, 20(2),
341–352. [Link]
46. Hintze, B. J., Lewis, S. M., Richardson, J. S., & Richardson, D. C. (2016). Molprobity’s ultimate rotamer-library distributio
ns for model validation. Proteins, 84(9), 1177–1189. [Link]
[Link] jCXRi4lLmjSpYsSElbKJa0Q== Page 12 of 15
eProofing 15/08/24, 7:57 PM
47. Hansen, D. F., & Kay, L. E. (2011). Determining valine side-chain rotamer conformations in proteins from methyl 13C che
mical shifts: Application to the 360 kDa half-proteasome. Journal of the American Chemical Society, 133(21), 8272–828
1. [Link]
48. Liu, H., Elstner, M., Kaxiras, E., Frauenheim, T., Hermans, J., & Yang, W. (2001). Quantum mechanics simulation of protei
n dynamics on long timescale. Proteins, 44(4), 484–489. [Link]
49. Krieger, E., Koraimann, G., & Vriend, G. (2002). Increasing the precision of comparative models with YASARA NOVA–a s
elf-parameterizing force field. Proteins, 47(3), 393–402. [Link]
50. Adcock, S. A., & McCammon, J. A. (2006). Molecular dynamics: Survey of methods for simulating the activity of protein
s. Chemical Reviews, 106(5), 1589–1615. [Link]
51. Hospital, A., Goñi, J. R., Orozco, M., & Gelpí, J. L. (2015). Molecular dynamics simulations: Advances and applications. A
dvances and Applications in Bioinformatics and Chemistry (AABC), 8, 37–47. [Link]
52. Hendlich, M., Lackner, P., Weitckus, S., et al. (1990). Identification of native protein folds amongst a large number of inc
orrect models. The calculation of low energy conformations from potentials of mean force. Journal of Molecular Biology,
216(1):167–180. [Link] .
53. Jones, D. T. (2001). Predicting novel protein folds by using FRAGFOLD. Proteins (Suppl 5), 127–132. [Link]
002/prot.1171 .
54. Jaroszewski, L., Rychlewski, L., Zhang, B., & Godzik, A. (1998). Fold prediction by a hierarchy of sequence, threading, a
nd modeling methods. Protein Science Publication Protein Society, 7(6), 1431–1440.
55. Altschul, S. F., Madden, T. L., Schäffer, A. A., et al. (1997). Gapped BLAST and PSI-BLAST: A new generation of protein
database search programs. Nucleic Acids Research, 25(17), 3389–3402. [Link]
56. Krogh, A., Brown, M., Mian, I. S., Sjölander, K., & Haussler, D. (1994). Hidden Markov models in computational biology: A
pplications to protein modeling. Journal of Molecular Biology, 235(5), 1501–1531. [Link]
57. Söding, J. (2005). Protein homology detection by HMM–HMM comparison. Bioinformatics, 21(7), 951–960. [Link]
rg/10.1093/bioinformatics/bti125
58. De Novo Protein Structure Prediction by QUARK. Retrieved May 6, 2024, from [Link] .
59. D-I-TASSER: deep learning-based protein structure prediction. Retrieved May 6, 2024, from [Link]
-TASSER/ .
60. Google-deepmind/alphafold. Published online May 5, 2024. Retrieved May 6, 2024, from [Link]
pmind/alphafold .
61. Jones, D. T., Buchan, D. W. A., Cozzetto, D., & Pontil, M. (2012). PSICOV: Precise structural contact prediction using spa
rse inverse covariance estimation on large multiple sequence alignments. Bioinformatics, 28(2), 184–190. [Link]
g/10.1093/bioinformatics/btr638
62. O S. Sokrypton/GREMLIN. Published online December 26, 2023. Retrieved May 6, 2024, from [Link]
ton/GREMLIN .
[Link] jCXRi4lLmjSpYsSElbKJa0Q== Page 13 of 15
eProofing 15/08/24, 7:57 PM
63. Pearce, R., & Zhang, Y. (2021). Toward the solution of the protein structure prediction problem. Journal of Biological Ch
emistry, 297(1), 100870. [Link]
64. CCMpred–fast and precise prediction of protein residue-residue contacts from correlated mutations-PubMed. Retrieve
d May 6, 2024, from [Link] .
65. Soedinglab/CCMpred. Published online April 9, 2024. Retrieved May 6, 2024, from [Link]
red .
66. NeBcon: protein contact map prediction using neural network training coupled with naïve Bayes classifiers. Bioinformati
cs. Oxford Academic. Retrieved May 6, 2024, from [Link]
0.
67. Li, Y., Hu, J., Zhang, C., Yu, D. J., & Zhang, Y. (2019). ResPRE: High-accuracy protein contact prediction by coupling pre
cision matrix with deep residual neural networks. Bioinformatics, 35(22), 4647–4655. [Link]
tics/btz291
68. Greener, J. G., Kandathil, S. M., & Jones, D. T. (2019). Deep learning extends de novo protein modelling coverage of gen
omes using iteratively predicted structural constraints. Nature Communications, 10(1), 3977. [Link]
467-019-11994-0
69. TripletRes: contact map prediction based on a triplet of coevolutionary features and deep residual neural networks. Retr
ieved May 6, 2024, from [Link] .
70. Simons, K. T., Bonneau, R., Ruczinski, I., & Baker, D. (1999). Ab initio protein structure prediction of CASP III targets usin
g ROSETTA. Proteins, (Suppl 3), 171–176. [Link]
q.
71. Cohen, M., Potapov, V., & Schreiber, G. (2009). Four distances between pairs of amino acids provide a precise descripti
on of their interaction. PLoS Computational Biology, 5(8), e1000470. [Link]
72. Han, K. F., & Baker, D. (1995). Recurring local sequence motifs in proteins. Journal of Molecular Biology, 251(1), 176–18
7. [Link]
73. Shortle, D., Simons, K. T., & Baker, D. (1998). Clustering of low-energy conformations near the native structures of small
proteins. Proceedings of the National Academy of Sciences USA, 95(19), 11158–11162. [Link]
9.11158
74. All-atom empirical potential for molecular modeling and dynamics studies of proteins. The Journal of Physical Chemistr
y B. Retrieved May 6, 2024, from [Link] , [Link] .
75. Brooks, B. R., et al. (1983). CHARMM: A program for macromolecular energy, minimization, and dynamics calculations. J
ournal of Computational Chemistry. Retrieved May 6, 2024, from [Link] , [Link]
10.1002/jcc.540040211 . (Wiley Online Library)
76. Neria, E., Fischer, S., & Karplus, M. (1996). Simulation of activation free energies in molecular systems. The Journal of C
hemical Physics, 105(5), 1902–1921. [Link]
[Link] jCXRi4lLmjSpYsSElbKJa0Q== Page 14 of 15
eProofing 15/08/24, 7:57 PM
77. Cornell, W. D., Cieplak, P., Bayly, C. I., et al. (1995, 1996). A second generation force field for the simulation of proteins,
nucleic acids, and organic molecules. Journal of the American Chemical Society, 117, 5179−5197. Journal of the America
n Chemical Society, 118(9), 2309–2309. [Link] .
78. Duan, Y., & Kollman, P. A. (1998). Pathways to a protein folding intermediate observed in a 1-microsecond simulation in
aqueous solution. Science, 282(5389), 740–744. [Link]
79. Weiner, S. J., Kollman, P. A., Case, D. A., et al. (1984). A new force field for molecular mechanical simulation of nucleic a
cids and proteins. ACS Publications. [Link]
80. Scott, W. R. P., Hünenberger, P. H., Tironi, I. G., et al. (1999). The GROMOS biomolecular simulation program package. J
ournal of Physical Chemistry A, 103(19), 3596–3607. [Link]
81. Atomic-level characterization of the structural dynamics of proteins. Science. Retrieved May 6, 2024, from [Link]
[Link]/doi/abs/ , [Link] .
82. Bowie, J. U., & Eisenberg, D. (1994). An evolutionary approach to folding small alpha-helical proteins that uses sequenc
e information and an empirical guiding fitness function. Proceedings of the National Academy of Sciences, 91(10), 443
6–4440. [Link]
83. Improved protein structure prediction using predicted interresidue orientations. PNAS. Retrieved May 6, 2024, from htt
ps://[Link]/doi/abs/ , [Link] .
84. Senior, A. W., Evans, R., Jumper, J., et al. (2020). Improved protein structure prediction using potentials from deep learn
ing. Nature, 577(7792), 706–710. [Link]
85. Pearce, R., & Zhang, Y. (2021). Deep learning techniques have significantly impacted protein structure prediction and pr
otein design. Current Opinion in Structural Biology, 68, 194–207. [Link]
86. Jumper, J., Evans, R., Pritzel, A., et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature, 596(
7873), 583–589. [Link]
87. Callaway, E. (2020). “It will change everything”: DeepMind’s AI makes gigantic leap in solving protein structures. Nature,
588(7837), 203–205.
88. Yang, J., Zhang, W., He, B., et al. (2016). Template-based protein structure prediction in CASP11 and retrospect of I-TA
SSER in the last decade. Proteins, 84(Suppl 1), 233–246. [Link] .
© Springer Nature
[Link] jCXRi4lLmjSpYsSElbKJa0Q== Page 15 of 15
View publication stats