AI in Protein Structure Analysis
AI in Protein Structure Analysis
CBSE ROLL NO :
CLASS : XII
SUBJECT : BIOLOGY
1
Arsha Vidya Mandir
BIOLOGY PROJECT
BONAFIDE CERTIFICATE
I hereby certify that this project titled “Unravelling Protein Structures with AI” is
a bonafide work done by Krithika Rajasekar of Grade XII of Arsha Vidya Mandir
during the academic year, 2025 – 26.
Dr. Aruna A
School Seal
2
ACKNOWLEDGEMENT
I thank the CBSE Board for giving me an opportunity to carry out a research project
and widen our knowledge on the chosen topic.
I thank my Biology teacher, for providing me the required guidance and support. It is
my humble pleasure to acknowledge my deep sense of gratitude to her, for the constant
help and suggestions at each and every stage, without which it wouldn’t have been
possible to complete this project.
Last but not the least I thank the Almighty for giving me patience and perseverance to
successfully complete this work.
3
TABLE OF CONTENTS
1. Introduction 5
2. Aim 6
3. Objectives 6
4. Methodology 6
13. Discussion 32
14. Conclusion 33
15. Bibliography 34
4
Introduction:
5
Aim:
To explore how AI, particularly AlphaFold 2 and RF Diffusion, has transformed our
ability to predict and design protein structures, and to understand the broader
implications of these advances on science.
Objectives:
Methodology:
This report employs a qualitative analysis of secondary data sourced from scientific
literature, research papers, and documentaries.
1. Literature Review:
A comprehensive examination of academic journals, review papers and
Nobel Prize citations that explore AlphaFold and advancements in synthetic
protein design.
2. Technical Breakdown:
Analyzing the architecture and training process of AlphaFold 2, with a focus on
6
its use of deep learning techniques, evolutionary data and attention-based
transformer models for protein structure prediction.
3. Case Studies and Application Exploration:
Investigating the real-world implementation of AlphaFold and RF Diffusion,
including their roles in vaccine development, anti-venom production and the
design of enzymes for plastic degradation.
7
Amino Acids and Peptide Bonds: The Chemistry of Proteins
Figure 1: Structure of a general amino acid showing the α-carbon, –NH₂, –COOH, –
H, and R group. Source: [Link]
The term ‘amino acid’ is derived from the presence of both an amine and a carboxyl
group in each molecule.
At physiological pH (around 7.4), amino acids generally do not exist in their neutral
form. Instead, they predominantly exist in the Zwitterionic state. In this form, the amine
group gains a proton to become –NH₃⁺, while the carboxyl group loses a proton to
become –COO⁻. This intra-molecular charge balancing makes the molecule
electrically neutral overall, but with both positive and negative charges present.
8
Figure 2: Zwitterionic form of an amino acid, showing –NH₃⁺ and –COO⁻ groups
with respective charges. Source: [Link]
The zwitterionic nature of amino acids is crucial for peptide bond formation. During
a condensation reaction, the amine group of one amino acid reacts with the carboxyl
group of another – forming a peptide bond. This covalent bond links the two amino
acids together while releasing one molecule of water (H2O).
Figure 3: Peptide bond formation between two amino acids, with a water
molecule shown as a by-product. Source: [Link]
This reaction can repeat multiple times, allowing amino acids to link together and form
polypeptide chains, which fold into complex three-dimensional structures to form
functional proteins.
9
Starting from the polypeptide chain, proteins exhibit four distinct levels of structural
organization, each contributing to their final shape and biological function: primary,
secondary, tertiary, and, in some cases, quaternary structures.
1. Primary Structure:
It refers to the polypeptide chain, with amino acids connected by peptide bonds.
This sequence is genetically determined and dictates all higher levels of protein
structure. The chain has free amine group (N-terminus) at one end and a free
carboxyl group (C-terminus) at the other.
Figure 4: Linear polypeptide chain showing peptide bonds and labeling of N-terminus
and C-terminus. Source: [Link]
2. Secondary Structure:
It arises from regular, repeating patterns formed by hydrogen bonding between the
backbone atoms of the polypeptide chain (not the side chains). The two most
common secondary structures are:
• α-Helix (Alpha Helix): A right-handed coil stabilized by hydrogen bonds
between the carbonyl oxygen of one amino acid and the amide hydrogen of
another four residues ahead in the chain. This results in a tightly coiled, spring-
like structure.
10
• β-Pleated Sheet (Beta Sheet): Formed when two or more strands of a
polypeptide chain lie side by side, connected by hydrogen bonds. These sheets
may be parallel or antiparallel, and the structure appears folded or pleated.
Figure 5: Alpha helix and beta pleated sheet structures, with hydrogen bonds shown.
Source: [Link]
Only right-handed α-helices are found in naturally occurring proteins due to steric
constraints and stability considerations.
3. Tertiary Structure:
It is the three-dimensional folding of a single polypeptide chain into a compact,
globular shape. This folding is stabilized by various interactions between side
chains (R groups), including:
• Hydrogen bonds
• Ionic interactions
• Hydrophobic interactions
• Disulfide bridges (in some proteins)
This structure can be visualized as a folded ball of yarn, with many internal
interactions holding it together. Not all proteins go beyond this level; for many, the
tertiary structure represents their functional form.
11
Examples include myoglobin (an oxygen-binding protein) and enzymes like
lysozyme or ribonuclease, which function independently in their tertiary form.
12
‘The Protein Folding Problem’ and Levinthal’s Paradox
The folding of a protein is determined by its amino acid sequence (the primary
structure), yet the journey from a linear chain to a functional protein remains one of
the greatest puzzles in molecular biology. This challenge, known as the ‘Protein
Folding Problem’, has both intrigued and frustrated scientists for over half a century.
So, why does this problem persist? If a protein's structure is determined by its primary
structure, shouldn’t it be straightforward to predict?
Apparently not. This paradox was first described by Cyrus Levinthal in 1969.
Levinthal argued that if a protein were to randomly sample all possible conformations,
even a small protein of 100 residues would take longer than the age of the universe to
fold correctly. Yet, in living cells, proteins consistently fold into their native structures
within milliseconds to seconds. This discrepancy, known as Levinthal’s Paradox,
revealed that folding is not a random process, but follows specific pathways and
principles — though the details remained elusive.
Figure 8: Comparing in vitro and in vivo protein folding situations using Levinthal's
Paradox. Source: [Link]
13
Also, understanding protein folding is not just an academic pursuit. Protein misfolding
is central to a range of human diseases, including Alzheimer’s, Parkinson’s,
Huntington’s, amyotrophic lateral sclerosis (ALS), and prion diseases like
Creutzfeldt–Jakob disease. In these conditions, proteins adopt abnormal
conformations that aggregate into toxic fibrils or plaques, damaging cells and tissues.
In prion diseases, misfolded proteins can even propagate their abnormal structure to
other molecules, making them particularly destructive.
The quest to determine protein structures began long before the advent of AI. For much
of the twentieth century, structural biology relied on painstaking experimental
techniques that required immense technical skill, sophisticated instruments, and
14
years of effort. While revolutionary for their time, these methods were also limited,
slowing the pace of discovery.
X-ray crystallography has been the dominant technique for protein structure
determination since the 1950s. It works on a straightforward but technically demanding
principle: proteins are crystallized into an ordered lattice, then bombarded with X-
rays. The way X-rays scatter or diffract off the crystal provides data that can be
mathematically transformed into an electron density map. By fitting the amino acid
sequence into this density, researchers can reconstruct the protein’s three-dimensional
atomic structure.
The first protein structures solved by X-ray crystallography were myoglobin and
haemoglobin, published in the late 1950s and early 1960s by John Kendrew and Max
Perutz, respectively. These achievements earned them the 1962 Nobel Prize in
Chemistry and demonstrated the enormous potential of crystallography to reveal the
molecular basis of life.
15
In the 1980s, NMR spectroscopy emerged as an alternative to crystallography,
especially for proteins that could not be crystallized. NMR exploits the magnetic
properties of atomic nuclei, particularly hydrogen, carbon (13C), and nitrogen (15N).
When placed in a strong magnetic field and exposed to radiofrequency pulses, these
nuclei resonate in ways that depend on their chemical environment. By analysing these
signals, researchers can infer interatomic distances and angles, helping to construct
three-dimensional structures.
16
complexes are flash-frozen in vitreous ice and imaged with an electron beam.
Thousands to millions of images of randomly oriented molecules are computationally
combined to reconstruct a three-dimensional structure.
17
complexes that were previously intractable by crystallography or NMR. Structures such
as ion channels, ribosomes, and membrane-bound receptors became accessible,
ushering in a golden age of structural biology.
In 2017, the Nobel Prize in Chemistry was awarded to Jacques Dubochet, Joachim
Frank, and Richard Henderson for their contributions to cryo-EM. Cryo-EM was
widely used to study large, flexible, and membrane-associated proteins,
complementing crystallography and NMR.
Other methods like Small-angle X-ray scattering (SAXS), Mass spectrometry (MS)
and FRET (Förster resonance energy transfer) were also developed. However, they
were not widely used and did not create as much impact as the three methods mentioned
above.
Determining protein structures has long been one of the most complex challenges in
biology. Though experimental techniques such as X-ray crystallography, NMR
spectroscopy and cryo-EM made significant strides in the 20th and 21st centuries,
numerous technical, computational and biological limitations continue to impede
scalability, accessibility and accuracy.
18
Experimental Challenges
Computational Challenges
1. Folding Complexity: Proteins are made of 20 amino acids, each with unique
chemical properties. A 100-residue chain could fold into an astronomical number
of conformations. This vast “folding space,” as pointed out by Levinthal's
paradox, makes it infeasible to simulate all possible structures through brute
force, even with powerful computational models.
2. Molecular Dynamics (MD) Limitations: MD simulations rely on computing
atomic motions based on physical laws but are constrained by computational
19
power. Simulating protein folding, which occurs on timescales of milliseconds
to seconds, requires enormous computational resources. Even supercomputers
can only model small proteins over limited timescales.
3. Incomplete Force Fields: Current computational models depend on force fields
to describe atomic interactions. While they have improved, these force fields are
still imperfect, failing to capture long-range electrostatic interactions, solvent
effects, or entropy contributions, leading to inaccuracies in simulating protein
behavior over time.
4. Limited Training Data: Machine learning models require large, high-quality
datasets, but until recently, structural databases were incomplete, especially for
membrane proteins, disordered regions, and multi-subunit complexes. This data
scarcity hampers the accuracy of predictions, especially for underrepresented
protein types.
Biological Challenges
20
This is a major hurdle in understanding cellular processes, as structural data is
often needed for complexes, not just individual proteins.
4. Environmental Effects: Proteins fold and function in the crowded,
heterogeneous environment of the cell, which is vastly different from the
purified, dilute conditions typically used in experimental setups. Factors such as
molecular crowding, pH, ionic strength, and interactions with chaperone proteins
all influence folding and stability, but they are rarely accurately modeled in
traditional experimental or computational approaches.
The challenges of protein structure determination are not only scientific but also
societal. The cost and complexity of these methods have historically limited access to
structural biology, creating disparities between well-funded institutions and those in
low-resource settings. The slow pace of structural determination can also delay
responses to urgent global health issues. For instance, during the early stages of the
COVID-19 pandemic, structural biologists raced to solve the SARS-CoV-2 spike
protein's structure using cryo-EM. AI-based tools like AlphaFold, however,
accelerated research by quickly providing high-quality predictive models of proteins.
21
Additionally, the inability to resolve protein structures at scale has hindered progress in
drug discovery, enzyme engineering and even basic biological research. With
millions of proteins in the human proteome alone, traditional methods would never
meet the growing demand for structural insights in an era of increasing scientific and
biomedical challenges.
This section explores how AlphaFold emerged, the innovations behind it, its validation
through the CASP competitions and the far-reaching implications of its success for
biology and beyond.
Throughout the 1980s and 1990s, researchers tried to predict protein structures from
sequence using physical chemistry rules or homology modelling, where a protein’s
structure was inferred from a similar, already-known one. While somewhat effective,
these approaches lacked broad accuracy. The Critical Assessment of Protein
Structure Prediction (CASP) competition, launched in 1994, regularly highlighted
just how far computational methods lagged behind experimental ones.
For over two decades, progress at CASP was slow and incremental. Traditional
models were hampered by limited computing power, imperfect force fields and a
lack of large, high-quality datasets. By the early 2010s, enthusiasm had faded—
some questioned whether accurate prediction was even possible. It was in this climate
of stagnation that DeepMind’s AlphaFold made its debut.
22
The CASP Competitions
CASP provided the testing ground that ultimately validated AlphaFold. In CASP11
(2014) and CASP12 (2016), machine learning began to show promise, particularly in
predicting residue–residue contact maps based on evolutionary covariance. These
early breakthroughs hinted that statistical patterns in sequence databases contained
hidden structural information.
However, the true revolution came in CASP14 (2020), when AlphaFold2 delivered
unprecedented accuracy. For many targets, its predictions matched experimental
structures within the margin of experimental error. The results were so striking that
organizers declared the problem “essentially solved”.
Figure 14: CASP Scores for the years 2012, 2013 and 2014; showing AlphaFold's
rise. Source: [Link]
23
Architecture of AlphaFold2: A Conceptual Leap
AlphaFold2 was not simply an incremental improvement over its predecessor; it was
a conceptual and architectural leap. Its design integrated insights from machine
learning, structural biology and physics into a cohesive system. The following
aspects aid its accuracy:
24
4. Recycling Mechanism
A unique innovation in AlphaFold2 is its recycling mechanism. Predictions are
fed back into the network, allowing iterative refinement. This mimics the way
scientists refine models by repeatedly comparing them to experimental data, but
it is executed automatically within the AI.
25
26
Figure 15: Working of AlphaFold 2. Source: [Link]
Impact of AlphaFold2’s Breakthrough
1. Validation against previously determined Experimental Structures from
Physics (using root-mean-square derivations or RMSD).
2. Democratization of Protein Structures through public release of the AlphaFold
Protein Structure Database.
3. Acceleration of Biological Discovery due to availability of accurate protein
structures.
4. Complementarity with Experimental Methods (without rendering them
obsolete).
27
picture. Building upon the successes of AlphaFold and advances in deep generative
modeling, RF Diffusion represents a new phase in structural biology: not just
understanding proteins as they are, but engineering them as we wish them to be.
28
29
Figure 16: Comparison of AlphaFold 2 (Protein Folding) with RF Diffusion (Protein
Design). Source: [Link]
The Principles of Diffusion Models
In the case of RF Diffusion, the data are not images but protein backbones — the
three-dimensional coordinates of atoms forming polypeptide chains. The model learns
to add and then remove noise from these structures, ultimately allowing it to sample
novel protein conformations consistent with the principles of folding.
RF Diffusion Architecture
RF Diffusion integrates several key innovations that make it uniquely suited for protein
design:
1. Rotationally and Translationally Equivariant Representations
Proteins are three-dimensional objects, and their orientation in space should not
affect how they are represented or understood. RF Diffusion employs SE(3)-
equivariant neural networks, which respect the symmetries of 3D space. This
ensures that the model treats a protein identically whether it is rotated,
translated or flipped, a critical feature for accurate structural learning.
30
2. Conditioning for Functional Design
RF Diffusion can be conditioned to design proteins with specific properties. For
example, researchers can instruct the model to create a binding pocket with a
desired shape, or to generate symmetric assemblies useful for nanomaterials.
This conditional generation opens the door to targeted protein design for
biomedical and industrial applications.
3. Integration with Sequence Design Models
Once a backbone structure is generated, additional tools such as
ProteinMPNN (a sequence design network also developed by Baker’s group)
can assign amino acid sequences likely to fold into that structure. These
sequences can then be validated experimentally, creating a full pipeline from
design to laboratory realization.
31
• Therapeutic Binders that attach to viral proteins such as SARS-CoV-2 spike
protein and neutralize infection. Unlike antibodies, which are large and
complex, these small designed binders are easier to manufacture and potentially
more stable, offering new modalities in antiviral therapy.
The advent of artificial intelligence (AI) in protein science has resolved theoretical
challenges and unlocked a vast array of practical applications with great potential.
AlphaFold2 and RF Diffusion have the potential to reshape biotechnology, medicine,
environmental science and materials engineering.
This section explores the applications of AI-driven protein science across multiple
domains, highlighting concrete examples:
32
b. RF Diffusion has aided protein-based therapeutics, such as the production
of novel antibodies, enzymes and cytokines to mitigate disease.
c. Personalized medicine using patient-specific genomic data.
2. Environmental Sustainability
a. Plastic degradation using PETase enzymes synthasized by RF Diffusion.
b. Carbon capture and greenhouse gas mitigation using enzymes with
enhanced carbonic anhydrase-like activity.
33
Discussion:
When I began this project, I believed that AI might one day replace scientists
altogether. With so much discussion about automation and machines performing tasks
faster and more accurately than humans, it seemed inevitable that AI would eventually
take over scientific research. Learning that AlphaFold2 could predict millions of
protein structures in such a short time only strengthened that concern. It appeared that
if a computer could solve one of biology’s greatest mysteries, there would be little left
for people to do.
However, as I explored the topic further, I realized that AI is not replacing scientists
but supporting them. AlphaFold2 and RF Diffusion have not removed the need for
experiments; instead, they make research more focused, efficient and creative. AI can
analyse immense data sets and uncover hidden patterns; but human insight,
validation and ethical judgment remain essential.
Through this project, I learned that science today is defined by collaboration between
humans and machines. Every breakthrough — whether in protein prediction, enzyme
design or disease research — still depends on curiosity, teamwork and
responsibility.
34
Conclusion:
The integration of AI into protein science marks a historic turning point in biology.
AlphaFold2’s success in predicting protein structures with near-experimental
accuracy solved a challenge that had frustrated scientists for decades, providing
structural insights across the entire proteome and democratizing access through the
AlphaFold Protein Structure Database. Meanwhile, RF Diffusion has extended the
frontier from prediction to creation, enabling the design of entirely new proteins with
functions tailored to medicine, sustainability and nanotechnology. Together, these
breakthroughs represent a paradigm shift from descriptive to predictive and
generative biology.
In sum, the rise of AI in protein science is not merely a technical achievement but a
profound reimagining of what it means to study and engineer life. Its future trajectory
will be defined not only by scientific ingenuity but also by the collective wisdom with
which humanity chooses to wield this transformative power.
35
Bibliography:
• [Link]
• [Link]
model/#life-molecules
• [Link]
• [Link]
• [Link]
• [Link]
• [Link]
• [Link]
• [Link]
• [Link]
• [Link]
• [Link]
• [Link]
• [Link]
• [Link]
• NCERT Biology – Textbook for Class XI (2025-26 Edition)
36
TEACHER’S OBSERVATION
I acknowledge that this project is the original work, done by Krithika Rajasekar of Class
XII as per the guidelines of CBSE requirement. I also record that this project is a result
of adequate research and laboratory activities/ field visits made, which laid the
foundation for the completion of this project.
Teacher’s sign
37