0% found this document useful (0 votes)
10 views32 pages

Protein Structure and Visualization Guide

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
10 views32 pages

Protein Structure and Visualization Guide

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

PROTEIN DATABASE,

STRUCTURE, VISUALISATION
& PREDICTION ALGORITHM
By
Maajjidah (298)
Madhuraa Sree (303)
Madhu Mitha (260)
Sai Mahathi (302)
Vaishnavi (296)
Protein synthesis

● Production of polypeptide chains(proteins)

● Two phases
Transcription & Translation

● m RNA must be processed before it leaves the


nucleus of eukaryotic cells
Three types of RNA
● Messenger RNA(mRNA) carrier genetic information to
the ribosomes.

● Ribosomal RNA(rRNA) along with protein,makes up the


ribosomes.

● Transversal RNA(tRNA) transfers amino acids to the


ribosomes where proteins are synthesized.
Genes & proteins
● Proteins are made of amino acids linked together by
peptide bonds

● 20 different amino acids exist

● Amino acids chains are called polypeptides

● Segment of DNA that codes for amino acid sequence in


a protein are called genes
Transcription & translation

● Transcription is the process of copying a segment of DNA


into RNA.

● Both DNA and RNA are nucleic acids.

● Some segments of DNA are transcribed into RNA


molecules that encode proteins called messenger RNA.

● Other segments of DNA are transcribed into RNA


molecules called non-coding RNAs
End product-the protein!
● The end products of protein synthesis is a primary
structure of a protein

● A sequence of amino acid bonded together by peptide


bonds

● Protein synthesis starts with transcription,which occurs


in the nucleus.
Introduction to Protein Databases
● Definition: A protein database contains detailed
information about the 3D structure of proteins and other
biological macromolecules.
● Purpose: These databases help scientists study the
structure, function, and interactions of proteins.
● Primary Example: The Protein Data Bank (PDB) is the
most well-known and comprehensive database for
protein 3D structures.
Experimental Methods for Determining Protein Structures

● X-ray Crystallography: Provides high-resolution 3D


structures by analyzing diffraction patterns of protein
crystals.
● Nuclear Magnetic Resonance (NMR) Spectroscopy:
Determines protein structures in solution by analyzing
nuclear magnetic resonance signals.
● Electron microscopes use signals from the interaction of
an electron beam with the sample to obtain information
about structure, morphology, and composition.
Key Features of PDB Files
● What PDB Files Contain:
○ Molecule name: Identifies the protein or macromolecule.
○ Primary and secondary structure: Sequences and
structural motifs.
○ Ligands: Small molecules that interact with the protein.
○ Atomic Coordinates: 3D position of each atom.
○ Crystallographic data and NMR experimental data.
● PDB ID: Each entry in the PDB is given a unique 4-character
alphanumeric identifier.
PDB Database Access and Use
● Search Capabilities: Users can search by structure,
sequence, function, or experimental technique.
● Visualization Tools: Allows users to visualize the 3D
structure of proteins.
● Download Options: Users can download PDB files for
further analysis.
● Global Contribution: Data are submitted by scientists
worldwide and maintained by the Worldwide Protein Data
Bank (wwPDB).
● Public Access: All data in the PDB are freely accessible.
Related Databases Derived from PDB
● SCOP (Structural Classification of Proteins): Groups
proteins based on their structural similarities.
● HSSP (Homology-Derived Secondary Structure of
Proteins): Provides both 3D structure and 1D sequence
information.
● CATH: Classifies proteins based on evolutionary
relationships and structural domains.
● Purpose: These databases aid in understanding the
evolution, function, and classification of protein
structures.
Introduction to Protein Structure

Proteins are essential macromolecules involved in virtually


every cellular process.

Protein function is directly linked to its three-dimensional


structure.

Four levels of protein structure: Primary, Secondary,


Tertiary, and Quaternary.
Primary Structure
● Definition: The amino acid sequence of a protein.

● Amino Acids: 20 standard amino acids.

● Peptide Bonds: Link amino acids through condensation


reactions, forming the polypeptide chain.

● Significance: The sequence determines the final folded shape


of the protein.

● Example: A mutation in the amino acid sequence can cause


dysfunction (e.g., sick)
Secondary Structure
● Definition: Localized folding of the polypeptide chain into
structures like alpha helices and beta sheets.

● Alpha Helix: Right-handed spiral, stabilized by hydrogen bonds.

● Beta Sheet: Parallel or antiparallel strands, stabilized by


hydrogen bonds.

● Turns and Loops: Flexible regions linking secondary structure


elements.

● Key Feature: Stabilized by hydrogen bonds between backbone


atoms.
Tertiary Structure

● Definition: The 3D shape of the entire polypeptide chain.


● Interactions:
● Hydrophobic interactions pull non-polar side chains
inward.
● Hydrogen bonds stabilize the shape.
● Disulfide bonds between cysteine residues add stability.
● Significance: The tertiary structure forms the protein's
active site (important for enzyme function).
Quaternary Structure
● Definition: The arrangement of multiple polypeptide
chains (subunits) in a multi-subunit protein.
● Subunits: Held together by non-covalent interactions
(hydrogen bonds, ionic bonds, hydrophobic
interactions).
● Example: Hemoglobin – four subunits working together
to transport oxygen.
● Significance: Interaction between subunits is essential
for functional activity.
Protein Folding and Denaturation
● Protein Folding:
The process by which proteins fold into their functional 3D
shape.
Guided by the primary structure (amino acid sequence).
Chaperone proteins help prevent misfolding and aggregation.
● Denaturation:
Loss of a protein's 3D structure due to external factors (heat, pH,
chemicals).
In most cases, denaturation results in loss of function.
Protein Structure-Function Relationship

● Enzymes: Active sites are shaped to bind substrates and


catalyze reactions.
● Antibodies: Bind to specific antigens for immune
defense.
● Transport Proteins: Hemoglobin changes shape to bind
and release oxygen.
● Mutation Impact: Small changes in structure can
dramatically alter function (e.g., sickle-cell disease).
Protein visualisation tools

● Protein visualization tools are essential in bioinformatics


and molecular biology for understanding the structure,
function, and interactions of proteins.

● These tools help researchers analyze and interpret


complex protein data by providing interactive and
graphical representations.
Uses of visualising tools
1. To see the protein's shape: Proteins are like tangled strings. Visual
tools untangle them and show us their structure.

2. To find important spots: They help locate areas on the protein where
other molecules can attach, like where a drug might bind.

3. To compare proteins: Scientists can use these tools to compare


similar proteins and find differences.

4. To study movement: Proteins move and change. Some tools let us


watch how proteins wiggle or interact over time.
PyMol -
● PyMol is a powerful molecule visualization software with the
following main features:
● Able to produce high-quality graphics ready for
publications.
● Able to create movies.
● Able to measure bond distances and angles.
● Has an extensive help system.
● Structures can be sliced, diced, and reassembled on the fly
and written out to standard files.
● Both command line interface and graphical user interface
are provided.
● Python API is provided to access all functionalities.
Rasmol/RasTop -
● RasMol is a molecular graphics program intended for the
visualisation of proteins, nucleic acids and small
molecules.
● The program reads in a molecule coordinate file and
interactively displays the molecule on the screen in a
variety of colour schemes and molecule representations.
● Currently available representations include depth cued
wireframes, 'Dreiding' sticks, space filling (CPK)
spheres, ball and stick, solid and strand biomolecular
ribbons, atom labels and dot surfaces.
JMol
Jmol creates a 3D graphical representation of molecules,
allowing users to:
● Rotate and zoom in on the molecule.
● View molecular structures in different styles (e.g.,
ball-and-stick, space-filling, ribbons).
● Highlight specific regions, like active sites or ligands.
● Measure distances, angles, and bond lengths.
● Interact with molecules on websites (via its
browser-based version, JSmol).
PROTEIN PREDICTION ALGORITHM
● Protein prediction algorithms are used to predict the structure, function, or
behavior of proteins based on their amino acid sequences.
● These algorithms are crucial in bioinformatics for understanding how proteins fold,
how they interact, and how mutations may affect their functionality.
● These algorithms are often combined with large protein sequence and structure
databases to improve prediction accuracy.

Here are some key types of protein prediction algorithms:

1. Sequence Alignment Algorithms 4. Protein-Protein Interaction Prediction

2. Secondary Structure Prediction 5. Functional Prediction

3. Tertiary Structure Prediction


SECONDARY STRUCTURE PREDICTION ALGORITHM
● Secondary structure prediction algorithms are used to predict
the local structures within a protein, such as alpha-helices,
beta-sheets, and coils, based on its amino acid sequence.

● Here are the key approaches and algorithms for secondary


structure prediction:
1. Ab Initio - Based methods
2. Homology - Based Methods
AB - INITIO - BASED METHODS
● The ab initio methods, which belong to early generation
methods, predict secondary structures based on statistical
calculations of the residues of a single query sequence.
● It measures the relative propensity(Natural tendency) of each
amino acid belonging to a certain secondary structure element.
● The propensity scores are derived from known crystal
structures.
● Examples of ab initio prediction are the Chou–Fasman and
Garnier, Osguthorpe, Robson (GOR) algorithms
The Chou–Fasman algorithm
● Developed by Chou and Fasman in 1974, the algorithm uses statistical data derived
from known protein structures to predict whether each segment of a protein will form
an alpha helix, beta sheet, or coil (random coil) in its final 3D structure.
● It determines the propensity or intrinsic tendency of each residue to be in the helix,
strand, and β-turn conformation using observed frequencies found in protein crystal
structures.
HOMOLOGY BASED METHODS
● Homology-based methods are a class of techniques in bioinformatics
and computational biology that use the evolutionary relationships
between proteins (or genes) to predict the structure or function of an
unknown protein based on known homologous sequences.
● These methods rely on the assumption that proteins with similar
sequences (homologs) share similar structures and functions, as they
are likely derived from a common ancestor.
● This type of method combines the ab initio secondary structure
prediction of individual sequences and alignment information from
multiple similar sequences (>35% identity).
● This homology based method has helped improve the prediction
accuracy by another 10% over the second-generation methods.
PHD - NEURAL NETWORK ALGORITHM

● The PHD (Profile Hidden Markov Model) Neural Network


Algorithm is a hybrid computational approach combining Hidden
Markov Models (HMMs) and neural networks to predict protein
secondary structures, protein folding, or other aspects of sequence
analysis.
● This approach leverages both the statistical power of HMMs and
the learning capacity of neural networks to improve prediction
accuracy in bioinformatics applications.
Comparison of the Three Methods

Method Accuracy Strengths Limitations

Chou-Fasman 60-70% Simple, fast, computationally Limited accuracy, does not consider
efficient sequence context

Ab Initio 60-80% No need for homologous Computationally expensive, lower accuracy


sequences, flexible compared to modern methods

PHD Neural 75-85% High accuracy, uses evolutionary Requires homologous sequences,
Networks
data and machine learning computationally intensive

You might also like