UNIT II: BIOCHEMISTRY & MOLECULAR BIOLOGY
Comprehensive Study Notes
TOPICS COVERED IN THIS UNIT
Carbohydrates, Lipids, Proteins (Enzymes & Hormones), DNA, RNA, Human Genome Project, Genomics,
Sequence Databases, BLAST Tool
1. CHEMISTRY OF LIFE & CHEMICAL BONDS
1.1 Basic Concepts
• Atom: Basic unit of matter (protons, neutrons, electrons)
• Elements: Substances made of the same kinds of atoms
• Compounds: Molecules made of different atoms joined in precise arrangements
• Metabolism: ALL chemical reactions occurring in the body
• Organic Molecules: Always contain CARBON; large molecules with covalent bonds
1.2 Types of Chemical Bonds
Bond Type Description Strength Example
Ionic (Polar) Bond Electron DONATED from Moderate NaCl (table salt)
one atom to another;
creates cation (+) and
anion (-)
Covalent Bond Electrons SHARED Strong Cyclohexane (CH2)6
between atoms; can be
single, double, or triple
Hydrogen Bond Weak attraction Weak (~1/20 of Water molecules, DNA
between H and covalent) base pairs
electronegative atom (O
or N)
Van der Waals Weak, dispersed Very Weak Non-polar molecules
electromagnetic
interactions
Hydrophobic Interaction Non-polar molecules Weak Lipid bilayer
clustering together in
water
Key Point: Non-covalent bonds (ionic, hydrogen) are MUCH WEAKER than covalent bonds but are CRITICAL for:
◦ Maintaining 3D structure of proteins and nucleic acids
◦ Holding DNA double helix together
◦ Enzyme-substrate binding and antibody-antigen association
2. THE FOUR MACROMOLECULES (Biomolecules)
The four main types of carbon-based molecules found in living things:
Macromolecule Monomer Function Examples
Carbohydrates Monosaccharide (e.g., Dietary energy, storage, Glucose, starch,
glucose) plant structure cellulose, glycogen
Lipids Fatty acid + Glycerol Long-term energy Fats, oils, steroids,
storage, hormones, cell waxes
membrane
Proteins Amino acids (20 types) Enzymes, structure, Hemoglobin, insulin,
transport, hormones lactase
Nucleic Acids Nucleotide (base + Information storage and DNA, RNA
sugar + phosphate) transfer
Key Reactions
• Dehydration Synthesis (Condensation): Joining monomers by REMOVING water (H2O) - builds polymers
• Hydrolysis: Breaking polymers by ADDING water - breaks down macromolecules
3. CARBOHYDRATES
3.1 Types of Carbohydrates
Type Description Formula Examples
Monosaccharides Simple sugars; main fuel C6H12O6 Glucose, Fructose,
for cells (ATP) Galactose (all isomers!)
Disaccharides Two monosaccharides C12H22O11 Sucrose
joined by a GLYCOSIDIC (glucose+fructose),
bond (via condensation) Lactose
(galactose+glucose),
Maltose
(glucose+glucose)
Type Description Formula Examples
Polysaccharides Many sugar monomers Variable Starch (plants, energy),
linked; complex Glycogen (animals,
carbohydrates energy), Cellulose
(plants, structure)
Important: Glucose, Fructose, and Galactose all have formula C6H12O6 but DIFFERENT structural formulas
(isomers).
4. LIPIDS
4.1 Key Properties
• Hydrophobic (water-fearing) - do NOT mix with water
• Non-polar molecules
• Include: fats, waxes, steroids, oils
4.2 Types of Lipids
Triglycerides (Fats & Oils)
• Structure: 1 glycerol backbone + 3 fatty acid chains
• Function: Energy storage, body insulation, organ protection
Type Bonds State at Room Temp Found In
Saturated Fatty Acids Only SINGLE bonds Solid Animal fats, butter
between carbons
(maximum H atoms)
Unsaturated Fatty Acids Has at least one Liquid Plant oils, fish
DOUBLE bond between
carbons (fewer H
atoms)
Phospholipids
• Major component of CELL MEMBRANES (bilayer)
• Hydrophilic HEAD (polar, phosphate-containing) - faces water
• Hydrophobic TAILS (2 nonpolar fatty acid chains) - face inward
Steroids
• Carbon skeleton bent to form 4 FUSED RINGS
• Cholesterol = 'base steroid' from which body makes other steroids
• Estrogen & Testosterone are steroids
• Synthetic anabolic steroids = variants of testosterone; pose serious health risks
Waxes
• Single complex alcohol + long-chain fatty acid (ester linkage)
• Structural lipids - protective coatings (leaves, skin, hair)
• Extremely hydrophobic; barriers against water loss
5. PROTEINS
5.1 Structure
• Polymers of AMINO ACIDS (monomers)
• All proteins made from 20 different amino acids in different orders
• Amino acids joined by PEPTIDE BONDS
• Breakdown: Proteins --[hydrolysis]--> Peptides --[hydrolysis]--> Amino Acids
5.2 Protein Classification
Classification Types
By Structure Fibrous proteins, Globular proteins, Intermediate
proteins
By Composition Simple proteins, Conjugated proteins
By Function Structural, Enzymes, Hormones, Pigments,
Transport, Contractile, Storage, Toxins
5.3 Enzymes
Enzymes are biological CATALYSTS (most are proteins with tertiary/quaternary structure).
Property Description
Catalyst Speed up reactions without being permanently
changed
Specificity Each enzyme acts on specific substrate(s)
Reusable Not consumed in the reaction
Speed Up to 10^16 times faster than uncatalyzed rates!
Property Description
Naming Enzyme names end in -ase (e.g., Sucrase, Lactase,
Maltase)
Active Site The specific region where substrate binds
How Enzymes Work (Enzyme-Substrate Interaction)
• Substrate binds to enzyme's active site forming enzyme-substrate complex
• Enzyme places stress on substrate bonds
• Bond breaks, products are released
• Enzyme is FREE to bind new substrates
Restriction Enzymes (Special Example)
• Recognize specific base pair sequences in DNA called RESTRICTION SITES
• Cleave DNA by hydrolyzing the phosphodiester bond
• Cut between 3' carbon of first nucleotide and phosphate of next
• Fragment ends have 5' phosphates and 3' hydroxyls
• Naturally occur in bacteria to protect against viruses
• Over 400 restriction enzymes have been isolated
• Named with 3 italicized letters: EcoRI (from E. coli), HindIII (from H. influenzae), BamHI
• Many restriction sites are palindromes of 4-, 6-, or 8-base pairs
Application: EcoRI cuts both DNA strands leaving 'sticky ends' - used to create RECOMBINANT DNA
Applications of Recombinant DNA Technology
• Pharmaceutical products: Insulin production, vaccine sub-units (safer, cheaper)
• Gene therapy: Replacing defective genes using adeno/retrovirus as vectors
• Gene silencing: RNA interference (RNAi) using siRNA to degrade target mRNA
5.4 Hormones
Hormones are extracellular signaling molecules that carry information from SENSOR CELLS to TARGET CELLS to
coordinate metabolic processes.
Insulin
• Type: Protein hormone; secreted by BETA cells in Islets of Langerhans (Pancreas)
• Function: LOWERS blood glucose level (regulates blood sugar)
• Nature: Hydrophilic (Lipophobic) - acts via membrane receptors
• Main target cells: Skeletal Muscle & Adipose tissue
• Mechanism: Acts as a 'key' that unlocks glucose channels on cell membrane
• Deficiency: Insufficient insulin causes elevated blood sugar = DIABETES
Glucagon
• Type: Protein hormone; secreted by ALPHA cells (Pancreas)
• Function: RAISES blood glucose level from low to normal
• Mechanism: Acts in LIVER to stimulate breakdown of glycogen to glucose
• Glucagon is an Insulin Counter-Regulatory Hormone
• Stimulated by: Hypoglycemia (low blood sugar), high amino acid absorption after protein meal
• Inhibited by: High blood glucose level
6. NUCLEIC ACIDS: DNA & RNA
6.1 Overview
• Nucleic acids store hereditary information
• Contain information for making ALL the body's proteins
• Polymers of NUCLEOTIDES (monomers)
• Each nucleotide = Phosphate group + Sugar + Nitrogenous base
6.2 DNA vs RNA Comparison
Feature DNA RNA
Full Name Deoxyribonucleic Acid Ribonucleic Acid
Sugar Deoxyribose Ribose (extra -OH group)
Strands Double-stranded (double helix) Single-stranded
Bases A, T, G, C A, U, G, C (Uracil instead of
Thymine)
Location Nucleus (mainly) Nucleus and cytoplasm
Function Stores genetic information Transfers genetic info for protein
synthesis
6.3 DNA Base Pairing Rules
• Adenine (A) pairs with Thymine (T) - connected by 2 hydrogen bonds
• Cytosine (C) pairs with Guanine (G) - connected by 3 hydrogen bonds
• The two strands are held together by NON-COVALENT (hydrogen) bonds
6.4 Three Types of RNA
Type Full Name Function Analogy
mRNA Messenger RNA Carries genetic info Blueprint for protein
(blueprint) from DNA to
ribosomes
rRNA Ribosomal RNA Makes up ribosomes Construction site
(along with protein)
tRNA Transfer RNA Delivers amino acids to Delivery truck
ribosome during protein
synthesis
7. HUMAN GENOME PROJECT (HGP)
7.1 Key Facts
• AIM: Sequence the ENTIRE human genome and provide data FREE to the world
• Duration: 13 years of work; rough draft published in 2003
• Scale: Global collaboration - thousands of staff in institutes across the globe
• Output: Information on 3 TRILLION base pairs; sequences of ~30,000 genes
• Data Access: Free and open access through online public databases
7.2 Applications of HGP
Application Area Description
Molecular Medicine Identify fundamental causes of diseases; genetic
screening for rapid diagnosis; DNA tests detect
carriers; predict future disease likelihood
Waste Control & Environmental Cleanup Microbes with unique protein structures can be
used for practical waste cleanup purposes
Energy Sources Methane-producing microorganism genomics could
lead to cheaper fuel-grade methane production
Risk Assessment Assess individual risks from environmental
exposure to toxic agents; study cancer risk from
radiation
8. GENOMICS
8.1 Definition
Genomics is the study of WHOLE GENOMES of organisms, incorporating elements from genetics. It is a
COMPUTER-AIDED study of structure and function of entire genome.
8.2 Types of Genomics
Type Definition Main Concern Steps Involved
Structural Genomics Initial phase: Sequencing and High-res
sequencing and mapping the genome genetic/physical maps,
mapping the whole sequencing, determine
genome; determines proteins and 3D
structure of every structures
protein
Functional Genomics Study of how genes and Studying expression and Determine when/where
intergenic regions function of genome genes are expressed,
contribute to biological mutate genes to find
processes functions, find protein
interactions
Comparative Genomics Comparison of whole Evolutionary Compare genomes of
genomes from different relationships and related and unrelated
organisms conserved regions organisms to find
unique/shared elements
8.3 Key Points on Genomics
• Deals with mapping and sequencing of genes on chromosomes
• Genomic techniques are indispensable in plant breeding and genetics
• Applications: Genetic paternity/ancestry/compatibility tests, genetic fingerprinting, personalized
medicine, genetic disease risk assessment
9. SEQUENCE DATABASES (Biological Databases)
9.1 Why Databases?
• Genomic research generates ENORMOUS amounts of raw sequence data
• Sophisticated computational methods needed to manage the 'data deluge'
• Chief objective: Organize data in structured records for easy retrieval
9.2 Uses of Biological Databases
• Helps researchers study available data for their hypotheses
• Helps scientists understand biological phenomena
• Acts as storage of information
• Removes data redundancy
9.3 Types of Databases
1. Primary Databases (Archival Databases)
• Archives EXPERIMENTALLY DERIVED data submitted by scientists
• Data is un-curated (raw, directly from lab)
• Contains unique data from laboratory experiments
Database Type URL
GenBank Nucleic Acid sequences [Link]
nucleotide/
PDB (Protein Data Bank) Protein structure database [Link]
UniprotKB Protein sequence database [Link]
2. Secondary Databases
• Data DERIVED from analyzing primary data
• Draw from multiple sources (other databases, scientific literature)
• HIGHLY CURATED using computational algorithms + manual analysis
• Generate NEW knowledge from public record of science
Database Content
InterPro Protein families, motifs and domains -
[Link]
Pfam Protein family database with domain info
Prosite Database of protein sites, patterns and profiles
9.4 Search Tools
• Smith-Waterman: Similarity Search Tool (exact but slow)
• FASTA: Heuristic Search Tool (faster)
• BLAST: Heuristic Search Tool (fastest - most widely used)
10. BLAST TOOL
10.1 What is BLAST?
• BLAST = Basic Local Alignment Search Tool
• Local alignment algorithm for aligning multiple sequences
• Finds similarity/dissimilarity among various species
• Heuristic method = faster and efficient but relatively less sensitive
• Compares nucleotide OR protein sequences
• Calculates the STATISTICAL SIGNIFICANCE of matches
10.2 BLAST Algorithm Steps
Step Action
1. Split query Split query into overlapping words of length W
(called W-mers)
2. Find neighborhood Find a 'neighborhood' of similar words for each W-
mer
3. Hash table lookup Lookup each neighborhood word in a hash table to
find location in database (these are 'seeds')
4. Extend seeds Extend seeds until alignment score drops below
threshold X
5. Report Report matches with overall highest scores (High-
scoring Segment Pairs - HSPs)
10.3 BLAST Program Types
Program Query Database Use Case
blastn (nucleotide blast) Nucleotide Nucleotide Find similar nucleotide
sequences
blastp (protein blast) Amino acid (protein) Protein Find similar protein
sequences
blastx Nucleotide (translated Protein Find translation
all frames) products of unknown
nucleotide seq
tblastn Protein Nucleotide (translated Find DNA encoding a
all frames) known protein
tblastx Nucleotide (6 frame Nucleotide (6 frame Most computationally
translation) translation) intensive; compare 2
nucleotide seqs at
protein level
10.4 Understanding BLAST Output
• Results shown as hit tables in DECREASING order of matched score
• Each hit shows: accession number, title, query coverage, identity, score, E-value
• E-value (Expect Value): Number of alignments with scores equivalent to or BETTER THAN S expected by
CHANCE in database search
• LOWER E-value = MORE significant result (better match)
• E-value close to 0 = excellent match; E-value near 1 = could be random
• Alignment scores use color coding: Black (<40, BAD) to Red (>=200, GOOD)
10.5 How to Run BLAST
Step Action
1 Enter the query sequence (FASTA format or
accession number)
2 Select a job title
3 Select the database to search
4 Select the BLAST algorithm to use
5 Adjust algorithm parameters as necessary
6 Click BLAST to start the search
QUICK REVISION: KEY TERMS
Term Definition
Atom Basic unit of all matter
Cation Positively charged ion (donated electrons)
Anion Negatively charged ion (received electrons)
Monomer Small repeating unit that builds polymers
Polymer/Macromolecule Large molecule built from monomers
Glycosidic Bond Bond joining monosaccharides in
disaccharides/polysaccharides
Hydrophilic Water-loving (polar)
Hydrophobic Water-fearing (non-polar)
Term Definition
Catalyst Substance that speeds up reaction without being
consumed
Active Site Region on enzyme where substrate binds
Restriction Site Specific DNA base pair sequence recognized by
restriction enzyme
Recombinant DNA DNA combining sequences from two different
organisms
siRNA Short interfering RNA used in gene silencing
Genome Complete set of genetic information of an organism
Genomics Computer-aided study of whole genomes
E-value Statistical measure of BLAST alignment significance;
lower = better
Palindrome Restriction site that reads the same on both strands
(5'→3')
HSP High-Scoring Segment Pair - best alignment region
in BLAST
Hypoglycemia Low blood glucose level (triggers glucagon release)
Diabetes Condition caused by insufficient insulin; elevated
blood sugar
BIOCHEMISTRY & DISEASE (Exam Favorite!)
Disease Cause Related Biomolecule
Scurvy Deficiency of Vitamin C Vitamin
Rickets Deficiency of Vitamin D Vitamin
Atherosclerosis Genetic, dietary, environmental Lipids
factors
Cystic Fibrosis Mutation in CFTR protein gene Protein
(chloride ion transport)
Cholera Exotoxin of Vibrio cholerae Protein (toxin)
Diabetes Mellitus Type I Genetic/environmental factors Protein (hormone)
causing insulin deficiency
Phenylketonuria (PKU) Mutation in gene coding Protein (enzyme)
phenylalanine hydroxylase
Disease Cause Related Biomolecule
Sickle Cell Anemia Mutation causing abnormal Protein
hemoglobin structure
Genetic Diseases (general) Mutations in nucleic acid Nucleic Acid
sequences
Good luck on your exam! You've got this!