0% found this document useful (0 votes)
2 views11 pages

Genomics Bioinformatics Notes

The document is a comprehensive guide on genomics and bioinformatics, covering topics such as genome mapping, sequence data, biological databases, BLAST, phylogenetic trees, comparative genomics, and systems biology. It explains key concepts, techniques, and applications in each area, emphasizing the importance of genomics in fields like medicine, agriculture, and evolutionary biology. The guide serves as a reference for beginners to intermediates in understanding the structure, function, and analysis of genomes.

Uploaded by

nas huss
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views11 pages

Genomics Bioinformatics Notes

The document is a comprehensive guide on genomics and bioinformatics, covering topics such as genome mapping, sequence data, biological databases, BLAST, phylogenetic trees, comparative genomics, and systems biology. It explains key concepts, techniques, and applications in each area, emphasizing the importance of genomics in fields like medicine, agriculture, and evolutionary biology. The guide serves as a reference for beginners to intermediates in understanding the structure, function, and analysis of genomes.

Uploaded by

nas huss
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Genomics & Bioinformatics —

Complete Study Notes


A Beginner-to-Intermediate Reference Guide

Genomics · Genome Mapping · Sequence Data · BLAST · Biological Databases · Phylogenetic Trees ·
Comparative Genomics · Systems Biology
Table of Contents
1. Introduction to Genomics
2. Genome Mapping
3. Sequence Data (Sequencing & File Formats)
4. Biological Databases
5. BLAST (Basic Local Alignment Search Tool)
6. Phylogenetic Trees
7. Comparative Genomics
8. Systems Biology
1. Introduction to Genomics

1.1 What is Genomics?


Genomics is the branch of molecular biology concerned with the structure, function, evolution, and mapping of
genomes — the complete set of DNA (or RNA, in some viruses) within an organism, including all of its genes
and non-coding sequences. Unlike classical genetics, which studies one gene at a time, genomics studies
genes and their interactions collectively, using high-throughput molecular techniques and computational
analysis.

1.2 Branches of Genomics


• Structural genomics: Determines the 3-D structure of every protein encoded by a genome; also covers
physical mapping and whole-genome sequencing.
• Functional genomics: Studies gene expression, gene function, and gene product interactions using
techniques such as microarrays, RNA-seq, and knockout studies.
• Comparative genomics: Compares genome sequences of different species/strains to understand
evolution, gene function, and conserved elements (detailed in Section 7).
• Epigenomics: Studies heritable changes in gene expression that do not involve changes in the DNA
sequence itself (e.g., DNA methylation, histone modification).
• Metagenomics: Studies genetic material recovered directly from environmental samples, capturing entire
microbial communities.

1.3 Why Genomics Matters


• Identifying disease-causing genes and mutations (e.g., BRCA1/2, CFTR).
• Personalized/precision medicine — tailoring treatment based on an individual's genome.
• Agricultural improvement — crop and livestock breeding using marker-assisted selection.
• Evolutionary biology — tracing relationships among species.
• Drug discovery and vaccine development (e.g., rapid genome sequencing of SARS-CoV-2).
Note: The Human Genome Project (1990–2003) was a landmark international effort that produced the first complete
sequence of the human genome (~3.2 billion base pairs), laying the foundation for modern genomics.
2. Genome Mapping

2.1 What is Genome Mapping?


Genome mapping is the process of determining the location of genes or genetic markers and the distances
between them on a chromosome. It creates a 'roadmap' of the genome that helps locate genes responsible
for traits or diseases and assists in genome sequencing/assembly.

2.2 Types of Genome Maps


(a) Genetic (Linkage) Maps
Based on the frequency of recombination (crossing over) between genetic markers during meiosis. Distances
are measured in centimorgans (cM); 1 cM ≈ 1% recombination frequency. Markers used include RFLPs,
microsatellites (SSRs), and SNPs. Genes that are close together on a chromosome are inherited together
more often (linked) than genes far apart.

(b) Physical Maps


Based on actual physical/molecular distance between markers, measured in base pairs (bp, kb, Mb).
Constructed using techniques such as restriction mapping, Fluorescence In Situ Hybridization (FISH),
Sequence Tagged Site (STS) mapping, and contig assembly. Physical maps are more precise than genetic
maps and form the scaffold for whole-genome sequencing.

(c) Sequence Maps


The most detailed map — the complete nucleotide sequence of the genome, produced by aligning and
assembling millions of short sequencing reads into a continuous sequence.

2.3 Common Mapping Techniques & Markers


Technique/Marker Basis Use

RFLP (Restriction Fragment Length Polymorphism)


Variation in restriction enzyme cut sites Early linkage mapping, forensic ID

RAPD (Random Amplified Polymorphic


PCR DNA)
with random primers Genetic diversity studies

AFLP (Amplified Fragment Length Polymorphism)


Restriction + selective PCR amplificationHigh-resolution fingerprinting

SSR / Microsatellites Repeats of short DNA motifs (1–6 bp) Highly polymorphic linkage markers

SNP (Single Nucleotide Polymorphism)


Single base-pair variation Dense modern maps, GWAS studies

STS (Sequence Tagged Site) Short unique sequence with known location
Physical map landmarks

FISH (Fluorescence In Situ Hybridization)


Fluorescent probe hybridization to chromosome
Direct visual gene localization

2.4 Applications of Genome Mapping


• Locating disease genes (positional cloning).
• Marker-assisted selection in plant/animal breeding.
• Serving as a scaffold for genome sequencing and assembly.
• Comparative studies of genome organization across species.
3. Sequence Data

3.1 Types of Biological Sequence Data


• DNA sequences: Strings of nucleotides (A, T, G, C) representing genomic or gene sequences.
• RNA sequences: Transcribed sequences (A, U, G, C); include mRNA, tRNA, rRNA, and non-coding
RNAs.
• Protein sequences: Strings of amino acids (20 standard residues) representing translated gene products.

3.2 How Sequence Data is Generated


• Sanger sequencing (first-generation): Chain-termination method using dideoxynucleotides; accurate but
low throughput; used for smaller-scale work and validation.
• Next-Generation Sequencing/NGS (second-generation): e.g., Illumina; massively parallel sequencing of
millions of short fragments; low cost per base, high throughput.
• Third-generation sequencing: e.g., PacBio SMRT and Oxford Nanopore; produce very long reads in real
time, useful for resolving repetitive regions and structural variants.

3.3 Common Sequence File Formats


Format Content Typical Use

FASTA (.fasta/.fa) Header line (>) + raw sequence Storing/sharing DNA, RNA, or protein sequences

FASTQ (.fastq) Sequence + per-base quality scores Raw NGS reads

GenBank (.gb) Sequence + rich annotation (features, references)


NCBI's annotated record format

GFF/GTF Tab-delimited genome feature/annotation coordinates


Gene models, exon/intron structure

SAM/BAM Aligned reads mapped to a reference genomePost-alignment storage & analysis

VCF (Variant Call Format) Genomic variants (SNPs, indels) vs. a reference
Variant calling & GWAS

PDB (.pdb) 3-D atomic coordinates of macromolecules Protein/nucleic acid structure data

3.4 Sequence Alignment


Sequence alignment arranges two or more sequences to identify regions of similarity that may reflect
functional, structural, or evolutionary relationships.

• Pairwise alignment: Compares exactly two sequences (e.g., Needleman-Wunsch for global alignment,
Smith-Waterman for local alignment).
• Multiple Sequence Alignment (MSA): Aligns three or more sequences simultaneously (e.g.,
ClustalW/Clustal Omega, MUSCLE, MAFFT); essential for phylogenetics and conserved-motif discovery.
4. Biological Databases

4.1 Why Databases Matter


Biological databases are organized, searchable repositories of biological data (sequences, structures,
pathways, literature) that allow researchers worldwide to store, retrieve, compare, and analyze data. They are
broadly classified as primary (original experimental data, e.g., raw sequences) or secondary
(curated/derived data, e.g., annotated protein families).

4.2 Major Databases


Database Type Key Content

GenBank (NCBI, USA) Primary, nucleotide Publicly available DNA/RNA sequences; part of INSDC

EMBL-Bank (EBI, Europe) Primary, nucleotide European nucleotide sequence archive; part of INSDC

DDBJ (Japan) Primary, nucleotide Japanese nucleotide sequence archive; part of INSDC

UniProt / Swiss-Prot Primary + curated, protein


Protein sequence and functional annotation

PDB (Protein Data Bank) Primary, structure 3-D structures of proteins & nucleic acids

Ensembl Secondary, genome Annotated vertebrate/eukaryotic genome browser

UCSC Genome Browser Secondary, genome Genome assemblies with tracks for genes, variants, etc.

KEGG Secondary, pathway Metabolic and signaling pathway maps

OMIM Secondary, disease Catalog of human genes and genetic disorders

PubMed Literature Biomedical research abstracts and citations

Note: GenBank, EMBL-Bank, and DDBJ together form the INSDC (International Nucleotide Sequence Database
Collaboration) and mirror each other's data daily.
5. BLAST — Basic Local Alignment Search Tool

5.1 What is BLAST?


BLAST is a widely used bioinformatics algorithm/tool for comparing a query biological sequence against a
database of sequences to find regions of local similarity. It estimates the statistical significance of matches,
helping identify homologous genes, predict function, and study evolutionary relationships.

5.2 Types of BLAST


Program Query Type Database Type Use

BLASTn Nucleotide Nucleotide Compare DNA/RNA sequences directly

BLASTp Protein Protein Compare protein sequences directly

BLASTx Nucleotide (translated) Protein Find protein matches for a DNA/RNA query

tBLASTn Protein Nucleotide (translated) Find a protein's matches within a DNA database

tBLASTx Nucleotide (translated) Nucleotide (translated) Compare translated DNA vs translated DNA

5.3 How BLAST Works (Algorithm Steps)


• 1. Seeding: The query sequence is broken into short 'words' (e.g., 11 bp for nucleotides, 3 residues for
proteins).
• 2. Seed matching: These words are compared to the database to quickly find short exact/high-scoring
matches (seeds).
• 3. Extension: Each seed match is extended in both directions to build a longer alignment, stopping when
the score starts to drop (High-scoring Segment Pairs, HSPs).
• 4. Scoring & statistics: Alignments are scored using substitution matrices (e.g., BLOSUM62 for proteins)
and evaluated for statistical significance.

5.4 Key Output Parameters


Term Meaning

E-value Expected number of chance matches with this score in a database of this size; lower E-value = mor

Bit score Normalized alignment score, independent of database size; higher = better match

Percent identity Percentage of identical residues/bases in the aligned region

Query coverage Percentage of the query sequence covered by the alignment

5.5 Applications of BLAST


• Identifying an unknown gene/protein by similarity to known sequences.
• Finding homologous genes across species (orthologs/paralogs).
• Designing PCR primers and checking sequence specificity.
• Supporting genome annotation and evolutionary studies.
6. Phylogenetic Trees

6.1 What is a Phylogenetic Tree?


A phylogenetic tree is a branching diagram that represents the evolutionary relationships among a set of
organisms or genes, based on similarities and differences in their genetic or physical characteristics. It is built
from sequence alignments and reflects hypothesized common ancestry.

6.2 Basic Terminology


• Node: A point representing either an ancestor (internal node) or an existing taxon (terminal/leaf node).
• Branch (edge): Connects nodes and represents the evolutionary lineage between them; branch length
often indicates the amount of evolutionary change.
• Root: The common ancestor of all taxa in the tree.
• Clade: A group consisting of an ancestor and all of its descendants (a monophyletic group).
• Taxon/OTU (Operational Taxonomic Unit): The individual organism, species, or sequence being
compared.

6.3 Types of Trees


• Rooted tree: Has a single ancestral node from which all other nodes descend, showing direction of
evolution; requires an outgroup for rooting.
• Unrooted tree: Shows relationships among taxa without specifying a common ancestor or evolutionary
direction.
• Cladogram: Shows branching order only; branch lengths do not represent evolutionary distance/time.
• Phylogram: Branch lengths are proportional to the amount of evolutionary change (genetic distance).
• Chronogram (timetree): Branch lengths are proportional to actual geological/evolutionary time.

6.4 Methods of Tree Construction


(a) Distance-based methods
• UPGMA (Unweighted Pair Group Method with Arithmetic mean): Simple clustering method; assumes a
constant molecular clock (equal rate of evolution in all lineages).
• Neighbor-Joining (NJ): Does not assume a constant rate of evolution; widely used, fast, and works well
with large datasets.

(b) Character-based methods


• Maximum Parsimony (MP): Finds the tree that requires the fewest evolutionary changes (mutations) to
explain the observed data.
• Maximum Likelihood (ML): Finds the tree that has the highest probability of producing the observed
sequence data, given a model of evolution.
• Bayesian Inference: Uses probability distributions and prior information to estimate the most probable
tree(s); often reported with posterior probability support values.

6.5 Tree Evaluation & Tools


• Bootstrapping: A resampling technique used to assess the confidence/support of branches in a tree
(values closer to 100% indicate stronger support).
• Common software: MEGA, PHYLIP, RAxML, MrBayes, ClustalW/Clustal Omega, iTOL (for visualization).

6.6 Applications
• Tracing the evolutionary history of species or genes.
• Studying the spread and mutation of pathogens (e.g., viral outbreak tracking).
• Classifying organisms and resolving taxonomic uncertainties.
• Identifying gene duplication events (paralogs) versus speciation events (orthologs).
7. Comparative Genomics

7.1 What is Comparative Genomics?


Comparative genomics is the study of the similarities and differences in the genome structure, organization,
and content of different species, strains, or individuals. By comparing genomes, researchers can identify
conserved (functionally important) regions and lineage-specific (evolutionarily divergent) regions.

7.2 Key Concepts


• Orthologs: Genes in different species that evolved from a common ancestral gene via speciation; usually
retain the same function.
• Paralogs: Genes related by duplication within a genome; may evolve new functions over time.
• Synteny: Conservation of gene order/blocks on chromosomes across different species, reflecting shared
ancestry.
• Conserved sequences: Regions of DNA/protein that remain similar across species because they are
functionally essential (e.g., coding regions, regulatory elements).

7.3 Techniques Used


• Whole-genome alignment: Aligning entire genomes to detect large-scale similarities, rearrangements,
insertions, and deletions.
• Synteny mapping: Comparing gene order across species to infer chromosomal evolution.
• Ortholog/paralog identification: Using sequence similarity tools (e.g., BLAST, OrthoFinder) to classify
gene relationships.
• Comparative annotation: Using known gene models from one species to help annotate genes in a
related, less-studied species.

7.4 Applications
• Identifying functionally important, evolutionarily conserved genes and regulatory elements.
• Understanding the genetic basis of species-specific traits and adaptations.
• Improving genome annotation for newly sequenced organisms by comparison to well-studied model
organisms.
• Studying the evolution of gene families and genome architecture.
• Supporting drug target discovery by comparing pathogen and host genomes.
8. Systems Biology

8.1 What is Systems Biology?


Systems biology is an interdisciplinary field that studies biological systems as integrated, interacting networks
of genes, proteins, and metabolites rather than analyzing individual components in isolation. It combines
experimental biology, computational modeling, and mathematics to understand how complex behaviors (e.g.,
cell signaling, disease states) emerge from these interactions.

8.2 Integration of 'Omics' Layers


• Genomics: The complete set of genes/DNA.
• Transcriptomics: The complete set of RNA transcripts (gene expression).
• Proteomics: The complete set of proteins expressed by a cell/tissue.
• Metabolomics: The complete set of small-molecule metabolites.
• Epigenomics: The complete set of epigenetic modifications (e.g., methylation patterns).
• Systems biology integrates data across these layers to build comprehensive models of cellular behavior.

8.3 Approaches
• Top-down approach: Starts with large-scale experimental data (e.g., gene expression across conditions)
and works backward to infer the underlying network/interactions.
• Bottom-up approach: Starts with known individual components (genes, proteins, reactions) and builds
mathematical/computational models to predict system-level behavior.

8.4 Tools & Methods


• Network modeling: Representing genes/proteins as nodes and their interactions as edges (e.g., gene
regulatory networks, protein-protein interaction networks).
• Computational simulation: Mathematical modeling (e.g., differential equations, Boolean networks) to
simulate system dynamics over time.
• Databases & software: KEGG (pathways), STRING (protein interactions), Cytoscape (network
visualization), COBRA toolbox (metabolic modeling).

8.5 Applications
• Understanding complex diseases (cancer, diabetes) as network-level disruptions rather than single-gene
defects.
• Identifying novel drug targets by studying pathway interactions.
• Synthetic biology — designing and engineering new biological circuits/pathways.
• Personalized medicine — integrating multi-omics data from patients to guide treatment decisions.

End of Notes — This document provides a foundational overview of each topic. For deeper study, consult primary
literature, NCBI/EBI documentation, and standard bioinformatics textbooks (e.g., Lesk's Introduction to
Bioinformatics, Pevsner's Bioinformatics and Functional Genomics).

You might also like