Introduction to Databases
(Biological Databases and
Resources)
Computational Biochemistry and Molecular
Biology (BCH 208)
Biological Databases and
Resources
• A database is an easily accessible collection of structured information
(data)stored electronically in a computer system
• A biological database is a digital repository of biological information (genes,
proteins, diseases, literature, structure) in an organises and searchable format
• Primary databases store raw sequence data while secondary databases provide
information on the annotation of the sequence data
• Databases can be
• Public or Private
• Generalised or Specialised
• Curated or Non-curated
Biological Databases and
Resources
• Genomic Databases – International databases that collaborate to collect store and share raw
nucleotide sequence data (DNA/RNA) submitted by researchers worldwide. They are the
foundational resources for genomics
• The public genomic databases that work together under the International Nucleotide
Sequence Database Collection (INSDC) are
• National Centre for Biotechnology information (NCBI) - GenBank
• European Molecular Biology Laboratory (EMBL) - European Neulotide Archive (ENA)
• DDBJ – DNA Data Bank of Japan
• NCBI and EMBL-EBI are the most popular databases for extracting data
• The organisations exchange data daily to ensure they hold the same collection of sequence
data
• INSDC covers the full spectrum of raw reads, alignment, assembly, and functional annotation
• Each database has their own accession number and tools
GenBank
• GebBank is the genetic sequence database for the
National institute of health (NIH)
• It contains an annotated collection of all publicly available
DNA sequences
• The database is updated bimonthly
• The link provides more insight to information found in
GenBank (
[Link]
NCBI
• The National Centre for Biological Information (NCBI) is a United States
government funded bioinformatics centre established in 1988 that provides free
access to biological data and tools to help scientists understand molecular and
genetic information
• It supports research in genomics, molecular biology, medicine and
computational biology
• It is currently the worlds leading platform for genomics, proteomics and medical
data integration
• It is updated daily, using automated tools and human curation to check data
quality.
• Website [Link]
Blue wavy box – Resources list; Green dashed box – Central window indicates main
features of the website; Yellow box – Indicates popular resources
EMBL - EBI
• EBI – European Bioinformatics Institute
• EMBL –EBI was founded in 1994 to provide biological data and tools for the
scientific community. They also develop cutting-edge research in
bioinformatics
• The database is currently maintained by the Wellcome Genome Campus in
the United Kingdom
• Website [Link]
• It has a better annotation quality and more advanced bioinformatic tools
than NCBI.
• It is less popular and has a more complex user interface than NCBI
Accessing Data
• You can use biological databases to
• Identify a gene sequence
• Identify variants in the sequence
• Compare your sequence to others
• Identify similar sequences
• Find diseases associated with variation in your gene of interest
Accessing Data
• Biological databases are easy to use as tools are provided
to link all annotations to a particular query
• NCBI & EBI allows scientist to search all available databases with a
single query
• It is important that the researcher provides clear
information in the query box.
• A search for Insulin or using its symbol INS would yield significant
results but a search for gene associated with low blood sugar or
hypoglycaemia will yield limited results
Popular NCBI Databases
• Gene – Provides information on gene annotation
• PubMed – Biomedical literature
• Nucleotide – DNA sequence data
• SNP – Single Nucleotide Polymorphism
• Protein – Protein sequence data
• RefSeq – Comprehensive, integrated, well- annotated set of (genomic,
transcript & protein) reference sequences
• OMIM (Online Mendelian Inheritance in Man) – Human genes and
genetic phenotypes
• ClinVar – Genomic variation and relationship to human health
Reference Sequence (RefSeq)
• RefSeq – A RefSeq (reference sequence) is a curated ,
high – quality, and non – redundant set of annotated DNA,
RNA, and protein sequences.
• They are the baseline or reference for
• Genetic research
• Disease mutation analysis
• Functional studies
• Clinical diagnostics
Accession Number
• An accession number is a unique identifier assigned to a biological
sequence submitted to a public database.
• It is easier to locate, retrieve, and cite sequences with known accession
numbers
• The accession number is always the same but the version number changes as
the version is updated
• INS NM_000207.3
• NM_ indicates a RefSeq mRNA
• 000207 is the unique ID
• 3 is the version number (this changes if the sequence is updated)
Accession Number
Accession Type Description
prefixes
NC_ Known RefSeq Complete genome or chromosome
NG_ Known RefSeq Genomic region
NM_ Known RefSeq mRNA
NP_ Known RefSeq Protein
XM_ Predicted/Model RefSeq mRNA
XP_ Predicted/Model RefSeq Protein
XR_ Predicted/Model RefSeq RNA
Popular EMBL - EBI Resources
• Ensembl – High-quality annotation data
• Uniprot – Universal resource for protein sequence and
functional annotation data
• PDBe –Protein Data Bank of Europe
• InterPro – Database of Protein Families, Domains and
Conserved sites
Specialised Databases
• Specialised databases focus on a narrow area of biology and
provide highly detailed or unique information on a
particular organism or infectious agent. Usually,
• Manually curated
• Specialised resources
• Specific tool for mining data
• Examples include - Malaria
• (MalariaGEN [Link]
• PlasmoDB ([Link]