0% found this document useful (0 votes)
27 views40 pages

Understanding GI Numbers in Bioinformatics

The document discusses the distinction between primary and secondary bioinformatics databases, detailing various identifiers such as locus names, accession numbers, and GI numbers. It also covers different protein databases, alignment methods, and the significance of multiple sequence alignments in protein analysis. Additionally, it introduces tools like BLAST and PSI-BLAST for protein-protein comparisons and highlights the importance of conserved regions in protein sequences.

Uploaded by

Mohamed Hasan
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
27 views40 pages

Understanding GI Numbers in Bioinformatics

The document discusses the distinction between primary and secondary bioinformatics databases, detailing various identifiers such as locus names, accession numbers, and GI numbers. It also covers different protein databases, alignment methods, and the significance of multiple sequence alignments in protein analysis. Additionally, it introduces tools like BLAST and PSI-BLAST for protein-protein comparisons and highlights the importance of conserved regions in protein sequences.

Uploaded by

Mohamed Hasan
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd

BASIC BIOINFORMATICS

AHMED GAMEEL
MSC, BIOTECHNOLOGY-BIOINFORMATICS
PRIMARY AND SECONDARY DATABASE

• There is an important distinction between primary (archival)


And secondary (curated) databases.
• The primary databases represent experimental results (with
some interpretation)
But are not a curated review.
• Curated reviews are found in the secondary databases.
NCBI FLAT FILE
LOCUS NAME

• LOCUS NAME :

• THE LOCUS NAME IS RESTRICTED TO TEN OR FEWER NUMBERS AND


UPPERCASE LETTERS.
• THE FIRST THREE LETTERS OF THE NAME WERE AN ORGANISM CODE AND
THE REMAINING LETTERS A CODE FOR THE GENE .
• THERE IS A PROBLEM WITH THE LOCUS NAME THAT THE LOCUS NAME MAY
CHANGES TO A NEW ONE BECAUSE THE NEW DISCOVERIS FOR GENE OR THE
REGION, WHICH MAK A PROBLEMS IN IDENTFING AND RETRIEVING SUCH
SEQUENCES.
ACCESSION NUMBER

• IT INTENTIONALLY CARRIES NO BIOLOGICAL MEANING, TO ENSURE THAT IT


WILL REMAIN STABLE.
• ITS CONSISTED OF ONE UPPERCASE LETTER FOLLOWED BY FIVE DIGITS.
• NEW ACCESSIONS CONSIST OF TWO UPPERCASE LETTERS FOLLOWED BY SIX
DIGITS.
• IT IS ALSO HAVE SOME PROBLEMS ON RETRIEVING , FOR EXAMPLE, IF THERE
IS ANY UBDATES AT THE SEQUENCE LIKE ADDITION OF NEW NUCLEOTIDES AT
SPECIFIC REGION, WHEN THE USER RETRIEVE THE SEQUENCE WITH THE
ACCESSION NUMBER HE WILL NOTICE DIFFERENCES BETWEEN THE OLD AND
NEW RETRIEVED SEQUENCE.
GI NUMBER

• GENINFO IDENTIFIERS (GI):


• NCBI BEGAN ASSIGNING (GI) TO ALL SEQUENCES PROCESSED INTO ENTREZ, INCLUDING
NUCLEOTIDE SEQUENCES FROM DDBJ/EMBL/GENBANK, THE PROTEIN SEQUENCES FROM
THE TRANSLATED CDS FEATURES, PROTEIN SEQUENCES FROM SWISS-PROT, PIR, PRF,
PDB, PATENTS, AND O
• THE GI NUMBER SERVES THREE MAJOR PURPOSES:
• IT PROVIDES A SINGLE IDENTIFIER ACROSS SEQUENCES FROM MANY SOURCES.
• IT PROVIDES AN IDENTIFIER THAT SPECIFIES AN EXACT SEQUENCE.
• IT IS STABLE AND RETRIEVABLE.
[Link] COMBINED IDENTIFIER

• THE MEMBERS OF THE INTERNATIONAL NUCLEOTIDE SEQUENCE DATABASE


COLLABORATION
• (GENBANK, EMBL, AND DDBJ) INTRODUCED A ‘‘BETTER’’ SEQUENCE
IDENTIFIER, ONE THAT COMBINES AN ACCESSION (WHICH IDENTIFIES A
PARTICULAR SEQUENCE RECORD) WITH A
• VERSION NUMBER (WHICH TRACKS CHANGES TO THE SEQUENCE ITSELF).
ACCESSION NUMBERS ON PROTEIN
SEQUENCES

• The international sequence database collaborators also


started assigning accession-version numbers to protein
sequences within the records.
REFSEQ

• The NCBI refseq project provides a curated, nonredundant set of reference


sequence standards for naturally occurring biological molecules, ranging
from chromosomes to transcripts to proteins.
• Refseq identifiers are in [Link] form but are prefixed with NC
(chromosomes), NM (mRNAs), NP (proteins), or NT (constructed genomic
contigs). The NG prefix will be used for genomic regions or gene clusters.
• Refseq records are a stable reference point for functional annotation, point
mutation analysis, gene expression studies, and polymorphism discovery.
FASTA FORMAT

• FASTA format contains a definition line and sequence characters.


• The definition line starts with a right angle bracket (>) and is
usually followed by the sequence identifiers in a parsable form, as
in this example:
>Gi|2352912|gb|af012433.1|hsddt2
• The remainder of the definition line, which is usually a title for the
sequence.
PROTEIN DATABASES
• Non-redundant protein sequences (nr)
• Kitchen-sink:
• Translations of genbank coding sequences (CDS)
• Refseq proteins
• PDB (RCSB protein data bank - 3d-structure)
• Swissprot
• Protein information resource (PIR)
• Protein research foundation (japanese DB)
• Reference proteins (refseq_protein)
• NCBI reference sequences: comprehensive,
integrated, non-redundant, well-annotated set
of sequences
• Swissprot protein sequences (swissprot)
• Swiss-prot: european protein database (no
PROTEIN DATABASES
• Patented protein sequences (pat)
• Patented sequences
• Protein data bank proteins (pdb)
• Sequences from RCSB protein data bank with
experimentally determined structures
• Environmental samples (env_nr)
• Protein sequences from environmental samples
(not associated with known organism)
PATTERNS DATABASE: PROSITE

• Patterns database: prosite


• Prosite is a database containing patterns and proles:
• WEB access: [Link]
• well documented.
• Easy to test new patterns.
• Patterns length typically around 10-20 aa.
PATTERNS DATABASE: PROSITE
PSSM DATABASES: PRINTS

• position-specific scoring matrix (PSSM) Databases: prints


• Collection of conserved motifs used to characterize a protein.
• Uses ngerprints (conserved motif groups).
• Very good to describe sub-families.
• Http://[Link]/dbbrowser/prints.
• Blocks is another pssms database similar to prints
• (Http://[Link]/).
PSSM DATABASES: PRINTS
PROTEIN DOMAIN DATABASES: PFAM

• Good links to structure, taxonomy.


• Http://[Link]/pfam. Domains and families
PROTEIN DOMAIN DATABASES: PFAM
PROTEIN DOMAIN DATABASES: PROSITE

• COLLECTION OF MOTIFS, PROTEIN DOMAINS, AND FAMILIES.


• HIGH QUALITY DOCUMENTATION.
• HTTP://[Link]/PROSITE.
PROSITE
PROTEIN DOMAIN DATABASES: SMART

• Collection of protein domains.


• Excellent graphic interface.
• Excellent taxonomic information.
• Easy to search meta-motifs.
• Http://[Link]
PROTEIN DOMAIN DATABASES: SMART
PROTEIN DOMAIN DATABASES: PRODOM

• COLLECTION OF PROTEIN MOTIFS OBTAINED AUTOMATICALLY USING PSI-


BLAST.
• VERY HIGH THROUGHPUT ... BUT NO ANNOTATION.
• HTTP://[Link]/PRODOM/DOC/[Link].
PROTEIN GROUPS
SIMILARITY VS HOMOLOGY

• Similarity is an observable quantity that might be expressed as,


say, percent identity or some other suitable measure.
• Homology refers to a conclusion drawn from these data that two
genes share a common evolutionary history.
• Genes either are or
• are not homologous—there are no degrees for homology as
there are for similarity.
GLOBAL VS LOCAL ALIGNMENT

• A global alignment is an alignment that essentially


spans the full extents of the input sequences is called.
• local alignment is an alignment consists of paired
subsequences
• that may be surrounded by residues that are
completely unrelated.
ALIGNMENT
AACGTTTCCAGTCCAAATAGCTAGGC
===--=== =-===-==-======
AACCGTTC TACAATTACCTAGGC

HITS(+1): 18
MISSES (-2): 5
GAPS (EXISTENCE -2, EXTENSION -1): 1 LENGTH:
3
SCORE = 18 * 1 + 5 * (-2) – 2 – 2 = 6
SUBSTITUTION MATRIX

• It is well known that certain amino acids can substitute


easily for one another in related proteins because of
their similar physicochemical properties.
• Examples of these ‘‘conservative substitutions’’ include
isoleucine for valine (both small and hydrophobic) and
serine for threonine (both polar)
• identical amino acids should be given greater value
than substitutions, but conservative substitutions
should also be greater than non-conservative changes.
SCORING MATRIX

SCORING MATRIX:
• PAM: accepted point mutation
• Empirically derived chance a substitution will be accepted, based on
closely related proteins
• Higher PAM numbers correspond to greater evolutionary distance
• BLOSUM: blacks substitution matrix
• Another empirically derived matrix, based on more distantly related
proteins
• Lower BLOSUM numbers correspond to greater evolutionary distance
PROTEIN MULTIPLE SEQUENCE ALIGNMENT

• A multiple sequence alignment is simply an alignment that


contains more than two sequences
• The alignment of multiple sequences is a method of choice to
detect conserved
regions in protein or DNA sequences
• The aim of protein alignment is to assign function to the newly
determined protein.
• Multiple sequence alignment is useful in prediction of secondary
structure.
• The conserved region in MSA are usually associated with:
 Signals (promoters, signatures for phosphorylation, cellular location, ...);
 Structure (correct folding, protein-protein interactions...);
 Chemical reactivity (catalytic sites,... ).

• The effective analysis if protein MSA is very important for many protein
annotation like the most important residue for activity and the amino acid
required for stability.
PROTEIN ALIGNMENT
HIERARCHICAL METHODS

• All pairs Of sequences in the set to be aligned are compared by a pairwise method of sequence

Comparison.
• This provides a set of pairwise similarity scores for the sequences that Can be fed into a cluster
analysis or tree calculating program.
• The tree is calculated to place more similar pairs of sequences closer together on the tree than
sequences that are less similar.
• The multiple alignment is then built by starting with the pair of sequences that is most similar
and aligning them and then aligning the next most similar pair, and so on.
• CLUSTAL W is an example for hierarchical method for protein MSA.
PROTEIN BALST

BALST BASIC LOCAL ALIGNMENT SEARCH TOOL


PSI-BLAST
• PROTEIN-PROTEIN BLAST
• POSITION-SPECIFIC ITERATED BLAST
• FINDS MORE DISTANTLY RELATED MATCHES
• ITERATES: INITIAL SEARCH RESULTS PROVIDE INFORMATION
ON “ALLOWED” MUTATIONS; SUBSEQUENT SEARCHES USE
THESE TO CREATE CUSTOM SUBSTITUTION MATRIX
PHI-BLAST
• PROTEIN-PROTEIN BLAST
• PATTERN HIT INITIATED BLAST
• VARIATION OF PSI-BLAST
• SPECIFY A PATTERN THAT HITS MUST MATCH
• USE WHEN YOU KNOW PROTEIN FAMILY HAS A
SIGNATURE PATTERN: ACTIVE SITE, STRUCTURAL
DOMAIN, ETC.
• BETTER CHANCE OF ELIMINATING FALSE POSITIVES
• EXAMPLE: VKAHGKKV
THANKS

You might also like