0% found this document useful (0 votes)
163 views19 pages

Understanding Expressed Sequence Tags (ESTs)

Expressed sequence tags (ESTs) are short sequences of cDNA derived from mRNA. They provide a snapshot of gene expression in a given tissue or developmental stage. ESTs are constructed by isolating mRNA from a cell/tissue, reverse transcribing it to cDNA, cloning the cDNA into a vector to make a library, and then sequencing the 5' and 3' ends. While ESTs only represent a small portion of the genome, they contain the majority of functional information. They are useful for gene discovery, mapping genes, and assessing gene expression levels. Common databases like UniGene cluster ESTs into gene families to organize the large amounts of raw EST sequence data.

Uploaded by

Anand Dangre
Copyright
© Attribution Non-Commercial (BY-NC)
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPT, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
163 views19 pages

Understanding Expressed Sequence Tags (ESTs)

Expressed sequence tags (ESTs) are short sequences of cDNA derived from mRNA. They provide a snapshot of gene expression in a given tissue or developmental stage. ESTs are constructed by isolating mRNA from a cell/tissue, reverse transcribing it to cDNA, cloning the cDNA into a vector to make a library, and then sequencing the 5' and 3' ends. While ESTs only represent a small portion of the genome, they contain the majority of functional information. They are useful for gene discovery, mapping genes, and assessing gene expression levels. Common databases like UniGene cluster ESTs into gene families to organize the large amounts of raw EST sequence data.

Uploaded by

Anand Dangre
Copyright
© Attribution Non-Commercial (BY-NC)
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPT, PDF, TXT or read online on Scribd

EST Expressed Sequence Tags

-Manali Mehendale

The ultimate goal of the genome project is to produce a complete and accurate sequence of the entire genetic material

What are ESTs ??

Expressed Sequence Tags (ESTs) are short (usually about 300-500 bp), single-pass sequence reads from mRNA (cDNA). Typically they are produced in large batches. They represent a snapshot of genes expressed in a given tissue and/or at a given developmental stage. They are tags (some coding, others not) of expression for a given cDNA library Expressed sequence tags (ESTs) - An expressed sequence tag (EST) is a small part of the active part of a gene which can be used to fish the rest of the gene out of the chromosome

Overview of how ESTs are constructed


Cell/Tissue
Isolate mRNA, Reverse transcription Deposit the EST sequences

Clone cDNA into a vector to make a library

Sequence 5 and 3 ends of cDNA insert

Pick individual clones

Debates over ESTs


FOR:

Merits of sequencing cDNA Represent only 3% of DNA, but represent the vast majority of information content Not easy to predict gene coding regions and their mRNA

AGAINST:

Difficult to find every mRNA

Relationship between EST sequence and mRNA transcript


Exon Intron

Splicing

Splice variants

3 EST sequences

Consequences of Methodology

Many sequences derived from 3 ends of mRNA thus mostly contain information about 3 untranslated regions. Average quality is low, errors are quite common. Genes that are highly expressed in tissues from which libraries have been made will be represented in many EST sequences. Genes that are expressed only in tissues that were not used for preparing cDNA will not be represented in the EST database. Substantial number of sequences derived from partially spliced RNA species. ESTs as good as the clones from which they are derived.

Why ESTs ?!

EST's can act as standard markers for the physical mapping of the genome. Additional advantage of pointing directly to an expressed gene.
Used intensively as a source of information for the discovery of new genes whose function can be tentatively deduced from sequence.

The DATA

Raw data is unorganised, unannotated, redundant and of low quality Present in flat file format Found at:
SAMPLE ENTRY (Genbank) CURRENT STATISTICS

EST clustering

UNIGENE:
-Experimental system for automatically partitioning Genbank sequences into a non redundant set of gene oriented clusters. -In addition to sequences of well characterized genes, hundreds of thousands EST sequences have been included

Unigene
Genbank mRNAs Genbank genomic CDSs dbESTs

Preliminary clusters
Unanchored clusters

UniGene

UniGene.

Querying Unigene result

OTHER SIMILAR DATABASES: TIGR Gene Indices BLAST -result

STACK BLAST-result

Querying the EST database

BLAST : result

FASTA (Only email submission allowed at [Link] Smith and Waterman

CAN ALSO PERFORM MOTIF SEARCHES

EST CONTIG ASSEMBLY

GCG package

FINDING CODING REGIONS

ESTScan : [Link]
Type of hidden Markov model that explicitly deals with the possibility of errors in the sequence to analyze, and incorporates a method for correcting these errors.

1) DNA (CDS) : result 2) Full DNA : result 3) Protein : result

TrEST
TrEST is an attempt to produce contigs from UniGene clusters (see below) and to translate them into proteins. This is a two-step process: (i) assembly of contigs from a collection of ESTs; (ii) translation of the assembled contigs into protein (using ESTscan)

Use's of ESTs

Hunting for novel genes

A large scale project generated 7,000 ESTs representing 4,000 sequences from [Link], sequence comparison with public databases identified potential for 500 novel genes.

Creating Gene Indices

Using EST and STS. Map location is provided by Unigene.

Uses of ESTs.

Gene Predication in genomic DNA

Reannotation of [Link]: - In about half the cases the computationally predicted genes were identical to the EST alignments; 25% of the genes were predicted with less accuracy, remaining were predicted poorly. - Error rate of worm genome is less than 0.0001 - Many of the alternative splices are not annotated on the genomic sequence - Computational methods may predict separate genes, whereas EST shows that these are exons of a single gene.

Uses of ESTs

Ideal source of polymorphic data

Since ESTs are sequenced redundantly from libraries prepared from different individuals, they seem an ideal source of polymorphic data.

Assessing the level of gene expression

Since ESTs are generated by random sequencing of clones from many different
[Link] is creation of CGAP.

In cDNA microarrays SAGE

Common questions

Powered by AI

ESTs contribute to new gene discovery by providing sequence tags that can indicate the presence of genes even when the full coding sequence is not known. They allow researchers to deduce gene function tentatively based on sequence similarities to known genes . They are also pivotal in large-scale projects, such as generating 7,000 ESTs that represent 4,000 sequences from T.gondii, which identified potential for 500 novel genes .

ESTs are crucial in creating gene indices and mapping gene locations due to their ability to represent gene expression across different tissues and developmental stages. They provide data for indices such as UniGene, which partitions sequences into non-redundant gene clusters, helping researchers identify chromosomal locations of genes through sequence alignment . Mapping EST data can thus enhance understanding of gene distribution and facilitate the construction of comprehensive genomic maps .

ESTscan and hidden Markov models improve the analysis of EST data by providing a statistical framework for predicting coding regions while accounting for sequencing errors present in ESTs. ESTscan specifically utilizes hidden Markov models to rectify errors and enhance prediction accuracy, which is crucial for assembling accurate contigs and effective translation to proteins . This method addresses one of the major challenges in EST data analysis, thereby enhancing the reliability of genetic information extracted from EST sequences.

Contig assembly from EST sequences is challenged by sequence errors, redundancy, and the presence of partially spliced RNA species. To address these challenges, bioinformatics tools like TrEST are employed to assemble contigs from UniGene clusters, followed by translating them into proteins using ESTscan, which corrects sequencing errors. By leveraging statistical models and sophisticated algorithms, these tools help mitigate sequencing inaccuracies and ensure reliable contig assembly .

The advantages of using ESTs in genome projects include their ability to act as standard markers for physical mapping of the genome, and to point directly to expressed genes. They are valuable for discovering new genes and assessing gene expression levels . However, the limitations are significant as well, such as low average quality and high error rates of the sequences, and the difficulty in finding every mRNA. Additionally, many ESTs represent highly expressed genes in tissues used for library preparation, while genes expressed only in unstudied tissues may be absent .

ESTs aid in gene prediction and annotation by providing empirical evidence of gene expression that supports computational predictions. During the re-annotation of C.elegans, it was observed that EST alignments matched computational predictions in half the cases, but in 25% of cases, ESTs provided more accurate annotations than computational methods. They also revealed instances where separate predicted genes were actually exons of a single gene .

ESTs are considered an ideal source of polymorphic data because they are sequenced redundantly from libraries prepared from different individuals, capturing variations at the sequence level between different DNA samples. This redundancy enables the detection of polymorphisms in the population, despite general limitations like sequencing errors and biased gene representation . Their utility in capturing genetic diversity makes them valuable in studying population genetics and disease associations .

Methodological approaches significantly impact the quality and utility of ESTs due to the single-pass nature of sequencing, leading to low average quality and frequent errors . Many ESTs are derived from 3' untranslated regions, limiting information about coding sequences. Additionally, the representation of genes is biased toward those expressed in sampled tissues, which affects the comprehensiveness of available gene data. The raw data are often unorganized and redundant, requiring sophisticated bioinformatics tools for analysis and error correction .

ESTs play a pivotal role in assessing gene expression levels as they are generated through random sequencing from various libraries representing different tissues and development stages. This diversity allows for the comparison of expression levels across different conditions, helping to identify tissue-specific or developmental stage-specific gene expression patterns . ESTs were used, for instance, in the creation of CGAP to assess cancer genome anatomy by comparing expression profiles .

Using ESTs in cDNA microarrays is beneficial for genomic research as it allows for the simultaneous monitoring of expression levels of thousands of genes. This high-throughput technique enables the identification of gene expression profiles under various conditions or treatment scenarios, offering insights into gene function, interactions, and the underlying genetic factors of diseases . By utilizing ESTs, researchers can leverage extensive expression data across multiple samples, enhancing the power and scope of genomic studies.

You might also like