Understanding Expressed Sequence Tags (ESTs)
Understanding Expressed Sequence Tags (ESTs)
ESTs contribute to new gene discovery by providing sequence tags that can indicate the presence of genes even when the full coding sequence is not known. They allow researchers to deduce gene function tentatively based on sequence similarities to known genes . They are also pivotal in large-scale projects, such as generating 7,000 ESTs that represent 4,000 sequences from T.gondii, which identified potential for 500 novel genes .
ESTs are crucial in creating gene indices and mapping gene locations due to their ability to represent gene expression across different tissues and developmental stages. They provide data for indices such as UniGene, which partitions sequences into non-redundant gene clusters, helping researchers identify chromosomal locations of genes through sequence alignment . Mapping EST data can thus enhance understanding of gene distribution and facilitate the construction of comprehensive genomic maps .
ESTscan and hidden Markov models improve the analysis of EST data by providing a statistical framework for predicting coding regions while accounting for sequencing errors present in ESTs. ESTscan specifically utilizes hidden Markov models to rectify errors and enhance prediction accuracy, which is crucial for assembling accurate contigs and effective translation to proteins . This method addresses one of the major challenges in EST data analysis, thereby enhancing the reliability of genetic information extracted from EST sequences.
Contig assembly from EST sequences is challenged by sequence errors, redundancy, and the presence of partially spliced RNA species. To address these challenges, bioinformatics tools like TrEST are employed to assemble contigs from UniGene clusters, followed by translating them into proteins using ESTscan, which corrects sequencing errors. By leveraging statistical models and sophisticated algorithms, these tools help mitigate sequencing inaccuracies and ensure reliable contig assembly .
The advantages of using ESTs in genome projects include their ability to act as standard markers for physical mapping of the genome, and to point directly to expressed genes. They are valuable for discovering new genes and assessing gene expression levels . However, the limitations are significant as well, such as low average quality and high error rates of the sequences, and the difficulty in finding every mRNA. Additionally, many ESTs represent highly expressed genes in tissues used for library preparation, while genes expressed only in unstudied tissues may be absent .
ESTs aid in gene prediction and annotation by providing empirical evidence of gene expression that supports computational predictions. During the re-annotation of C.elegans, it was observed that EST alignments matched computational predictions in half the cases, but in 25% of cases, ESTs provided more accurate annotations than computational methods. They also revealed instances where separate predicted genes were actually exons of a single gene .
ESTs are considered an ideal source of polymorphic data because they are sequenced redundantly from libraries prepared from different individuals, capturing variations at the sequence level between different DNA samples. This redundancy enables the detection of polymorphisms in the population, despite general limitations like sequencing errors and biased gene representation . Their utility in capturing genetic diversity makes them valuable in studying population genetics and disease associations .
Methodological approaches significantly impact the quality and utility of ESTs due to the single-pass nature of sequencing, leading to low average quality and frequent errors . Many ESTs are derived from 3' untranslated regions, limiting information about coding sequences. Additionally, the representation of genes is biased toward those expressed in sampled tissues, which affects the comprehensiveness of available gene data. The raw data are often unorganized and redundant, requiring sophisticated bioinformatics tools for analysis and error correction .
ESTs play a pivotal role in assessing gene expression levels as they are generated through random sequencing from various libraries representing different tissues and development stages. This diversity allows for the comparison of expression levels across different conditions, helping to identify tissue-specific or developmental stage-specific gene expression patterns . ESTs were used, for instance, in the creation of CGAP to assess cancer genome anatomy by comparing expression profiles .
Using ESTs in cDNA microarrays is beneficial for genomic research as it allows for the simultaneous monitoring of expression levels of thousands of genes. This high-throughput technique enables the identification of gene expression profiles under various conditions or treatment scenarios, offering insights into gene function, interactions, and the underlying genetic factors of diseases . By utilizing ESTs, researchers can leverage extensive expression data across multiple samples, enhancing the power and scope of genomic studies.