0% found this document useful (0 votes)
5 views3 pages

TranscriptClean: Long-Read Correction Tool

Uploaded by

wang142857
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views3 pages

TranscriptClean: Long-Read Correction Tool

Uploaded by

wang142857
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Bioinformatics, 35(2), 2019, 340–342

doi: 10.1093/bioinformatics/bty483
Advance Access Publication Date: 15 June 2018
Applications Note

Gene expression

TranscriptClean: variant-aware correction of

Downloaded from [Link] by Johns Hopkins University user on 07 July 2024


indels, mismatches and splice junctions in
long-read transcripts
Dana Wyman1,2 and Ali Mortazavi1,2,*
1
Department of Developmental and Cell Biology and 2Center for Complex Biological Systems, UC Irvine, Irvine, CA,
92697, USA
*To whom correspondence should be addressed.
Associate Editor: Bonnie Berger
Received on March 13, 2018; revised on May 18, 2018; editorial decision on June 10, 2018; accepted on June 13, 2018

Abstract
Motivation: Long-read, single-molecule sequencing platforms hold great potential for isoform
discovery and characterization of multi-exon transcripts. However, their high error rates are an obs-
tacle to distinguishing novel transcript isoforms from sequencing artifacts. Therefore, we devel-
oped the package TranscriptClean to correct mismatches, microindels and noncanonical splice
junctions in mapped transcripts using the reference genome while preserving known variants.
Results: Our method corrects nearly all mismatches and indels present in a publically available
human PacBio Iso-seq dataset, and rescues 39% of noncanonical splice junctions.
Availability and implementation: All Python and R scripts used in this paper are available at
[Link]
Contact: [Link]@[Link]

1 Introduction To address this problem, various PacBio-specific tools have been


Conventional short-read RNA sequencing is widely used to quantify developed to correct transcripts downstream of the ToFU pipeline.
gene expression in a variety of applications. While cost-effective and TAPIS, HapIso and SQANTI use a reference-guided approach to
accurate, short reads lack the ability to resolve full-length mamma- correct indels within exons (Abdel-Ghany, 2016; Mangul et al.,
lian isoforms, which are commonly multiple kilobases long (Conesa, 2017; Tardaguila, 2018). HapIso distinguishes single nucleotide var-
2016). Long-read sequencing platforms such as Pacific Biosciences iants from errors in a haplotype-aware manner by phasing long
(PacBio) and Oxford Nanopore bypass the transcript reconstruction reads. TAPIS and SQANTI deal with remaining errors by removing
challenges of short reads, but have substantially higher error rates. affected transcripts, the former using a splice junction quality filter,
Raw PacBio reads have a stochastic error rate of 11–15%, including and the latter using a random forest classifier. While these methods
single-base mismatches and microindel errors (Eid, 2009). Microindels produce cleaner PacBio datasets, none of them attempt to correct
are especially problematic during isoform mapping because they can noncanonical splice junctions arising from microindel errors.
misrepresent splice junction locations. Furthermore, HapIso requires multiple transcripts per gene in order
Circular consensus correction and read polishing steps in the for the phasing to work, which is not a given depending on sequenc-
PacBio ToFU analysis pipeline can substantially reduce the error ing depth and gene expression level.
rate for most transcripts once raw reads are processed (Eid, 2009; We present TranscriptClean, a program that uses the reference
Gordon, 2015). However, this correction process is only effective genome, splice annotation and a variant file to correct mismatches,
when multiple sequencing passes over the same insert molecule are microindels and noncanonical splice junctions in PacBio transcripts
available, which becomes less likely as transcript length increases while preserving known variants. Running TranscriptClean on a
(Rhoads and Au, 2015). publicly available PacBio human transcriptome from GM12878

C The Author(s) 2018. Published by Oxford University Press.


V 340
This is an Open Access article distributed under the terms of the Creative Commons Attribution License ([Link] which permits
unrestricted reuse, distribution, and reproduction in any medium, provided the original work is properly cited.
TranscriptClean 341

Table 1. Summary of GM12878 TranscriptClean results

No TC TC with GENCODE Corrected TC with GM12878 Corrected


splice junctions Illumina SJs and variants

Total Transcripts 568048 568048 – 568048 –


Canon. Transc. 479005 512092 – 511541 –
Noncan. Transc. 89043 55956 37% 56507 37%
Deletions 3133172 23047 99% 29883 99%
Insertions 1901787 20175 99% 21816 99%
Mismatches 14380068 0 100% 295547 98%

Downloaded from [Link] by Johns Hopkins University user on 07 July 2024


NCSJ 109268 66304 39% 66784 39%

(Tilgner, 2014), we corrected 99% of indels, 98% of mismatches high-confidence splice junctions (derived from same-sample mapped
and 39% of noncanonical splice junctions present in these tran- short RNA-seq reads or a reference annotation) and is changed to
scripts. This allowed us to salvage 32 536 transcripts that would match the known junction when the distance between the NCSJ and
have been discarded under previous workflows because of nonca- its nearest high-confidence junction is microindel-sized.
nonical splice junctions.

3 Results
2 Materials and methods
We performed two TranscriptClean runs on CCS-processed circular
2.1 Indel and mismatch correction consensus GM12878 PacBio transcripts from Tilgner (2014) (Table
TranscriptClean processes transcripts in the SAM format, scanning 1). In the first, we used known human splice junction annotations
each entry to look for insertions, deletions and mismatches relative from GENCODE v24 and no variant file. Next we provided
to the reference genome. Indels less than or equal to the size thresh- GM12878-specific variants and splice junctions derived from
old (default  5 bp) are modified to match the reference sequence. GM12878 short reads. When provided GM12878-specific referen-
Mismatches in the transcripts are replaced with the reference base. ces for correction, TranscriptClean corrected 99% of indels and
Indel and mismatch correction can also be run in variant-aware 39% of NCSJs, rescuing 32 536 transcripts no longer considered
mode to avoid removing variants of interest to the user. In this noncanonical. 98% of mismatches were corrected, with the remain-
mode, mismatches and indels are changed to the reference sequence ing 2% representing known NA12878 SNPs.
only if they do not match the position and sequence of a known vari- A major goal of long-read isoform characterization is to provide
ant in a user-provided VCF file. A potential downside of running a higher-quality reference transcriptome for short-read quantitation.
mismatch correction is that it will remove novel SNPs or RNA edit- If such a reference contains frequent sequencing errors, reads will
ing events not provided in the VCF. not map well to it, defeating its purpose. Furthermore, downstream
TranscriptClean outputs a SAM file of corrected transcripts with analysis programs commonly ignore transcripts with one or more
updated CIGAR, sequence and MD/NM fields. It also provides a fasta NCSJs, effectively throwing out long-read data that could provide
file of corrected sequences alongside log files tracking changes to indi- interesting isoform information. Repairing errors where possible
vidual errors and transcripts. The accessory script generate_report.R allows more data to be used, particularly for longer transcripts.
produces figures summarizing the TranscriptClean results, and can While the current version of variant-aware TranscriptClean does
also be used to choose an appropriate indel size threshold for a given not account for the case of a sequencing error converting a real SNP
dataset, as the size distribution may vary across different PacBio to the reference base, nor the case where a real indel is disguised by
chemistries. one or more sequencing errors, we hope to improve correction for
special cases like these in future versions and to support transcript
correction of Oxford Nanopore reads.
2.2 Noncanonical splice junction correction
TranscriptClean also provides the option of correcting noncanonical
splice junctions. During pre-mRNA splicing, dinucleotides at the
start and end of the intron form highly conserved canonical motifs Funding
GTAG, GCAG and ATAC, with GTAG accounting for 98.9% of This work was supported by the National Human Genome Research Institute
known human splice junctions (Dobin, 2013; Parada, 2014). to AM [UM1 HG009443].
Noncanonical splice junctions (NCSJs) are very rare events, which
Conflict of Interest: none declared.
suggests that most NCSJs in long-read transcripts are likely to be
sequencing errors. Typically 10–20% of PacBio transcripts contain
at least one NCSJ (Tardaguila, 2018). References
When a microindel error disrupts a splice boundary, the read
Abdel-Ghany,S.E. et al. (2016) A survey of the sorghum transcriptome using
mapping can be affected in a variety of ways. In one scenario, the
single-molecule long reads. Nat. Commun., 7, 11706.
entire junction is shifted upstream or downstream of its original lo-
Conesa,A. et al. (2016) A survey of best practices for RNA-seq data analysis.
cation. In another, the error is split across the junction, resulting in a
Genome Biol., 17, 13.
smaller indel on each side. Finally, the error may only affect one side Dobin,A. et al. (2013) STAR: ultrafast universal RNA-seq aligner.
of the junction. Bioinformatics (Oxford, England), 29, 15–21.
To identify NCSJs, TranscriptClean checks the intron motif of Eid,J. et al. (2009) Real-time DNA sequencing from single polymerase mole-
each transcript splice site. Each NCSJ is compared to user-provided cules. Science, 323, 133–138.
342 [Link] and [Link]

Gordon,S.P. et al. (2015) Widespread polycistronic transcripts in fungi Rhoads,A. and Au,K., F. (2015) PacBio sequencing and its applications.
revealed by single-molecule mRNA sequencing. PLoS One, 10, e0132628. Genomics Proteomics Bioinf., 13, 278–289.
Mangul,S. et al. (2017) HapIso: an accurate method for the haplotype-specific Tardaguila,M., d., l. et al. (2018) SQANTI: extensive characterization of
isoforms reconstruction from long single-molecule reads. IEEE Trans. long-read transcript sequences for quality control in full-length transcrip-
NanoBioscience, 16, 108–115. tome identification and quantification. Genome Res., 28, 396–411.
Parada,G.E. et al. (2014) A comprehensive survey of non-canonical splice sites Tilgner,H. et al. (2014) Defining a personal, allele-specific, and single-molecule
in the human transcriptome. Nucleic Acids Res., 42, 10564–10578. long-read transcriptome. Proc. Natl. Acad. Sci. USA., 111, 9869–9874.

Downloaded from [Link] by Johns Hopkins University user on 07 July 2024

You might also like