2013/05/29 by Yinlong Xie, Xie, Yinlong, Gengxiong Wu +29 · 3 citations
Biochemistry, Genetics and Molecular Biology · #FOS: Biological sciences #Genomics (q-bio.GN) #Genomics and Phylogenetic Studies #RNA and protein synthesis mechanisms #RNA modifications and cancer
paper · pdf · doi:10.48550/arxiv.1305.6760
openalex publication_date 2013/05/29 · openalex created_date 2019/06/27 · openalex updated_date 2026/07/28
Motivation: Transcriptome sequencing has long been the favored method for quickly and inexpensively obtaining the sequences for a large number of genes from an organism with no reference genome. With the rapidly increasing throughputs and decreasing costs of next generation sequencing, RNA-Seq has gained in popularity; but given the typically short reads (e.g. 2 x 90 bp paired ends) of this technol- ogy, de novo assembly to recover complete or full-length transcript sequences remains an algorithmic challenge. Results: We present SOAPdenovo-Trans, a de novo transcriptome assembler designed specifically for RNA-Seq. Its performance was evaluated on transcriptome datasets from rice and mouse. Using the known transcripts from these well-annotated genomes (sequenced a decade ago) as our benchmark, we assessed how SOAPdenovo- Trans and two other popular software handle the practical issues of alternative splicing and variable expression levels. Our conclusion is that SOAPdenovo-Trans provides higher contiguity, lower redundancy, and faster execution. Availability and Implementation: Source code and user manual are at http://sourceforge.net/projects/soapdenovotrans/ Contact: [email protected] or [email protected]