vix.ing · top · new · best · stats · spec

Fast and accurate approximate inference of transcript expression from\n RNA-seq data

2014/12/18 by James Hensman, Hensman, James, Panagiotis Papastamoulis +7
Biochemistry, Genetics and Molecular Biology · #Cancer-related molecular mechanisms research #FOS: Biological sciences #Genomics (q-bio.GN) #Genomics and Phylogenetic Studies #Molecular Biology Techniques and Applications #Quantitative Methods (q-bio.QM)

paper · pdf · doi:10.48550/arxiv.1412.5995

openalex publication_date 2014/12/18 · openalex created_date 2022/09/27 · openalex updated_date 2026/07/28

Abstract

Motivation: Assigning RNA-seq reads to their transcript of origin is a\nfundamental task in transcript expression estimation. Where ambiguities in\nassignments exist due to transcripts sharing sequence, e.g. alternative\nisoforms or alleles, the problem can be solved through probabilistic inference.\nBayesian methods have been shown to provide accurate transcript abundance\nestimates compared to competing methods. However, exact Bayesian inference is\nintractable and approximate methods such as Markov chain Monte Carlo (MCMC) and\nVariational Bayes (VB) are typically used. While providing a high degree of\naccuracy and modelling flexibility, standard implementations can be\nprohibitively slow for large datasets and complex transcriptome annotations.\n Results: We propose a novel approximate inference scheme based on VB and\napply it to an existing model of transcript expression inference from RNA-seq\ndata. Recent advances in VB algorithmics are used to improve the convergence of\nthe algorithm beyond the standard Variational Bayes Expectation Maximisation\n(VBEM) algorithm. We apply our algorithm to simulated and biological datasets,\ndemonstrating a significant increase in speed with only very small loss in\naccuracy of expression level estimation. We carry out a comparative study\nagainst seven popular alternative methods and demonstrate that our new\nalgorithm provides excellent accuracy and inter-replicate consistency while\nremaining competitive in computation time.\n Availability: The methods were implemented in R and C++, and are available as\npart of the BitSeq project at urlhttps://github.com/BitSeq. The method is\nalso available through the BitSeq Bioconductor package. The source code to\nreproduce all simulation results can be accessed via\n urlhttps://github.com/BitSeq/BitSeqVBbenchmarking.\n

Citations

Related