2014/02/26 by Chandrima Sarkar, Sarkar, Chandrima, Raamesh Deshpande +4
Biochemistry, Genetics and Molecular Biology · Computer Science · Environmental Science · #Computational Engineering #Environmental DNA in Biodiversity Studies #FOS: Biological sciences #FOS: Computer and information sciences #Finance #Genomics and Phylogenetic Studies #Identification and Quantification in Food #Molecular Biology Techniques and Applications #Quantitative Methods (q-bio.QM) #and Science (cs.CE) #cs.CE #q-bio.QM
paper · pdf · doi:10.48550/arxiv.1402.6775
openalex publication_date 2014/02/26 · arxiv created 2014/02/27 · arxiv updated 2014/02/28 · openalex created_date 2022/10/01 · openalex updated_date 2026/07/28
In this paper we aim at investigating whether barcode sequence features can predict the read count ambiguities caused during PCR based next generation sequencing techniques. The methodologies we used are mutual information based motif discovery and Lasso regression technique using features generated from the barcode sequence. The results indicate that there is a certain degree of correlation between motifs discovered in the sequences and the read counts. Our main contribution in this paper is a thorough investigation of the barcode features that gave us useful information regarding the significance of the sequence features and the sequence containing the discovered motifs in prediction of read counts.