2013/07/30 by Michał Modzelewski, Modzelewski, Michał, Norbert Dojer +1
Biochemistry, Genetics and Molecular Biology · #Bioinformatics and Genomic Networks #FOS: Biological sciences #Genetic diversity and population structure #Genomics and Phylogenetic Studies #Quantitative Methods (q-bio.QM) #q-bio.QM
paper · pdf · doi:10.48550/arxiv.1307.7844
Peer-reviewed and presented as part of the 13th Workshop on Algorithms in Bioinformatics (WABI2013)
arxiv created 2013/07/30 · openalex publication_date 2013/07/30 · arxiv updated 2013/07/31 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Progressive methods offer efficient and reasonably good solutions to the multiple sequence alignment problem. However, resulting alignments are biased by guide-trees, especially for relatively distant sequences. We propose MSARC, a new graph-clustering based algorithm that aligns sequence sets without guide-trees. Experiments on the BAliBASE dataset show that MSARC achieves alignment quality similar to best progressive methods and substantially higher than the quality of other non-progressive algorithms. Furthermore, MSARC outperforms all other methods on sequence sets with the similarity structure hardly represented by a phylogenetic tree. Furthermore, MSARC outperforms all other methods on sequence sets whose evolutionary distances are hardly representable by a phylogenetic tree. These datasets are most exposed to the guide-tree bias of alignments. MSARC is available at http://bioputer.mimuw.edu.pl/msarc