2026/06/24 by Papri Saha, Sudipta Kumar Das, Anonnya Sarkar +1
#q-bio.PE
Alignment-free methods in phylogenetic tree construction have major benefits in computational efficiency over alignment-based methods, but most sacrifice sequence information to pairwise distances, losing the statistical power of maximum likelihood (ML) inference. We describe ML-MAWS, an algorithm that fills this gap by encoding Minimal Absent Words (MAWs) as a binary presence/absence character matrix and estimating using an ML tree under the Lewis Mkv model using ascertainment bias correction. MAWs are obtained in linear time through the traversal of a suffix automaton. The pipeline incorporates strand-aware intersection filtering that retains only MAWs absent from both DNA orientations, entropy-based multi-length selection via Shannon entropy maximization to select the most informative lengths of MAWs, and parsimony-informative character capping to retain the most discriminative columns. We tested ML-MAWS on 14 benchmark datasets of bacterial, mitochondrial, viral, and simulated genomes with normalized Robinson-Foulds distances and matching split distances against published reference trees. The results show that while the binary encoding of MAWs can lead to higher topological error than continuous-valued distance baselines on closely related genomes, ML-MAWS is the first MAW-based method to provide per-branch bootstrap support and a rigorous probabilistic framework with ascertainment bias correction capabilities lacking from all existing alignment-free methods.