2016/02/03 by Joey McCollum, Stephen Brown, McCollum, Joey +1 · 1 citation
Computer Science · #6U815 #FOS: Computer and information sciences #I.2.7 #Machine Learning (cs.LG) #acm:6U815 #cs.LG #msc:6U815
paper · pdf · doi:10.48550/arxiv.1602.01323
31 pages, 2 figures, 42 tables
arxiv created 2016/02/03 · arxiv updated 2016/02/04
The text-critical practice of grouping witnesses into families or texttypes often faces two obstacles: Contamination in the manuscript tradition, and co-dependence in identifying characteristic readings and manuscripts. We introduce non-negative matrix factorization (NMF) as a simple, unsupervised, and efficient way to cluster large numbers of manuscripts and readings simultaneously while summarizing contamination using an easy-to-interpret mixture model. We apply this method to an extensive collation of the New Testament epistle of Jude and show that the resulting clusters correspond to human-identified textual families from existing research.