vix.ing · top · new · best · stats · spec

Improving Statistical Language Model Performance with Automatically\n Generated Word Hierarchies

1995/03/09 by John McMahon, McMahon, John, F.J. Smith +1
Computer Science · #Natural Language Processing Techniques #Text and Document Classification Technologies #Advanced Text Analysis Techniques

paper · pdf · doi:10.48550/arxiv.cmp-lg/9503011

Abstract

An automatic word classification system has been designed which processes\nword unigram and bigram frequency statistics extracted from a corpus of natural\nlanguage utterances. The system implements a binary top-down form of word\nclustering which employs an average class mutual information metric. Resulting\nclassifications are hierarchical, allowing variable class granularity. Words\nare represented as structural tags --- unique n-bit numbers the most\nsignificant bit-patterns of which incorporate class information. Access to a\nstructural tag immediately provides access to all classification levels for the\ncorresponding word. The classification system has successfully revealed some of\nthe structure of English, from the phonemic to the semantic level. The system\nhas been compared --- directly and indirectly --- with other recent word\nclassification systems. Class based interpolated language models have been\nconstructed to exploit the extra information supplied by the classifications\nand some experiments have shown that the new models improve model performance.\n

Citations

Related