2001/05/07 by Mark Johnson, Johnson, Mark
Computer Science · #AI-based Problem Solving and Planning #Computation and Language (cs.CL) #FOS: Computer and information sciences #H.5.2 #Natural Language Processing Techniques #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.cs/0105012
8 pages, Proceedings of the ACL 2001
arxiv created 2001/05/07 · openalex publication_date 2001/05/07 · arxiv updated 2009/11/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/31
This paper compares two different ways of estimating statistical language models. Many statistical NLP tagging and parsing models are estimated by maximizing the (joint) likelihood of the fully-observed training data. However, since these applications only require the conditional probability distributions, these distributions can in principle be learnt by maximizing the conditional likelihood of the training data. Perhaps somewhat surprisingly, models estimated by maximizing the joint were superior to models estimated by maximizing the conditional, even though some of the latter models intuitively had access to ``more information''.