1997/06/13 by Doug Beeferman, Beeferman, Doug, Adam Berger +3
Computer Science · Mathematics · Physics and Astronomy · #Bayesian Methods and Mixture Models #Computation and Language (cs.CL) #FOS: Computer and information sciences #Opinion Dynamics and Social Influence #Statistical Methods and Bayesian Inference
paper · pdf · doi:10.48550/arxiv.cmp-lg/9706018
openalex publication_date 1997/06/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
This paper introduces new methods based on exponential families for modeling the correlations between words in text and speech. While previous work assumed the effects of word co-occurrence statistics to be constant over a window of several hundred words, we show that their influence is nonstationary on a much smaller time scale. Empirical data drawn from English and Japanese text, as well as conversational speech, reveals that the ``attraction'' between words decays exponentially, while stylistic and syntactic contraints create a ``repulsion'' between words that discourages close co-occurrence. We show that these characteristics are well described by simple mixture models based on two-stage exponential distributions which can be trained using the EM algorithm. The resulting distance distributions can then be incorporated as penalizing features in an exponential language model.