2004/05/12 by Oren Kurland, Lillian Lee, Kurland, Oren +1
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #H.3.3 #I.2.7 #Information Retrieval (cs.IR) #cs.CL #cs.IR
paper · pdf · doi:10.48550/arxiv.cs/0405044
To appear, SIGIR 2004
arxiv created 2004/05/12 · arxiv updated 2009/12/01
Most previous work on the recently developed language-modeling approach to information retrieval focuses on document-specific characteristics, and therefore does not take into account the structure of the surrounding corpus. We propose a novel algorithmic framework in which information provided by document-based language models is enhanced by the incorporation of information drawn from clusters of similar documents. Using this framework, we develop a suite of new algorithms. Even the simplest typically outperforms the standard language-modeling approach in precision and recall, and our new interpolation algorithm posts statistically significant improvements for both metrics over all three corpora tested.