2021/09/27 by Laura Aina, Aina, Laura, Xixian Liao +5
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech and dialogue systems
paper · pdf · doi:10.48550/arxiv.2109.13105
It is often posited that more predictable parts of a speaker's meaning tend\nto be made less explicit, for instance using shorter, less informative words.\nStudying these dynamics in the domain of referring expressions has proven\ndifficult, with existing studies, both psycholinguistic and corpus-based,\nproviding contradictory results. We test the hypothesis that speakers produce\nless informative referring expressions (e.g., pronouns vs. full noun phrases)\nwhen the context is more informative about the referent, using novel\ncomputational estimates of referent predictability. We obtain these estimates\ntraining an existing coreference resolution system for English on a new task,\nmasked coreference resolution, giving us a probability distribution over\nreferents that is conditioned on the context but not the referring expression.\nThe resulting system retains standard coreference resolution performance while\nyielding a better estimate of human-derived referent predictability than\nprevious attempts. A statistical analysis of the relationship between model\noutput and mention form supports the hypothesis that predictability affects the\nform of a mention, both its morphosyntactic type and its length.\n