2019/06/07 by Abhik Jana, Dmitry Puzyrev, Jana, Abhik +9
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1906.03007
openalex publication_date 2019/06/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
The compositionality degree of multiword expressions indicates to what extent\nthe meaning of a phrase can be derived from the meaning of its constituents and\ntheir grammatical relations. Prediction of (non)-compositionality is a task\nthat has been frequently addressed with distributional semantic models. We\nintroduce a novel technique to blend hierarchical information with\ndistributional information for predicting compositionality. In particular, we\nuse hypernymy information of the multiword and its constituents encoded in the\nform of the recently introduced Poincar 'e embeddings in addition to the\ndistributional information to detect compositionality for noun phrases. Using a\nweighted average of the distributional similarity and a Poincar 'e similarity\nfunction, we obtain consistent and substantial, statistically significant\nimprovement across three gold standard datasets over state-of-the-art models\nbased on distributional information only. Unlike traditional approaches that\nsolely use an unsupervised setting, we have also framed the problem as a\nsupervised task, obtaining comparable improvements. Further, we publicly\nrelease our Poincar 'e embeddings, which are trained on the output of\nhandcrafted lexical-syntactic patterns on a large corpus.\n