2002/07/23 by Patrick Pantel, Dekang Lin · 3 citations
Computer Science · Mathematics · #Artificial intelligence #Centroid #Cluster (spacecraft) #Cluster analysis #Computer science #Domain (mathematical analysis) #Element (criminal law) #Feature (linguistics) #Feature vector #Image (mathematics) #Information retrieval #Linguistics #Mathematics #Natural Language Processing Techniques #Natural language processing #Precision and recall #Recall #Semantic Web and Ontologies #Set (abstract data type) #Similarity (geometry) #Space (punctuation) #Topic Modeling #Word (group theory)
paper · doi:10.1145/775047.775138
openalex publication_date 2002/07/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
Inventories of manually compiled dictionaries usually serve as a source for word senses. However, they often include many rare senses while missing corpus/domain-specific senses. We present a clustering algorithm called CBC (Clustering By Committee) that automatically discovers word senses from text. It initially discovers a set of tight clusters called committees that are well scattered in the similarity space. The centroid of the members of a committee is used as the feature vector of the cluster. We proceed by assigning words to their most similar clusters. After assigning an element to a cluster, we remove their overlapping features from the element. This allows CBC to discover the less frequent senses of a word and to avoid discovering duplicate senses. Each cluster that a word belongs to represents one of its senses. We also present an evaluation methodology for automatically measuring the precision and recall of discovered senses.