2020/01/01 by Wendi Ren, Yinghao Li, Hanting Su +4 · 46 citations
Computer Science · #Artificial intelligence #Artificial neural network #Classifier (UML) #Code (set theory) #Computer science #Data mining #Labeled data #Machine Learning and Data Classification #Machine learning #Noise (video) #Noise reduction #Noisy data #Pattern recognition (psychology) #Reliability (semiconductor) #Source code #Text and Document Classification Technologies #Topic Modeling #cs.CL
paper · pdf · doi:10.18653/v1/2020.findings-emnlp.334
16 pages, 7 figures
openalex publication_date 2020/01/01 · arxiv created 2020/10/09 · arxiv updated 2021/03/12 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
We study the problem of learning neural text classifiers without using any labeled data, but only easy-to-provide rules as multiple weak supervision sources. This problem is challenging because rule-induced weak labels are often noisy and incomplete. To address these two challenges, we design a label denoiser, which estimates the source reliability using a conditional soft attention mechanism and then reduces label noise by aggregating rule-annotated weak labels. The denoised pseudo labels then supervise a neural classifier to predicts soft labels for unmatched samples, which address the rule coverage issue. We evaluate our model on five benchmarks for sentiment, topic, and relation classifications.