2020/08/17 by Dara Bahri, Bahri, Dara, Yi Tay +9 · 2 citations
Computer Science · #Topic Modeling #Computational Physics and Python Applications #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.2008.13533
Large generative language models such as GPT-2 are well-known for their\nability to generate text as well as their utility in supervised downstream\ntasks via fine-tuning. Our work is twofold: firstly we demonstrate via human\nevaluation that classifiers trained to discriminate between human and\nmachine-generated text emerge as unsupervised predictors of "page quality",\nable to detect low quality content without any training. This enables fast\nbootstrapping of quality indicators in a low-resource setting. Secondly,\ncurious to understand the prevalence and nature of low quality pages in the\nwild, we conduct extensive qualitative and quantitative analysis over 500\nmillion web articles, making this the largest-scale study ever conducted on the\ntopic.\n