2021/09/27 by Adaku Uchendu, Zeyu Ma, Uchendu, Adaku +7 · 11 citations
Computer Science · #Authorship Attribution and Profiling #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2109.13296
openalex publication_date 2021/09/27 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
Recent progress in generative language models has enabled machines to\ngenerate astonishingly realistic texts. While there are many legitimate\napplications of such models, there is also a rising need to distinguish\nmachine-generated texts from human-written ones (e.g., fake news detection).\nHowever, to our best knowledge, there is currently no benchmark environment\nwith datasets and tasks to systematically study the so-called "Turing Test"\nproblem for neural text generation methods. In this work, we present the\nTuringBench benchmark environment, which is comprised of (1) a dataset with\n200K human- or machine-generated samples across 20 labels Human, GPT-1,\nGPT-2small, GPT-2medium, GPT-2large, GPT-2xl, GPT-2PyTorch, GPT-3,\nGROVERbase, GROVERlarge, GROVERmega, CTRL, XLM, XLNETbase, XLNETlarge,\nFAIRwmt19, FAIRwmt20, TRANSFORMERXL, PPLMdistil, PPLMgpt2, (2) two\nbenchmark tasks -- i.e., Turing Test (TT) and Authorship Attribution (AA), and\n(3) a website with leaderboards. Our preliminary experimental results using\nTuringBench show that FAIRwmt20 and GPT-3 are the current winners, among all\nlanguage models tested, in generating the most human-like indistinguishable\ntexts with the lowest F1 score by five state-of-the-art TT detection models.\nThe TuringBench is available at: https://turingbench.ist.psu.edu/\n