GPT-4 is judged more human than humans in displaced and inverted Turing tests
2024/07/11 by Ishika Rathi, Rathi, Ishika, Sydney Taylor +5 · 7 voices · 2 citations
Engineering · #Ferroelectric and Negative Capacitance Devices
paper · pdf · doi:10.48550/arxiv.2407.08853
Abstract
Everyday AI detection requires differentiating between people and AI in informal, online conversations. In many cases, people will not interact directly with AI systems but instead read conversations between AI systems and other people. We measured how well people and large language models can discriminate using two modified versions of the Turing test: inverted and displaced. GPT-3.5, GPT-4, and displaced human adjudicators judged whether an agent was human or AI on the basis of a Turing test transcript. We found that both AI and displaced human judges were less accurate than interactive interrogators, with below chance accuracy overall. Moreover, all three judged the best-performing GPT-4 witness to be human more often than human witnesses. This suggests that both humans and current LLMs struggle to distinguish between the two when they are not actively interrogating the person, underscoring an urgent need for more accurate tools to detect AI in conversations.
Cited by
Discussions
- Yet more evidence that people can’t accurately detect well-prompted AI writing (and AI can’t accurately detect well-prompted AI writing, either). arxiv.org/pdf/2407.08853 [bsky, 116 points, 5 comments]
- GPT-4 is judged more human than humans in displaced and inverted Turing tests [hn, 3 points, 1 comments]
- Interestingly, there seem to be no published tests of this. There is a literature on AI adjudicating Turing tests, but apparently always as a passive classifier, never as an active interrogator (so, n [bsky, 2 points, 1 comments]
- Ehrm.. But the test doesn't have any architectural requirement, and they are already passing them? It's a bit cheeky, but they can pretty much score better than humans, just by getting told to be some [bsky, 1 points, 1 comments]
- Unlike GPT4 GPT-4 is judged more human than humans in displaced and inverted Turing tests: arxiv.org/pdf/2407.08853 A GPT-4 persona is judged to be human BY A HUMAN in 50.6% of cases of live dialogu [bsky, 0 points, 0 comments]
- But when it was studied scientifically, humans were still better at appearing human than AIs, at least until July 2024, when a GPT-4 based model significantly outperformed humans at a 2-party Turing T [bsky, 0 points, 1 comments]
- "This suggests that neither AI nor humans are reliable with detecting AI-contributions to online conversations." Paper (PDF): https://arxiv.org/pdf/2407.08853 [bsky, 0 points, 0 comments]
Related