Large Language Models Pass the Turing Test
2025/03/31 by Cameron R. Jones, Benjamin K. Bergen, Jones, Cameron R. +1 · 73 voices · 24 citations
Computer Science · Social Sciences · Medicine · #AI in Service Interactions #Language and cultural evolution #Artificial Intelligence in Healthcare and Education
paper · pdf · doi:10.48550/arxiv.2503.23674
Abstract
We evaluated 4 systems (ELIZA, GPT-4o, LLaMa-3.1-405B, and GPT-4.5) in two randomised, controlled, and pre-registered Turing tests on independent populations. Participants had 5 minute conversations simultaneously with another human participant and one of these systems before judging which conversational partner they thought was human. When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time: significantly more often than interrogators selected the real human participant. LLaMa-3.1, with the same prompt, was judged to be the human 56% of the time -- not significantly more or less often than the humans they were being compared to -- while baseline models (ELIZA and GPT-4o) achieved win rates significantly below chance (23% and 21% respectively). The results constitute the first empirical evidence that any artificial system passes a standard three-party Turing test. The results have implications for debates about what kind of intelligence is exhibited by Large Language Models (LLMs), and the social and economic impacts these systems are likely to have.
Citations
Cited by
Discussions
- UCSD: Large Language Models Pass the Turing Test [hn, 91 points, 106 comments]
- 👀This paper finds "the first robust evidence that any system passes the original three-party Turing test" People had a five minute, three-way conversation with another person & an AI. They picked GPT [bsky, 67 points, 2 comments]
- 🧪 Yes, LLMs can now pass the Turing test, but don’t confuse this with AGI, which is a long way off. arxiv.org/abs/2503.23674 [bsky, 48 points, 5 comments]
- GPT-4.5 ja LLaMa 3.1 läpäisivät Turingin testin. Siirretäänkö maalitolppia vai myönnetäänkö, että nyt tekoäly on aidosti ihmismäinen? arxiv.org/abs/2503.23674 #AGI #Tekoäly #TuringTest [bsky, 15 points, 4 comments]
- arxiv.org/abs/2503.236... www.nature.com/articles/d41... www.ie.edu/uncover-ie/h... [bsky, 6 points, 1 comments]
- arxiv.org/abs/2503.23674 [bsky, 5 points, 1 comments]
- arxiv.org/pdf/2503.23674 the full table is on page 6 and it seems to show low to no correlation with general llm familiarity etc. which ... could be under sensitive to this factor, but [bsky, 5 points, 2 comments]
- LLMs formally pass the Turing test arxiv.org/abs/2503.23674 [bsky, 4 points, 0 comments]
- How to pass a test that is just a thought experiment? "Large Language Models Pass the Turing Test" Cameron R. Jones, Benjamin K. Bergen arxiv.org/abs/2503.23674 [bsky, 4 points, 2 comments]
- Yes already there. arxiv.org/abs/2503.23674 More humans than humans fwiw :) [bsky, 3 points, 0 comments]
- Curtesy of Copilot arxiv.org/abs/2503.23674 [bsky, 3 points, 2 comments]
- arxiv.org/abs/2503.23674 Like for real... the Turing Test wasn't designed for a world where human brains collectively turn into goo. [bsky, 2 points, 1 comments]
- 🤔 Large Language Models Pass the Turing Test: arxiv.org/abs/2503.23674 [bsky, 2 points, 0 comments]
- #LLM pass the Turing Test arxiv.org/pdf/2503.23674 [bsky, 2 points, 1 comments]
- “When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time: significantly more often than interrogators selected the real human participant.” arxiv.org/pdf/2503.23 [bsky, 2 points, 0 comments]
- A quienes piensan que la IA nunca podrá traducir o hacer cosas como una persona de verdad, siempre os lo digo: cuidado con esgrimir ese argumento, porque la cosa se está complicando a pasos de gigante [bsky, 1 points, 0 comments]
- arxiv.org/pdf/2503.23674 leo que sí fueron 5 minutos con medias de 8 mensajes de aprox 40 caracteres en cada sesión... Bueno... Aceptemos pulpo 😅 [bsky, 1 points, 1 comments]
- *I get that an LLM sounds more human than a human, but how much MORE like a human can an LLM get *Not a Strong-AI superhuman, but some kinda Universal Everywoman, like everybody's Grandma https://arxi [bsky, 1 points, 0 comments]
- When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time: significantly more often than interrogators selected the real human participant. arxiv.org/pdf/2503.2367 [bsky, 1 points, 0 comments]
- Ortaçağ sorunlarıyla uğraşan ülkemizde maalesef pek yer bulmasa da, epey önemli bir gelişme. Videosu gelecek. Okumak isteyenler için makale, şurada; arxiv.org/pdf/2503.23674 [bsky, 1 points, 0 comments]
- This study was convincing to me: arxiv.org/pdf/2503.23674 And you're right, the consciousness aspect is an additional consideration. [bsky, 1 points, 0 comments]
- #LLMs Pass the #TuringTest: Interrogators mistook GPT-4.5 for a human 73% of the time—far more than they did the actual human participant🤖🔍 arxiv.org/abs/2503.23674 #AI #LLM #GenAI [bsky, 1 points, 0 comments]
- Big week in AI safety PR. Monday GPT-4.5 crushed the Turing test (arxiv.org/abs/2503.23674), Wednesday DeepMind’s 100+ page “An Approach to Technical AGI Safety and Security” (arxiv.org/abs/2504.01849 [bsky, 1 points, 0 comments]
- sorry to report, but both GPT-4.5 and LLaMa-3.1 have passed a standard three-party Turing Test (details here: arxiv.org/abs/2503.23674) [bsky, 1 points, 1 comments]
- [arXiv] LLMs pass the Turing Test arxiv.org/pdf/2503.23674 This paper presents empirical evidence from two pre-registered studies evaluating whether contemporary Large Language Models (LLMs) can pass [bsky, 1 points, 0 comments]
- arxiv.org/abs/2503.23674 [bsky, 1 points, 0 comments]
- The I in IQ also doesn't reflect intelligent intelligence. Re Turing tests, the most reasons studies I've seen implicitly show that those humans who can't tell the difference tend not to know the weak [bsky, 1 points, 2 comments]
- ChatGPT 4.5 (a částečně i Llama) prošla velmi přesvědčivě Turing testem. Měla prompt imitovat mladého introverta. Vyhrála v 73 % oproti reálnému člověku. AGI je tady. arxiv.org/pdf/2503.23674 [bsky, 1 points, 1 comments]
- Large Language Models Pass the Turing Test arxiv.org/abs/2503.23674 [bsky, 1 points, 0 comments]
- UCSD: Large Language Models Pass the Turing Test (arxiv.org) Main Link | Discussion [bsky, 1 points, 0 comments]
- Some folks did an actual double blind Turing test with GPT 4.5, arxiv.org/abs/2503.23674, and GPT 4.5 was selected as the human 73% of the time. More human than human! What a brave new world we live i [bsky, 1 points, 0 comments]
- LLMがチューリングテスト通ったとかいうやつ、PERSONAと呼ばれる指示が効いているらしい。 日本語に通じる要素を抽出して試すと確かに、これは人間とあんまり違いわかんないかも。 arxiv.org/abs/2503.23674 [bsky, 1 points, 1 comments]
- Sorry, wasn't planning on bugging you so soon after putting my foot in it the other day. But today I learned of this preprint from UCSD which gets at much of what was on my mind. With THAT title, thou [bsky, 1 points, 1 comments]
- Yo! #AI just passed the Turing Test in a UC San Diego study by Jones and Cameron arxiv.org/pdf/2503.23674 [bsky, 1 points, 0 comments]
- LLMがチューリングテストに合格する(どころか人間より人間らしいと評価される)ということは、円城塔『エピローグ』のオーバー・チューリング・クリーチャということになってくる arxiv.org/abs/2503.23674 [bsky, 1 points, 0 comments]
- Large Language Models Pass the Turing Test: "When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time: significantly more often than interrogators selected the re [bsky, 1 points, 0 comments]
- Large Language Models Pass the Turing Test Cameron R. Jones, Benjamin K. Bergen ArXiv:2503.23674 (cs) [Submitted on 31 Mar 2025] arxiv.org/abs/2503.23674 [bsky, 1 points, 0 comments]
- @phr34ky-c.artcru.sh I added some responses to your comments. I think we're at an epistemic point where we don't know what we don't know. Another brand new preprint that's awaiting peer review is: arx [bsky, 0 points, 3 comments]
- arxiv.org/pdf/2503.23674 [bsky, 0 points, 0 comments]
- Large Language Models Pass the Turing Test #llms #turing #artificialintelligence #test #benchmark [bsky, 0 points, 0 comments]
- is this the success of the AI or the failure of the human ? arxiv.org/pdf/2503.23674 [bsky, 0 points, 0 comments]
- Large Language Models Pass the Turing Test arxiv.org/abs/2503.23674 [bsky, 0 points, 0 comments]
- Researchers at UC San Diego have shown that AI systems can consistently pass the Turing test (where machines try to convince human judges they're human through text-only conversations). Chat GPT-4.5 w [bsky, 0 points, 0 comments]
- arxiv.org/pdf/2503.23674 [bsky, 0 points, 0 comments]
- LLMS passed the Turing Test and no one cares 🧵👇 First: You can read the paper here arxiv.org/abs/2503.23674 There has been a history of critiques of the test over the decades. The most common includ [bsky, 0 points, 1 comments]
- 大規模言語モデルがチューリング・テストに合格 arxiv.org/abs/2503.23674 ELIZA、GPT-4o、LLaMa-3.1-405B、GPT-4.5の4つを3者チューリングテストで評価。GPT-4.5は73%の確率で、LLaMa-3.1は56%の確率で人間だと判定された。 [bsky, 0 points, 0 comments]
- Large language models pass the Turing test. Granted the best only 73% of the time and posing as a nerdy 19 year old into video games, but still. Study from UC San Diego. #AI #artificialintelligence #m [bsky, 0 points, 0 comments]
- arxiv.org/abs/2503.23674 [bsky, 0 points, 0 comments]
- There's lots more detail in the paper arxiv.org/abs/2503.23674. We also release all of the data (including full anonymized transcripts) for further scrutiny/analysis/to prove this isn't an April Fools [bsky, 0 points, 0 comments]
- GPT-4.5首次正式通过图灵测试。https://arxiv.org/abs/2503.23674 [bsky, 0 points, 0 comments]
- It’s official: LLMs have passed The Turing Test. arxiv.org/pdf/2503.23674 [bsky, 0 points, 0 comments]
- O GPT 4.5 passou no teste de Turing com 73% de aproveitamento. arxiv.org/abs/2503.23674 [bsky, 0 points, 0 comments]
- arxiv.org/abs/2503.23674 5 minute conversations simultaneously with human participant and AI system before judging which conversational partner was human. GPT-4.5 was judged to be the human 73% of the [bsky, 0 points, 0 comments]
- Das hätte ich jetzt für gegeben angenommen arxiv.org/abs/2503.23674 [bsky, 0 points, 0 comments]
- Large Language Models Pass the Turing Test arxiv.org/abs/2503.23674 [bsky, 0 points, 0 comments]
- Quite a significant achievement, surprised this hasn’t received more press coverage. “Turing test passed”. #AI #LLMs arxiv.org/pdf/2503.23674 [bsky, 0 points, 0 comments]
- UCSD: Large Language Models Pass the Turing Test https://arxiv.org/abs/2503.23674 (https://news.ycombinator.com/item?id=43555248) [bsky, 0 points, 0 comments]
- In the study, participants conversed for 5 min with an LLM and a real person and had to figure out which was the real human. “When prompted to adopt a human-like persona, GPT-4.5 was judged to be the [bsky, 0 points, 1 comments]
- Asked to select between a human and GPT 4.5, users believed GPT-4.5 was the human 73% of the time. [bsky, 0 points, 0 comments]
- Researchers at UC San Diego just demonstrated that AI systems can consistently pass Alan Turing's famous test of machine intelligence, with OpenAI's GPT-4.5 being mistaken for human nearly three-quart [bsky, 0 points, 0 comments]
- This is 'not even wrong' in the sense that 'whether an interlocutor can noticed the difference' is an absolutely TERRIBLE criterion for 'substitutability.' arxiv.org/pdf/2503.23674 [h/t @bruces | mast [bsky, 0 points, 1 comments]
- Anyway, you can try it for yourself here: turingtest.live/ And find the full preprint here: arxiv.org/abs/2503.23674 Props to Cameron Jones and Benjamin Bergen for actually testing this! [bsky, 0 points, 1 comments]
- Meanwhile, LLaMa 3.1, when given a persona prompt, also performed well with a 56% win rate, but did not significantly outperform humans. Without a persona prompt, both GPT-4.5 and the LLaMa models per [bsky, 0 points, 0 comments]
- 3/4 📈 Implications sociales et économiques Avec des impacts potentiels sur l'emploi et les interactions sociales. Les systèmes d'IA pourraient remplacer des conversations humaines, ce qui pourrait tr [bsky, 0 points, 1 comments]
- UCSD: Large Language Models Pass the Turing Test [bsky, 0 points, 0 comments]
- UCSD: Large Language Models Pass the Turing Test #HackerNews https://arxiv.org/abs/2503.23674 [bsky, 0 points, 0 comments]
- La IA ya pasa el test de Turing lo que quiere decir que ya es posible hablar con ella y pensar que es un ser humano lo que está al otro lado. arxiv.org/pdf/2503.23674 [bsky, 0 points, 0 comments]
- So, gestern eine sehr schöne Diskussion zum Thema gehabt. Ich glaube, wir erleben da gerade einen echten "Don't look up"-Moment. Das klassische Kriterium - Turing Test - ist durch. Schlimmer: AI kann [bsky, 0 points, 1 comments]
- UCSD: Large Language Models Pass the Turing Test https://arxiv.org/abs/2503.23674 [bsky, 0 points, 0 comments]
- UCSD: Large Language Models Pass the Turing Test https://arxiv.org/abs/2503.23674 (https://news.ycombinator.com/item?id=43555248) [bsky, 0 points, 0 comments]
- actually, your are wrong: arxiv.org/abs/2503.23674 [bsky, 0 points, 1 comments]
- "The results constitute the first empirical evidence that any artificial system passes a standard three-party Turing test. The results have implications for debates about what kind of intelligence is [bsky, 0 points, 0 comments]
- The first most superficial version that pass your definition was ELIZA in the 70’s. That’s why I asked “in any setting”. I think it was a reasonable extension of the test to move away from that settin [bsky, 0 points, 0 comments]
Related