Data and trained models for "Empirical Evidence of Large Language Model's Influence on Human Spoken Communication"
2024/09/03 by Hiromu Yakura, Yakura, Hiromu, Ezequiel Lopez-Lopez +18 · 44 voices · 13 citations
Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #cs.AI #cs.CL #cs.CY #cs.HC
paper · pdf · doi:10.48550/arxiv.2409.01754
openalex publication_date 2024/09/03 · openalex created_date 2024/09/29 · openalex updated_date 2026/07/28
Abstract
Data accompanying Yakura, Lopez-Lopez, Brinkmann et al., "Empirical Evidence of Large Language Model's Influence on Human Spoken Communication" (arXiv:2409.01754). The reproduction and analysis code is archived separately (Zenodo 10.5281/zenodo.21296093). This deposit contains three files: manuscript-data-figures.tar.gz (~0.4 GB) — precomputed figure-source outputs: synthetic-control results, in-space placebo time series, and Bayesian change-point (Stan) results for each manuscript figure and its robustness variants (Fig. 1, Fig. 2, Fig. 3, the YouTube appendix, and the Supplementary figures). Extracts into observationalpipeline/runs/; sufficient to render the paper figures without re-running the pipeline. manuscript-data-substrate.tar.gz (~2.0 GB) — the analysis substrate: per-episode word-count matrices (counts50k.parquet, raw; counts50kauditedmerged.parquet, sense-audited) and the trained spontaneity-classifier weights (model81284.pt). Input for re-running the sense audit, synthetic-control, and change-point analyses from counts. manuscript-data-ngrams.parquet (~5.0 GB) — monthly podcast n-gram frequency tables, one row per (analysis group, month, n-gram): surface-form counts for n = 1–5 with a ≥5-occurrence cutoff (~445 million rows). Derived aggregate lexical statistics. See REPRODUCE.md in the code repository for how each file is used. Verbatim transcripts are not included and are available through an institutional data-use agreement.
Citations
Cited by
Discussions
- Everyone is starting to sound like AI, even in spoken language Analysis of 280,000 transcripts of videos of talks & presentations from academic channels finds they increasingly used words that are fav [bsky, 201 points, 17 comments]
- Chatbots are changing how we talk! An analysis of 360,445 academic talks and 771,591 podcast episodes across multiple disciplines revealed an abrupt increase in the use of words preferentially generat [bsky, 17 points, 0 comments]
- *Delvish as an AI-generated dialect used by humans. *I get it about the word-use stats, but I also wonder about humans "thinking" like generated chatbot content appears to think. How would one measure [bsky, 11 points, 2 comments]
- The secret third option of humans changing their writing patterns to sound like an LLM 👀 arxiv.org/abs/2409.017... [bsky, 10 points, 2 comments]
- Studie über veränderte Worthäufigkeiten in akademischen Youtube-Videos und Podcasts seit ChatGPT arxiv.org/pdf/2409.01754 [bsky, 9 points, 1 comments]
- You might want to delve into this paper. I want to underscore, that's a joke you'll comprehend only with meticulous reading of it. arxiv.org/pdf/2409.01754 [bsky, 6 points, 2 comments]
- Empirical evidence of LLM's influence on human spoken communication [hn, 6 points, 1 comments]
- arxiv.org/abs/2409.01754 Scary! Authors show the emergence of a bidirectional cultural feedback loop where AI systems trained on human data are now reshaping human language and culture. [bsky, 5 points, 0 comments]
- Researchers at the Max Planck Institute for Human Development, who analyzed close to 280,000 YouTube videos from academic channels, found that AI is influencing how humans write and speak, homogenizin [bsky, 5 points, 0 comments]
- ChatGPT 如何影響人類的語言習慣(已經不是很新的論文了) arxiv.org/abs/2409.01754 [bsky, 4 points, 1 comments]
- Empirical evidence of Large Language Model’s influence on human spoken communication (🤔 But aren't LLMs built upon human communication?) arxiv.org/pdf/2409.017... [bsky, 4 points, 0 comments]
- Empirical evidence of LLM's influence on human spoken communication [hn, 3 points, 0 comments]
- Just when you thought the world couldn’t get more topsy-turvy, new research finds that humans are starting to sound like ChatGPT. 😳🤦🤷 arxiv.org/pdf/2409.017... [bsky, 3 points, 0 comments]
- I wonder what is going on with the publication process for this paper. It was last revised in July 2025, but it still hasn't been published, as far as I can tell. Here is version 3 (instead of version [bsky, 2 points, 0 comments]
- New study from the Max‑Planck Institute: ChatGPT has begun influencing human speech. Analysis of over 740,000 hours of academic podcasts and YouTube lectures shows a significant surge in ChatGPT‑prefe [bsky, 2 points, 0 comments]
- 🚨 preprint 🚨 Can ChatGPT influence the way we speak to each other? I.e. can they shape human culture? We transcribed and analyzed 300k YouTube videos to see if humans increased the use of words l [bsky, 2 points, 0 comments]
- Interesting result, but the ivory tower of academia seems alive and well: "Our result [raises] concerns about the potential of AI ... to be deliberately misused for mass manipulation." That's the poin [bsky, 2 points, 0 comments]
- 🗣️ Au quotidien, nos façons de parler et d’écrire évoluent aussi. Des chercheurs ont observé une hausse des « tics de langage » de ChatGPT chez les youtubeurs et créateurs de balados. 👉🏼 « Empirica [bsky, 2 points, 1 comments]
- NB: "25" verweist darauf: H. Yakura et al., Empirical evidence of Large Language Model’s influence on human spoken communication (2024), arxiv.org/abs/2409.01754 [bsky, 1 points, 0 comments]
- achei arxiv.org/abs/2409.01754 vi pela primeira vez nesse tiktok vt.tiktok.com/ZSmr5w6hV/ [bsky, 1 points, 0 comments]
- Noting with amusement: In Star Trek, being part of The Borg means that "you" think like "it". (In seriousness, though, this doesn't seem novel; humans have been thinking like the stuff we expose ourse [bsky, 1 points, 0 comments]
- New evidence shows chatbots are not only changing the way we write✍️, but the way we talk 🎙️too! arxiv.org/abs/2409.01754 [bsky, 1 points, 0 comments]
- oh yeah, the study says science academics. I guess I was leaning into LA&S folks. arxiv.org/abs/2409.01754 but yep, seems like the science folks are bypassing communicators to go straight to AI. awful [bsky, 1 points, 1 comments]
- Empirical evidence of Large Language Model’s influence on human spoken communication arxiv.org/pdf/2409.01754 [bsky, 1 points, 0 comments]
- There’s already some evidence of this: ‘Empirical evidence of Large Language Model’s influence on human spoken communication’ arxiv.org/pdf/2409.01754 [bsky, 1 points, 0 comments]
- Humans copying AI that’s copying humans. We’ve now gone full circle. I’ve delved into this swiftly but meticulously and think I can boast that I fully comprehend it. arxiv.org/abs/2409.01754 [bsky, 1 points, 1 comments]
- You sound like ChatGPT [lemmy, 1 points, 0 comments]
- And the preprint about humans emulating AI speech patterns. Except what if a lot of the YouTube videos that use the word "delve" more after ChatGPT was introduced are AI-generated? arxiv.org/abs/2409. [bsky, 0 points, 0 comments]
- @emollick Everyone is starting to sound like AI, even in spoken language Analysis of 280,000 transcripts of videos of talks & presentations from academic channels finds they increasingly used words [bsky, 0 points, 0 comments]
- Ou comment le langage des LLM commence à contaminer le langage des humains ? arxiv.org/pdf/2409.01754 [bsky, 0 points, 0 comments]
- Empirical evidence of Large Language Model's influence on human spoken communication #preprint arxiv.org/abs/2409.017... [bsky, 0 points, 0 comments]
- Well before the study it was a lot too- social media videos have gradually then all of a sudden turned over to generative scripting so they all sound alike. I'm wondering when ppl will see it and star [bsky, 0 points, 0 comments]
- Forget slop, remember my fears of unwittingly training myself to write like an AI, where it becomes a style of choice for humans as that is the content we’re exposed to? We’re there. arxiv.org/pdf/240 [bsky, 0 points, 0 comments]
- Link to paper arxiv.org/abs/2409.01754 [bsky, 0 points, 1 comments]
- Empirical evidence of Large Language Model’s influence on human spoken communication arxiv.org/pdf/2409.01754 [bsky, 0 points, 0 comments]
- Estudo mostra como cada vez mais acadêmicos estão falando como chatgpt. arxiv.org/pdf/2409.01754 [bsky, 0 points, 1 comments]
- This study suggests we are starting to use AI language, and sound more like AI: Empirical evidence of Large Language Model’s influence on human spoken communication arxiv.org/pdf/2409.01754 This was b [bsky, 0 points, 1 comments]
- arxiv.org/abs/2409.01754 Yes, AI generated text is changing how we speak and write. So did the invention of writing itself, radio, television, and telephones. The way we use language is constantly evo [bsky, 0 points, 0 comments]
- Empirical evidence of Large Language Model's influence on human spoken communication arxiv.org/abs/2409.01754 [bsky, 0 points, 0 comments]
- Boomerang is coming back to haunt us once more, poisoning real human language: 'Empirical evidence of Large Language Model’s influence on human spoken communication' (Max Planck Institute) arxiv.org/p [bsky, 0 points, 0 comments]
- arxiv.org/pdf/2409.01754 [bsky, 0 points, 0 comments]
- Mensen praten steeds vaker als ChatGPT. arxiv.org/pdf/2409.01754 [bsky, 0 points, 0 comments]
- here is what I am reading currently: arxiv.org/abs/2409.01754 I don't think its proof that everything I said is correct. The data that they are extracting the words from seem to be from thousands of h [bsky, 0 points, 1 comments]
- Namely https://arxiv.org/abs/2409.01754 https://arxiv.org/abs/2404.01268 https://www.nature.com/articles/s41598-023-30938-9 #AI #LLMs #AIHarms [bsky, 0 points, 0 comments]
Related