On the Dangers of Stochastic Parrots
2021/03/01 by Emily M. Bender, Timnit Gebru, Angelina McMillan-Major +1 · 6 voices · 5,980 citations
Computer Science · Engineering · #Artificial intelligence #Ask price #Big data #Business #Computer science #Data mining #Data science #Engineering #Finance #Management #Natural Language Processing Techniques #Software Engineering Research #Software deployment #Software engineering #Stakeholder #Topic Modeling #Work (physics)
paper · doi:10.1145/3442188.3445922
openalex publication_date 2021/03/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/06
Abstract
The past 3 years of work in NLP have been characterized by the development and deployment of ever larger language models, especially for English. BERT, its variants, GPT-2/3, and others, most recently Switch-C, have pushed the boundaries of the possible both through architectural innovations and through sheer size. Using these pretrained models and the methodology of fine-tuning them for specific tasks, researchers have extended the state of the art on a wide array of tasks as measured by leaderboards on specific benchmarks for English. In this paper, we take a step back and ask: How big is too big? What are the possible risks associated with this technology and what paths are available for mitigating those risks? We provide recommendations including weighing the environmental and financial costs first, investing resources into curating and carefully documenting datasets rather than ingesting everything on the web, carrying out pre-development exercises evaluating how the planned approach fits into research and development goals and supports stakeholder values, and encouraging research directions beyond ever larger language models.
Citations
Cited by
- A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation
- Spellburst: A Node-based Interface for Exploratory Creative Coding with Natural Language Prompts
- State media control influences large language models
- Enhancing linguistic research through critical use of race and ethnicity information
- Shared sensitivity to data distribution during learning in humans and transformer networks
- Delving into LLM-assisted writing in biomedical publications through excess vocabulary
- Rethinking Search: Making Domain Experts out of Dilettantes
- Multimodal large language models can make context-sensitive hate speech evaluations aligned with human judgement
- Agentic LLMs in the supply chain: towards autonomous multi-agent consensus-seeking
- What Makes a Good Theory, and How Do We Make a Theory Good?
- To Improve Literacy, Improve Equality in Education, Not Large Language Models
- ChatGPT’s interpretation of contested memoryscapes favours the voice of current governments, capitalism and the far-right
- Resisting Dehumanization in the Age of “AI”
- A Pedagogy of the Inevitable
- AI generates covertly racist decisions about people based on their dialect
- Hallucinating with AI: Distributed Delusions and “AI Psychosis”
- Critical, but constructive: defining, detecting, and addressing bias in Computational Social Science
- A Taxonomy of Linguistic Expressions That Contribute To Anthropomorphism of Language Technologies
- Why mental metaphors do not help us understand chatbot mistakes
- Whose news? Critical methods for assessing bias in large historical datasets
- Generative AI for Economic Research: Use Cases and Implications for Economists
- Gender bias and stereotypes in Large Language Models
- Generative AI Meets Open-Ended Survey Responses: Research Participant Use of AI and Homogenization
- Assisting or resisting patriarchy? a critical discourse analysis of chatgpt’s responses on feminism
- AI refusal in higher education: the right to refuse, the duty to understand and the diagnostic value of non-use
- Data Feminism for AI
- Human tests for machine models: What lies “Beyond the Imitation Game”?
- The Carbon Footprint of Machine Learning Training Will Plateau, Then Shrink
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- Simulacra as conscious exotica
- AI and Blackness: Towards moving beyond bias and representation
- AI Art and its Impact on Artists
- Reclaiming AI as a Theoretical Tool for Cognitive Science
- How widespread use of generative AI for images and video can affect the environment and the science of ecology
- Is Ockham’s razor losing its edge? New perspectives on the principle of model parsimony
- Large Language Models for Text Classification: From Zero-Shot Learning to Instruction-Tuning
- Machine Bias. How Do Generative Language Models Answer Opinion Polls? <sup/>
- “I have always found the whole area a minefield”: Wikidata, historical lives, and knowledge infrastructure
- Using Artificial Intelligence in TESOL: Some Ethical and Pedagogical Considerations
- A tutorial on open-source large language models for behavioral science
- Does ChatGPT have semantic understanding? A problem with the statistics-of-occurrence strategy
- Contextualizing predictive minds
- Medical large language models are vulnerable to data-poisoning attacks
- Von Menschen und Maschinen: Psychologiehistorische Reflexionen über Künstliche Intelligenz
- How to train your stochastic parrot: large language models for political texts
- Role play with large language models
- Math Education Digital Shadows for Investigating Learning with GenAI: Mathematics Performance, Anxiety, and Confidence in LLMs
- Leveraging AI for democratic discourse: Chat interventions can improve online political conversations at scale
- Analyzing the Ethical Logic of Eight Large Language Models
- Simulating Subjects: The Promise and Peril of Artificial Intelligence Stand-Ins for Social Agents and Interactions
- Updating “The Future of Coding”: Qualitative Coding with Generative Large Language Models
- Integrating Generative Artificial Intelligence into Social Science Research: Measurement, Prompting, and Simulation
- Teachy Mini: Development and Preliminary Evaluation of a Knowledge-Based Generative Social Robot for Higher Education
- A Unified Moral-Value Dataset for Instruction Tuning
- From Static Bibliometrics to Dynamic Knowledge Graphs: An LLM-Powered Framework for Modernizing Science, Technology, and Innovation (STI) Analytics
- What Does the Credential Still Certify? Cognitive Stewardship for AI-Mediated Education
- FMRP-LEAN: A HIPAA-Compliant AI-Augmented LIMS Architecture for End-to-End Clinical Assay Workflow Optimization
- Toward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference Alignment
- To Police or to Guide: How Higher Education Computer Science Instructors Design and Implement Generative AI Policies
- Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions
- Conjuring algorithms: Understanding the tech industry as stage magicians
- AutoJourn: Multi-Perspective Summarisation, Bias Detection and Bias Neutralisation for LLM-Generated News in Automated Journalism
- The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits
- Participatory provenance as representational auditing for AI-mediated public consultation
- Enhancing Small Language Models Reasoning through Knowledge Graph Grounding
- When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation
- Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins
- Auditing Differential Visibility of Political Content on TikTok
- Breakdowns for Human-Machine Creative Reflexivity
- Mapping the Narrow Corridor with Large Language Models
- Auditing Retrieval-Augmented LLM Hypotheses for Longitudinal Cell Painting Morphology
- A Human-Centric Evaluation of a Retrieval-Augmented Generation System for Explaining Quebec Insurance Contracts
- Linear representations of grammaticality in neural language models
- Data and trained models for "Empirical Evidence of Large Language Model's Influence on Human Spoken Communication"
- Neutrality Bites: Gender Representation in AI-Generated Animal Stories
- Directional Hallucinations: Ideological Drift in News-Grounded LLM Question Answering
- Creating Group Rules with AI: Human-AI Collaboration in WhatsApp Moderation
- Democratizing AI with Small Language Models: Structured Benchmarking and Parameter-Efficient Fine-Tuning for Local Deployment
- Semantic Field Theory: Historical Origin, Higher-Order Interaction, and Stabilized Semantic Inference
- Machine understanding
- Anthropomorphic Behaviors of AI
- Brainrot: Deskilling and Addiction are Overlooked AI Risks
- Reckoning with the Political Economy of AI: Avoiding Decoys in Pursuit of Accountability
- Terms of (Ab)Use: An Analysis of GenAI Services
- Failure of contextual invariance in large language models
- Generative AI & Fictionality: How Novels Power Large Language Models
- Subliminal Effects in Your Data: A General Mechanism via Log-Linearity
- How Human is AI? Examining the Impact of Emotional Prompts on Artificial and Human and Responsiveness
- Big AI is accelerating the metacrisis: What can we do?
- Irresponsible AI: big tech's influence on AI research and associated impacts
- The ‘design features’ of language revisited
- Epistemological Fault Lines Between Human and Artificial Intelligence
- Empowerment Gain and Causal Model Construction: Children and adults are sensitive to controllability and variability in their causal interventions
- The Risks of Industry Influence in Tech Research
- Surface Reading LLMs: Synthetic Text and its Styles
- Continuous Autoregressive Language Models
- Everyone prefers human writers, including AI
- AI, Digital Platforms, and the New Systemic Risk
- A Taxonomy of Transcendence
- Large Language Models Do Not Simulate Human Psychology
- What Does 'Human-Centred AI' Mean?
- The Homogenizing Effect of Large Language Models on Human Expression and Thought
- Misinformation by Omission: The Need for More Environmental Transparency in AI
- AI as Governance
- Stochastic Chameleons: Irrelevant Context Hallucinations Reveal Class-Based (Mis)Generalization in LLMs
- A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms
- Not Minds, but Signs: Reframing LLMs through Semiotics
- True Zero-Shot Inference of Dynamical Systems Preserving Long-Term Statistics
- Archaeology in the AI Era: Demystifying powerful and problematic systems shaping the future of the past
- LLM Social Simulations Are a Promising Research Method
- The ethics of AI or techno-solutionism? UNESCO’s policy guidance on AI in education
- Generative AI and its dilemmas: exploring AI from a translanguaging perspective
- Palatable Conceptions of Disembodied Being: Terra Incognita in the Space of Possible Minds
- Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
- The Widespread Adoption of Large Language Model-Assisted Writing Across Society
- "Don't Forget the Teachers": Towards an Educator-Centered Understanding of Harms from Large Language Models in Education
- Culture is Not Trivia: Sociocultural Theory for Cultural NLP
- Emergent Symbolic Mechanisms Support Abstract Reasoning in Large Language Models
- Provocations from the Humanities for Generative AI Research
- Stop treating `AGI' as the north-star goal of AI research
- Unmasking Conversational Bias in AI Multiagent Systems
- Actions Speak Louder than Words: Agent Decisions Reveal Implicit Biases in Language Models
- Aetheria: A multimodal interpretable content safety framework based on multi-agent debate and collaboration
- From Accuracy to Impact: The Impact-Driven AI Framework (IDAIF) for Aligning Engineering Architecture with Theory of Change
- The Future of NLP may not be at NLP Conferences: Scholarly Migration Patterns in Natural Language Processing
- Misclassification in Automated Content Analysis Causes Bias in Regression. Can We Fix It? Yes We Can!
- Co-constructing a virtual translanguaging space: an interpretative phenomenological analysis of human–AI communication
- When Machines Get It Wrong: Large Language Models Perpetuate Autism Myths More Than Humans Do
- Stand-Alone Complex or Vibercrime? Exploring the adoption and innovation of GenAI tools, coding assistants, and agents within cybercrime ecosystems
- Developing evaluative judgement for a time of generative artificial intelligence
- A few notes on the scalar foundations of foundation models
- From revolution to evolution: What generative AI really means for language learning
- The Law of Multi-Model Collaboration: Scaling Limits of Model Ensembling for Large Language Models
- Computational structuralism: Toward a formal theory of meaning in the age of digital intelligence
- Generative Aesthetics: On formal stuckness in AI verse
- Alignment Is Not Enough: A Relational Framework for Moral Standing in Human-AI Interaction
- Neutralising the propaganda machine: proposals for cyber-propaganda de-escalation
- The carbon emissions of writing and illustrating are lower for AI than for humans
- Language machines: Toward a linguistic anthropology of large language models
- Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics
- BenCSSmark: Making the Social Sciences Count in LLM Research
- Between fact and fairy: tracing the hallucination metaphor in AI discourse
- Addressing "Documentation Debt" in Machine Learning Research: A Retrospective Datasheet for BookCorpus
- Lessons from Neuroscience for AI: How integrating Actions, Compositional Structure and Episodic Memory could enable Safe, Interpretable and Human-Like AI
- BitFlipScope: Scalable Fault Localization and Recovery for Bit-Flip Corruptions in LLMs
- The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards
- Sociotechnical Implications of Generative Artificial Intelligence for Information Access
- Generative AI and Research Integrity
- Looking Beyond the Hype: Understanding the Effects of AI on Learning
- Reimagining Open Source and Openness in AI: Co-Creating Responsible Technological Futures
- False understanding in AI-assisted physics problem solving: a theoretical framework
- Efficient Multi-Modal Embeddings from Structured Data
- The New Positivism: A Hegelian Critique
- Debating AI in Archaeology: applications, implications, and ethical considerations
- Generative AI for Requirements Engineering: A Systematic Literature Review
- Testing of detection tools for AI-generated text
- Heads we win, tails you lose: AI detectors in education
- The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play
- On the Opportunities and Risks of Foundation Models
- Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed
- Generative Artificial Intelligence in Scientific Research: Individual Benefits, Collective Risks, and a Framework for Responsible Research with AI
- Reading Without a Reader: Large Language Models Collapse Reading and Writing into a Single Entangled Code
- Some hypotheses on how chatbots work in problem-solving-driven conversations. Large Language Models as confirmation of the Innovation Illusion
- TimeCapsule: Generative Hallucination as a Method for Historical Sensemaking
- Coherent without Grounding, Grounded without Success: The Bidirectional Coherence Paradox in Artificial Epistemic Agents
- Responsible Intelligence in Practice: A Fairness Audit of Open Large Language Models for Library Reference Services
- Auto-Prompting with Retrieval Guidance for Frame Detection in Logistics
- Few-Shot Learning of a Graph-Based Neural Network Model Without Backpropagation
- Plausibility as Failure: How LLMs and Humans Co-Construct Epistemic Error
- Blog Data Showdown: Machine Learning vs Neuro-Symbolic Models for Gender Classification
- Epistemic diversity across language models mitigates knowledge collapse
- A Multifaceted Analysis of Social Biases in Large Language Models
- Evaluation of AI Ethics Tools in Language Models: A Developers' Perspective Case Stud
- Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
- SafeGen: Embedding Ethical Safeguards in Text-to-Image Generation
- Mitigating Social Bias in English and Urdu Language Models Using PRM-Guided Candidate Selection and Sequential Refinement
- Mind the Gap! Pathways Towards Unifying AI Safety and Ethics Research
- Evaluation of Text Generation: A Survey
- A Categorical Analysis of Large Language Models and Why LLMs Circumvent the Symbol Grounding Problem
- Short-Context Dominance: How Much Local Context Natural Language Actually Needs?
- LUNE: Efficient LLM Unlearning via LoRA Fine-Tuning with Negative Examples
- Recent Advances in Natural Language Processing via Large Pre-trained Language Models: A Survey
- The Affective Scaffolding of Grief in the Digital Age: The Case of Deathbots
- AGI Requires a Coordination Layer on Top of Pattern Repositories
- AfriStereo: A Culturally Grounded Dataset for Evaluating Stereotypical Bias in Large Language Models
- Finetuned Language Models Are Zero-Shot Learners
- Group Selection as a Safeguard Against AI Substitution
- Is Lying Only Sinful in Islam? Exploring Religious Bias in Multilingual Large Language Models Across Major Religions
- Understanding Down Syndrome Stereotypes in LLM-Based Personas
- Advancing Academic Chatbots: Evaluation of Non Traditional Outputs
- ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization
- Silhouette-based Gait Foundation Model
- Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation Benchmarking
- Writing in Symbiosis: Mapping Human Creative Agency in the AI Era
- Rethinking AI Evaluation in Education: The TEACH-AI Framework and Benchmark for Generative AI Assistants
- AI Consciousness and Existential Risk
- Understanding the Staged Dynamics of Transformers in Learning Latent Structure
- Empathetic Cascading Networks: A Multi-Stage Prompting Technique for Reducing Social Biases in Large Language Models
- The Semiotic Channel Principle: Measuring the Capacity for Meaning in LLM Communication
- Foundations of Artificial Intelligence Frameworks: Notion and Limits of AGI
- Proactive Defense: Compound AI for Detecting Persuasion Attacks and Measuring Inoculation Effectiveness
- AI Art is Theft: Labour, Extraction, and Exploitation: Or, On the Dangers of Stochastic Pollocks
- Alignment Faking - the Train -> Deploy Asymmetry: Through a Game-Theoretic Lens with Bayesian-Stackelberg Equilibria
- Narratives to Numbers: Large Language Models and Economic Policy Uncertainty
- Generative AI in Sociological Research: State of the Discipline
- Stable diffusion models reveal a persisting human and AI gap in visual creativity
- Unsupervised and Distributional Detection of Machine-Generated Text
- A time for monsters: Organizational knowing after LLMs
- Efficiency Will Not Lead to Sustainable Reasoning AI
- Writing With Machines and Peers: Designing for Critical Engagement with Generative AI
- Strategic Innovation Management in the Age of Large Language Models Market Intelligence, Adaptive R&D, and Ethical Governance
- Bias in, Bias out: Annotation Bias in Multilingual Large Language Models
- GPT models for text annotation: An empirical exploration in public policy research
- Just Asking Questions: Doing Our Own Research on Conspiratorial Ideation by Generative AI Chatbots
- Developing a Grounded View of AI
- Bias and Fairness in Large Language Models: A Survey
- Pretrained Transformers as Universal Computation Engines
- Cultural Cartography with Word Embeddings
- The Fallacy of AI Functionality
- How Scientists Use Large Language Models to Program
- Competing narratives in AI ethics: a defense of sociotechnical pragmatism
- Language variation and algorithmic bias: understanding algorithmic bias in British English automatic speech recognition
- Towards a Posthumanist Critique of Large Language Models
- Reducing Hallucinations in LLM-Generated Code via Semantic Triangulation
- How Language Directions Align with Token Geometry in Multilingual LLMs
- From Passive to Persuasive: Localized Activation Injection for Empathy and Negotiation
- Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders
- On the Notion that Language Models Reason
- On the Measure of a Model: From Intelligence to Generality
- SlideBot: A Multi-Agent Framework for Generating Informative, Reliable, Multi-Modal Presentations
- Automatic Minds: Cognitive Parallels Between Hypnotic States and Large Language Model Processing
- Equilibrium Dynamics and Mitigation of Gender Bias in Synthetically Generated Data
- Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
- AI and extended authenticity: autism as a case study
- Unreliable minds, unreliable machines: dyslexic memory, ChatGPT, and the epistemic disobedience of generative AI
- Going with the Mainstream: Exploring GPT Representation of Journalistic Culture
- Sustainable AI: Environmental Implications, Challenges and Opportunities
- LLMs and the increasing role of the humanities in the digital age
- Faithful-First Reasoning, Planning, and Acting for Multimodal LLMs
- FLEX: Unifying Evaluation for Few-Shot NLP
- Generative AI poses ethical challenges for open science
- What can LLMs tell us about the mechanisms behind polarity illusions in humans? Experiments across model scales and training steps
- Agentic AI Sustainability Assessment for Supply Chain Document Insights
- Setting ε is not the Issue in Differential Privacy
- Multi-Reward GRPO Fine-Tuning for De-biasing Large Language Models: A Study Based on Chinese-Context Discrimination Data
- Who Gets Heard? Rethinking Fairness in AI for Music Systems
- Large language models replicate and predict human cooperation across experiments in game theory
- Design principles for text-to-image generative artificial intelligence creativity support tools for visual design
- AI as We Describe It: How Large Language Models and Their Applications in Health are Represented Across Channels of Public Discourse
- Promoting Sustainable Web Agents: Benchmarking and Estimating Energy Consumption through Empirical and Theoretical Analysis
- From Five Dimensions to Many: Large Language Models as Precise and Interpretable Psychological Profilers
- Examining Algorithms in the Light of their Ground Truth Datasets: Results, Objections, and Avenues of Reflection
- Vibe Learning: Education in the age of AI
- A Detailed Study on LLM Biases Concerning Corporate Social Responsibility and Green Supply Chains
- AI Progress Should Be Measured by Capability-Per-Resource, Not Scale Alone: A Framework for Gradient-Guided Resource Allocation in LLMs
- Sustainability of Machine Learning-Enabled Systems: The Machine Learning Practitioner's Perspective
- Understanding, Demystifying and Challenging Perceptions of Gig Worker Vulnerabilities
- Characterizing Selective Refusal Bias in Large Language Models
- Cross-Platform Evaluation of Reasoning Capabilities in Foundation Models
- Inverse Knowledge Search over Verifiable Reasoning: Synthesizing a Scientific Encyclopedia from a Long Chains-of-Thought Knowledge Base
- Understanding Hardness of Vision-Language Compositionality from A Token-level Causal Lens
- Position: Biology is the Challenge Physics-Informed ML Needs to Evolve
- Cognivia: A Cognitive Behavioral Therapy Copilot for Evidence-Based Mental Healthcare
- Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing
- Faster, Higher, Stronger? The Impact of GenAI on Knowledge Work Productivity - Evidence from the Field
- AI readiness is not enough: Towards purpose-led AI governance in projects
- Insolvent
- Factuality challenges in the era of large language models and opportunities for fact-checking
- (A)I Cannot See Them: A Situated Reflection on the Simulation of Historical Figures
- AI ethical challenges: a perspective of AI developers in postcolonial countries
- Overcoming distortion in multidimensional predictive representation
- Spectral imaginings and sympoietic creativity: AI hallucinations and the ethics of posthuman creativity
- What Is Rationality, Whom Is It Ascribed To, and Why Does It Matter? Evidence From Internet Text for 66 Social Groups and 101 Occupations
- Grounding large language models in an anthropological knowledge graph: A neuro-symbolic approach
- Log in, lie down: ethics and the digital turn in psychotherapy
- Cybernetic Resurgences: Machine Music Beyond AI Slop
- FinBERT: A Large Language Model for Extracting Information from Financial Text*
- Large language models and the problem of rhetorical debt
- Which educational approaches predict students’ generative AI confidence and responsibility?
- The ethics of ChatGPT in medicine and healthcare: a systematic review on Large Language Models (LLMs)
- Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition
- When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
- Studying Reddit: A Systematic Overview of Disciplines, Approaches, Methods, and Ethics
- MERLOT: Multimodal Neural Script Knowledge Models
- Generative AI and the politics of visibility
- A Mixed‐Methods Needs Assessment of Frontline Communities: Insights for Engagement and Partnerships Between Communities and Intermediary Organizations
- Large Language Models: A Paradigm Shift for Dementia Diagnosis and Care
- Using generative AI responsibly in writing and publishing: the NZVJ's policy and recommendations
- Large models of what? Mistaking engineering achievements for human linguistic agency
- Metaethical perspectives on ‘benchmarking’ AI ethics
- Reframing Diversity in Computing on the Basis of Genders
- Equipping Speech-Language Clinicians for the Critical Appraisal of an Artificial Intelligence–Driven, Evidence-Based Future
- Large language models propagate race-based medicine
- A Copious Void: Rhetoric as Artificial Intelligence 1.0
- How Does Narrow AI Impact Human Creativity?
- Training Computational Social Science PhD Students for Academic and Non-Academic Careers
- A reporting checklist for large language models in behavioural science
- AI Narrative Breakdown. A Critical Assessment of Power and Promise
- Collage is the New Writing: Exploring the Fragmentation of Text and User Interfaces in AI Tools
- The TESCREAL bundle: Eugenics and the promise of utopia through artificial general intelligence
- Can AI systems have free will?
- AI and Image
- The social AI author: modeling creativity and distinction in simulated cultural fields
- How to Use Generative AI in Educational Research
- Hijacking algorithmic bias: analyzing the political discourse around ChatGPT on social media
- Large language models, social demography, and hegemony: comparing authorship in human and synthetic text
- Survey of Cultural Awareness in Language Models: Text and Beyond
- Why ‘open’ AI systems are actually closed, and why this matters
- Implicit Values Embedded in How Humans and LLMs Complete Subjective Everyday Tasks
- Beyond AI as an environmental pharmakon: Principles for reopening the problem-space of machine learning's carbon footprint
- Generative AI and the Automating of Academia
- Pixels and Predictions: Potential of GPT-4V in Meteorological Imagery Analysis and Forecast Communication
- “Your friendly AI assistant”: the anthropomorphic self-representations of ChatGPT and its implications for imagining AI
- Language in the age of AI technology: From human to non-human authenticity, from public governance to privatised assemblages
- ATLAS: Harnessing retrieval-augmented generation (RAG)
- AI as Compensatory Infrastructure: Sensemaking, Responsibility, and Labour in Value-Driven Organisations
- Using AI-assisted coding to build digital tools for natural history collections
- Truth-Aware Decoding: A Program-Logic Approach to Factual Language Generation
- Can Large Language Models Transform Computational Social Science?
- The cognitive biases that may exacerbate inflationary and deflationary positions about large language models
- What if AI systems weren't chatbots?
- The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment
- Beyond ChatGPT: a review of the use of AI tools in biological education
- What Vision-Language Models `See' when they See Scenes
- GeDi: Generative Discriminator Guided Sequence Generation
- Use and usability: concepts of representation in philosophy, neuroscience, cognitive science, and computer science
- How Does Machine Learning Manage Complexity?
- The Rise of AI in Weather and Climate Information and its Impact on Global Inequality
- Transformers Can Learn Rules They've Never Seen: Proof of Computation Beyond Interpolation
- Toward Guarantees for Clinical Reasoning in Vision Language Models via Formal Verification
- Program Synthesis with Large Language Models
- On the Use of Large Language Models for Qualitative Synthesis
- Political censorship in large language models originating from China
- Enabling Sustainable Clouds: The Case for Virtualizing the Energy System
- LAMUS: A Large-Scale Corpus for Legal Argument Mining from U.S. Caselaw using LLMs
- Enhancing Linguistic Competence of Language Models through Pre-training with Language Learning Tasks
- Multitask Prompted Training Enables Zero-Shot Task Generalization
- The use of LLMs to annotate data in management research: Foundational guidelines and warnings
- Do Large Language Models Grasp The Grammar? Evidence from Grammar-Book-Guided Probing in Luxembourgish
- Politically Speaking: LLMs on Changing International Affairs
- A word association network methodology for evaluating implicit biases in LLMs compared to humans
- HiMAE: Hierarchical Masked Autoencoders Discover Resolution-Specific Structure in Wearable Time Series
- Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
- Human-Level Reasoning: A Comparative Study of Large Language Models on Logical and Abstract Reasoning
- From Stochasticity to Signal: A Bayesian Latent State Model for Reliable Measurement with LLMs
- The seven roles of generative AI: Potential & pitfalls in combatting misinformation
- “It doesn't give a s*** about Arabic or English”: Semiotic ideologies and demarcation among LLM engineers in Amman, Jordan
- Power and politics in framing bias in Artificial Intelligence policy
- Assessing the Relational Abilities of Large Language Models and Large Reasoning Models
- Beyond Understanding: Evaluating the Pragmatic Gap in LLMs' Cultural Processing of Figurative Language
- Generative AI and English language teaching: A global Englishes perspective
- Mapping the unseen in practice: comparing latent Dirichlet allocation and BERTopic for navigating topic spaces
- Critical GenAI Literacy: Postdigital Configurations
- Rethinking Error: “Hallucinations” and Epistemological Indifference
- “Who's Doing the Thinking Here?”: A Pedagogy-First Approach to Integrating Large Language Models in Higher Education
- Education Paradigm Shift To Maintain Human Competitive Advantage Over AI
- Assessing the Human-Likeness of LLM-Driven Digital Twins in Simulating Health Care System Trust
- In Generative AI We (Dis)Trust? Computational Analysis of Trust and Distrust in Reddit Discussions
- Actionable Cybersecurity Notifications for Smart Homes: A User Study on the Role of Length and Complexity
- Opening up ChatGPT: Tracking openness, transparency, and accountability in instruction-tuned text generators
- Race and Gender in LLM-Generated Personas: A Large-Scale Audit of 41 Occupations
- Preventing Catastrophic Forgetting: Behavior-Aware Sampling for Safer Language Model Fine-Tuning
- What's in the Box? A Preliminary Analysis of Undesirable Content in the Common Crawl Corpus
- Integrating Machine Learning into Belief-Desire-Intention Agents: Current Advances and Open Challenges
- Black Box Absorption: LLMs Undermining Innovative Ideas
- On the Detectability of LLM-Generated Text: What Exactly Is LLM-Generated Text?
- EQPO: Equitable Group Relative Policy Optimization for Clinical Reasoning
- Synthetic social data: trials and tribulations
- TheMCPCompany: Creating General-purpose Agents with Task-specific Tools
- Large language models in medicine
- ActivationReasoning: Logical Reasoning in Latent Activation Spaces
- Identity-Aware Large Language Models require Cultural Reasoning
- Cultural Alien Sampler: Open-ended art generation balancing originality and coherence
- Generative AI, Quadruple Deception & Trust
- Alterity and kinship: co-writing posthumanist speculative nonfiction with AI
- Techno-material entanglements and the social organisation of difference
- Grover's Algorithm for Question Answering
- AFRICAPTION: Establishing a New Paradigm for Image Captioning in African Languages
- The Evolving Nature of Latent Spaces: From GANs to Diffusion
- Knowing the Facts but Choosing the Shortcut: Understanding How Large Language Models Compare Entities
- Zero-Shot Performance Prediction for Probabilistic Scaling Laws
- The debate over understanding in AI’s large language models
- DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation
- On the controversiality of AI: The controversy is not the situation
- Artificial creativity: can there be creativity without cognition?
- Product Manager Practices for Delegating Work to Generative AI: "Accountability must not be delegated to non-human actors"
- How Social Should AI Be?
- On the Design and Evaluation of Human-centered Explainable AI Systems: A Systematic Review and Taxonomy
- Too Open for Opinion? Embracing Open-Endedness in Large Language Models for Social Simulation
- Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning
- Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and Rectification
- Attacks by Content: Automated Fact-checking is an AI Security Issue
- Detecting Gender Stereotypes in Scratch Programming Tutorials
- A Two-Step, Multidimensional Account of Deception in Language Models
- Making Power Explicable in AI: Analyzing, Understanding, and Redirecting Power to Operationalize Ethics in AI Technical Practice
- ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test
- The Achilles' Heel of LLMs: How Altering a Handful of Neurons Can Cripple Language Abilities
- Subspace Regularizers for Few-Shot Class Incremental Learning
- CauchyNet: Compact and Data-Efficient Learning using Holomorphic Activation Functions
- The Dawn of the Human-Machine Era: A forecast of new and emerging language technologies
- Coding Inequity: Assessing GPT-4’s Potential for Perpetuating Racial and Gender Biases in Healthcare
- Contemplative Superalignment
- The Evolution of Artificial Intelligence Paradigms – Implications for Performance, Scalability, and Responsible Deployment
- Inverse Language Modeling towards Robust and Grounded LLMs
- The social consequences of AI delegation
- MarxistLLM: Fine-tuning a language model with a Marxist worldview
- Viola: A Topic Agnostic Generate-and-Rank Dialogue System
- The Geometry of Reasoning: Flowing Logics in Representation Space
- Toward a biosemiotics framework for AI: Folding and the dynamics of meaning
- CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMs
- The Environmental Impacts of Machine Learning Training Keep Rising Evidencing Rebound Effect
- Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models
- Creating a large language model of a philosopher
- Contrastive Language-Image Pre-training for the Italian Language
- ISMIE: A Framework to Characterize Information Seeking in Modern Information Environments
- Contrastive Decoding for Synthetic Data Generation in Low-Resource Language Modeling
- Role-Conditioned Refusals: Evaluating Access Control Reasoning in Large Language Models
- All Claims Are Equal, but Some Claims Are More Equal Than Others: Importance-Sensitive Factuality Evaluation of LLM Generations
- Prompt Optimization Across Multiple Agents for Representing Diverse Human Populations
- Archival Reference Services in an Age of AI
- From authority to similarity: How Google transformed its knowledge infrastructure using computer vision
- What is a fact in the age of generative AI? Fact-checking as an epistemological lens
- Embracing Dialectic Intersubjectivity: Coordination of Differential Perspectives in Content Analysis With LLM Persona Simulation
- ParadAIse L0st?
- Envisioning Information Access Systems: What Makes for Good Tools and a Healthy Web?
- Revealing emergent human-like conceptual representations from language prediction
- Exposing Citation Vulnerabilities in Generative Engines
- PTEB: Towards Robust Text Embedding Evaluation via Stochastic Paraphrasing at Evaluation Time with LLMs
- The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives
- EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models
- Limitations of Current Evaluation Practices for Conversational Recommender Systems and the Potential of User Simulation
- InvThink: Premortem Reasoning for Safer Language Models
- A Set of Quebec-French Corpus of Regional Expressions and Terms
- COLE: a Comprehensive Benchmark for French Language Understanding Evaluation
- Large Language Models Hallucination: A Comprehensive Survey
- Mechanistic Interpretability of Socio-Political Frames in Language Models
- The Illusion of Artificial Inclusion
- Evaluating Large Language Models Trained on Code
- Experimental narratives: A comparison of human crowdsourced storytelling and AI storytelling
- Posthuman Cartography? Rethinking Artificial Intelligence, Cartographic Practices, and Reflexivity
- ChatGPT: deconstructing the debate and moving it forward
- Fairness Testing in Retrieval-Augmented Generation: How Small Perturbations Reveal Bias in Small Language Models
- Extreme Self-Preference in Language Models
- Event Tokenization and Next-Token Prediction for Anomaly Detection at the Large Hadron Collider
- QFrBLiMP: a Quebec-French Benchmark of Linguistic Minimal Pairs
- Toxicity in Online Platforms and AI Systems: A Survey of Needs, Challenges, Mitigations, and Future Directions
- Comparing Open-Source and Commercial LLMs for Domain-Specific Analysis and Reporting: Software Engineering Challenges and Design Trade-offs
- KI og effektivisering: Hvordan GPT-4 fant 155.000 overflødige årsverk i norsk offentlig sektor
- Bridging the behavior-neural gap: A multimodal AI reveals the brain's geometry of emotion more accurately than human self-reports
- VeriLLM: A Lightweight Framework for Publicly Verifiable Decentralized Inference
- The problem of alignment
- AI innovation at the boundaries: Justifying a generative AI decision support tool
- Synthetic media and computational capitalism: towards a critical theory of artificial intelligence
- Science as a vocation redux: outsourcing the logic of discovery to AI
- Using conversational AI to reduce science skepticism
- Fostering Robots: A Governance-First Conceptual Framework for Domestic, Curriculum-Based Trajectory Collection
- A Cross-Lingual Analysis of Bias in Large Language Models Using Romanian History
- Transphobia Is in the Eye of the Prompter: Trans-Centered Perspectives on Large Language Models
- Virtues for AI
- Exploring the scope of generative AI in literature review development
- Googling Politics? Comparing Five Computational Methods to Identify Political and News-related Searches from Web Browser Histories
- The promise and challenges of generative AI in education
- The linguistic dead zone of value-aligned agency, natural and artificial
- RefAM: Attention Magnets for Zero-Shot Referral Segmentation
- AI as Agency Without Intelligence: on ChatGPT, Large Language Models, and Other Generative Models
- Encountering Artificial Intelligence: Ethical and Anthropological Investigations
- Copilots for Linguists
- Automated Visual Analysis for the Study of Social Media Effects: Opportunities, Approaches, and Challenges
- Advancing Automated Content Analysis for a New Era of Media Effects Research: The Key Role of Transfer Learning
- Ethische Betrachtungen der automatisierten Textanalyse : Fortschritte, Risiken und die Notwendigkeit eines Gleichgewichts
- A TECNOLOGIA NÃO É UMA FERRAMENTA: CONTRIBUIÇÕES DA PEDAGOGIA CRÍTICA PARA A COMPREENSÃO DA NÃO NEUTRALIDADE TECNOLÓGICA NA EDUCAÇÃO
- Closed Rooms and Oracles: the poetics of vector space in Sam Riviere’s Conflicted Copy
- The Philosophy of Language Models
- Green Prompt Engineering: Investigating the Energy Impact of Prompt Design in Software Engineering
- From Superficial Outputs to Superficial Learning: Risks of Large Language Models in Education
- The Bias is in the Details: An Assessment of Cognitive Bias in LLMs
- Culture machines
- Does AI Coaching Prepare us for Workplace Negotiations?
- Quantum est in Libris: Navigating Archives with GenAI, Uncovering Tension Between Preservation and Innovation
- Beyond Chat: a Framework for LLMs as Human-Centered Support Systems
- Longitudinal Monitoring of LLM Content Moderation of Social Issues
- Artificial Intelligence’s new clothes? A system technology perspective
- The Knowledge-Behaviour Disconnect in LLM-based Chatbots
- GPT and Prejudice: A Sparse Approach to Understanding Learned Representations in Large Language Models
- Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports
- Confidence Calibration in Large Language Model-Based Entity Matching
- Generative AI and linguistic diversity in academic writing and publishing
- We are Building Gods: AI as the Anthropomorphised Authority of the Past
- Using degree apprenticeships to shape the future of traditional undergraduate degrees
- When AI Does the Work, What Is Learning For? Post-Instrumental Learning and the Risk of Capacity Dissolution
- Asymmetric Communication: Large Language Models and Language Games
- Cut the bullshit: why GenAI systems are neither collaborators nor tutors
- Newer, Larger, Better? A Critique of the Unreflective LLM Adoption in Communication Research
- Scaling, Lock-In, and Proxy Compliance: A Political Economy of Responsible AI
- The seer and the seen: Surveying Palantir’s surveillance platform
- Generative AI and academic scientists in US universities: Perception, experience, and adoption intentions
- Navigating ethical challenges in generative AI-enhanced research: The ETHICAL framework for responsible generative AI use
- Investigating machine moral judgement through the Delphi experiment
- Auditing the bias of conversational AI systems in occupational recommendations: a novel approach to bias quantification via Holland’s theory
- The science publishing manifesto: AI moves fast, science publishing must too
- Pleidooi voor nutteloze beschaving
- Endangered Judgment: Joseph Weizenbaum, Artificial Intelligence, and the Imperialism of Instrumental Reason
- Why language clouds our ascription of understanding, intention and consciousness
- Beyond path dependence: a methodological review for studying regional futures
- Blowing artificial intelligence bubbles: studies of technological hype
- Development and content validation of the CAREFUL-AI framework for evaluating AI-generated scientific manuscripts: an exploratory cross-platform study
- Language Models and Externalism: A Reply to Mandelkern and Linzen
- AI, autism, and the architecture of voice: from engineered exclusion to designed dignity
- Máquinas que piensan
- Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources
- When AI Meets Science: Research Diversity, Interdisciplinarity, Visibility, and Retractions across Disciplines in a Global Surge
- MS MARCO: Benchmarking Ranking Models in the Large-Data Regime
- The Ouroboros effect and heterodox domains
- ChatGPT: More Than a “Weapon of Mass Deception” Ethical Challenges and Responses from the Human-Centered Artificial Intelligence (HCAI) Perspective
- Artificial intelligence and the death of the academic author
- The uncontroversial ‘thingness’ of AI
- ED-AI Lit: An Interdisciplinary Framework for AI Literacy in Education
- The interplay of learning, analytics and artificial intelligence in education: A vision for hybrid intelligence
- Gen AI and research integrity: Where to now?
- Mussolini and ChatGPT. Examining the Risks of A.I. writing Historical Narratives on Fascism
- The steep cost of capture
- Minds
- A Feminist Account of Intersectional Algorithmic Fairness
- Auditing large language models: a three-layered approach
- Medicine as an Information Industry in the Age of Language Models
- Pensar con la mirada: cocreación humano-algorítmica y agencia distribuida en Visions of Destruction
- From moral panic to pragmatic governance: reframing AI’s societal impacts in employment, education, and ethics
- Bypassing Guardrails: Lessons Learned from Red Teaming ChatGPT
- Psychiatric Risk Associated With Large Language Model Chatbots: An Emerging Concern
- Borges and AI
- Anticipating Safety Issues in E2E Conversational AI: Framework and Tooling
- Between Imperative and Panic: Navigating AI's Educational Moment with Leon Furze
- Generative artificial intelligence and engineering education
- A content audit case study: AI-assisted processes for student-produced local newsrooms
- Ethics for the majority world: AI and the question of violence at scale
- Ghost in the cache: How data decay shapes the unseen landscape of AI memory
- Artificial intelligence for global health: cautious optimism with safeguards
- Towards Open-Ended Discovery for Low-Resource NLP
- DecipherGuard: Understanding and Deciphering Jailbreak Prompts for a Safer Deployment of Intelligent Software Systems
- Decoding Uncertainty: The Impact of Decoding Strategies for Uncertainty Estimation in Large Language Models
- Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
- The Even Sheen of AI: Kitsch, LLMs, and Homogeneity
- Designing Culturally Aligned AI Systems For Social Good in Non-Western Contexts
- Designing with Culture: How Social Norms Shape Trust and Preference in Health Chatbots
- Large Language Model probabilities cannot distinguish between possible and impossible language
- Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
- Trust Me, I Know This Function: Hijacking LLM Static Analysis using Bias
- Simulating a Bias Mitigation Scenario in Large Language Models
- Risk Assessment and Security Analysis of Large Language Models
- On the Creativity of Large Language Models
- Is Stack Overflow Obsolete? An Empirical Study of the Characteristics of ChatGPT Answers to Stack Overflow Questions
- Investigating ReLoRA: Effects on the Learning Dynamics of Small Language Models
- Harnessing the Power of AI in Qualitative Research: Role Assignment, Engagement, and User Perceptions of AI-Generated Follow-Up Questions in Semi-Structured Interviews
- Don't Change My View: Ideological Bias Auditing in Large Language Models
- Large Language Models Imitate Logical Reasoning, but at what Cost?
- Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
- Quantifying Language Disparities in Multilingual Large Language Models
- Human resource management in the age of generative artificial intelligence: Perspectives and research directions on ChatGPT
- Qualitative Research in an Era of AI: A Pragmatic Approach to Data Analysis, Workflow, and Computation
- Prompt Commons: Collective Prompting as Governance for Urban AI
- Uncertainty in Authorship: Why Perfect AI Detection Is Mathematically Impossible
- BioMetaphor: AI-Generated Biodata Representations for Virtual Co-Present Events
- Does Language Model Understand Language?
- The Prompt Engineering Report Distilled: Quick Start Guide for Life Sciences
- Bridging Cultural Distance Between Models Default and Local Classroom Demands: How Global Teachers Adopt GenAI to Support Everyday Teaching Practices
- Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
- GenAI Voice Mode in Programming Education
- Out of One, Many: Using Language Models to Simulate Human Samples
- Towards Trustworthy AI: Characterizing User-Reported Risks across LLMs "In the Wild"
- PromptGuard: An Orchestrated Prompting Framework for Principled Synthetic Text Generation for Vulnerable Populations using LLMs with Enhanced Safety, Fairness, and Controllability
- JUDGEBERT: Assessing Legal Meaning Preservation Between Sentences
- QFrCoLA: a Quebec-French Corpus of Linguistic Acceptability Judgments
- DExperts: Decoding-Time Controlled Text Generation with Experts and Anti-Experts
- Bias after Prompting: Persistent Discrimination in Large Language Models
- Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare
- That's So FETCH: Fashioning Ensemble Techniques for LLM Classification in Civil Legal Intake and Referral
- Neuro-Symbolic Frameworks: Conceptual Characterization and Empirical Comparative Analysis
- Measuring and mitigating overreliance to build human-compatible AI
- "Abuse Risks are Often Inherent to Product Features": Exploring AI Vendors' Bug Bounty and Responsible Disclosure Policies
- Stabilizing Deep Q-Learning with ConvNets and Vision Transformers under Data Augmentation
- Using generative artificial intelligence/ChatGPT for academic communication: Students' perspectives
- Drivel-ology: Challenging LLMs with Interpreting Nonsense with Depth
- Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain
- Do Language Models Follow Occam's Razor? An Evaluation of Parsimony in Inductive and Abductive Reasoning
- El peaje de los de abajo: midiendo el peaje lingüístico (comprensión, calidad y costo) del español colombiano frente a los LLM
- Generative AI in Heritage Practice: Improving the Accessibility of Heritage Guidance
- Mitigation of Gender and Ethnicity Bias in AI-Generated Stories through Model Explanations
- Should LLMs be WEIRD? Exploring WEIRDness and Human Rights in Large Language Models
- On the Alignment of Large Language Models with Global Human Opinion
- Automatic Identification and Description of Jewelry Through Computer Vision and Neural Networks for Translators and Interpreters
- Belief updating in AI‐risk debates: Exploring the limits of adversarial collaboration
- Who Gets Left Behind? Auditing Disability Inclusivity in Large Language Models
- Transforming Agency. On the mode of existence of Large Language Models
- Synthetic Founders: AI-Generated Social Simulations for Startup Validation Research in Computational Social Science
- Comparative Analysis of Large Language Models for the Machine-Assisted Resolution of User Intentions
- When Models Refuse: Political Steerability and Feature Richness as Measures of Ideological Depth
- When Does Pretraining Help? Assessing Self-Supervised Learning for Law and the CaseHOLD Dataset
- ExT5: Towards Extreme Multi-Task Scaling for Transfer Learning
- Automated Quality Assessment for LLM-Based Complex Qualitative Coding: A Confidence-Diversity Framework
- Correct-By-Construction: Certified Individual Fairness through Neural Network Training
- On the Future of Software Reuse in the Era of AI Native Software Engineering
- Benchmarking Hindi LLMs: A New Suite of Datasets and a Comparative Analysis
- Democracy-in-Silico: Institutional Design as Alignment in AI-Governed Polities
- A perishable ability? The future of writing in the face of generative artificial intelligence
- Insights into User Interface Innovations from a Design Thinking Workshop at deRSE25
- Synthetic Replacements for Human Survey Data? The Perils of Large Language Models
- What do language models model? Transformers, automata, and the format of thought
- The AI gambit: leveraging artificial intelligence to combat climate change—opportunities, challenges, and recommendations
- A Theory of Information, Variation, and Artificial Intelligence
- LLMs and Agentic AI in Insurance Decision-Making: Opportunities and Challenges For Africa
- In-Context Iterative Policy Improvement for Dynamic Manipulation
- Comparing energy consumption and accuracy in text classification inference
- A Word on Machine Ethics: A Response to Jiang et al. (2021)
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- The Cultural Gene of Large Language Models: A Study on the Impact of Cross-Corpus Training on Model Values and Biases
- A Multi-Task Evaluation of LLMs' Processing of Academic Text Input
- Tähenduste seletamine leksikograafias: kuivõrd on abi suurtest keelemudelitest?
- Becoming, Doing, Being: GenAI and the Promise of Professional Identity in Law
- Detecting Hate Speech with GPT-3
- Why Report Failed Interactions With Robots?! Towards Vignette-based Interaction Quality
- Diversity First, Quality Later: A Two-Stage Assumption for Language Model Alignment
- Computational Economics in Large Language Models: Exploring Model Behavior and Incentive Design under Resource Constraints
- Yet another algorithmic bias: A Discursive Analysis of Large Language Models Reinforcing Dominant Discourses on Gender and Race
- Wisdom of the Crowd, Without the Crowd: A Socratic LLM for Asynchronous Deliberation on Perspectivist Data
- BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation
- Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization
- Activation Steering for Bias Mitigation: An Interpretable Approach to Safer LLMs
- Generative artificial intelligence as an enabler of student feedback engagement: a framework
- Augmenting Bias Detection in LLMs Using Topological Data Analysis
- Progressive Depth Up-scaling via Optimal Transport
- Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution
- The Epistemic Power of Human-Machine Communication
- The NordDRG AI Benchmark for Large Language Models
- LLMs, Turing tests and Chinese rooms: the prospects for meaning in large language models
- The Problem of Atypicality in LLM-Powered Psychiatry
- Intuition emerges in Maximum Caliber models at criticality
- Navigating the Risks of Using Large Language Models for Text Annotation in Social Science Research
- Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
- Situated Epistemic Infrastructures: A Diagnostic Framework for Post-Coherence Knowledge
- I Think, Therefore I Am Under-Qualified? A Benchmark for Evaluating Linguistic Shibboleth Detection in LLM Hiring Evaluations
- Why are LLMs' abilities emergent?
- Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models
- Generative AI in Training and Coaching: Redefining the Design Process of Learning Materials
- Ethics through the Facets of Artificial Intelligence
- Analyzing Prominent LLMs: An Empirical Study of Performance and Complexity in Solving LeetCode Problems
- Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models
- LaTCoder: Converting Webpage Design to Code with Layout-as-Thought
- Investigating Gender Bias in LLM-Generated Stories via Psychological Stereotypes
- Machine culture
- Can LLMs Generate High-Quality Task-Specific Conversations?
- Dynaword: From One-shot to Continuously Developed Datasets
- TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
- A Confidence-Diversity Framework for Calibrating AI Judgement in Accessible Qualitative Coding Tasks
- A Survey on Data Security in Large Language Models
- Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback
- Quantum-RAG and PunGPT2: Advancing Low-Resource Language Generation and Retrieval for the Punjabi Language
- Polymorphic Combinatorial Frameworks (PCF): Guiding the Design of Mathematically-Grounded, Adaptive AI Agents
- Applying Psychometrics to Large Language Model Simulated Populations: Recreating the HEXACO Personality Inventory Experiment with Generative Agents
- Measuring and Analyzing Intelligence via Contextual Uncertainty in Large Language Models using Information-Theoretic Metrics
- Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English
- Lessons from complex systems science for AI governance
- Multilingual Political Views of Large Language Models: Identification and Steering
- Why can't Epidemiology be automated (yet)?
- Rote Learning Considered Useful: Generalizing over Memorized Data in LLMs
- Collaborative Storytelling with Human Actors and AI Narrators
- Modern Uyghur Dependency Treebank (MUDT): An Integrated Morphosyntactic Framework for a Low-Resource Language
- What Does it Mean for a Neural Network to Learn a "World Model"?
- Do Large Language Models Understand Morality Across Cultures?
- Coupled Gradient Estimators for Discrete Latent Variables
- The Xeno Sutra: Can Meaning and Value be Ascribed to an AI-Generated "Sacred" Text?
- The Carbon Cost of Conversation, Sustainability in the Age of Language Models
- The Algorithmic Imprint
- Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes
- Hate Speech Classifiers Learn Human-Like Social Stereotypes
- Small-Bench NLP: Benchmark for small single GPU trained models in Natural Language Processing
- Evaluating Large Language Models (LLMs) in Financial NLP: A Comparative Study on Financial Report Analysis
- Multiple text comprehension in contemporary contexts: Integrating broader understandings of affect, culture, and technology
- Why large language models are poor theories of human linguistic cognition: A reply to Piantadosi
- Generative AI in Qualitative Research and Related Transparency Problems: A Novel Heuristic for Disclosing Uses of AI
- Machine-assisted quantitizing designs: augmenting humanities and social sciences with artificial intelligence
- Friend or foe? Exploring the implications of large language models on the science system
- The Copernican Argument for Alien Consciousness; The Mimicry Argument Against Robot Consciousness
- Automating civilian harm: On Israel's use of the AI-enabled targeting system Lavender in Gaza and International Humanitarian Law
- The Promise of Foundational Large Language Models in Analysis and Interpretation of Wearable Data: Implications for Physical Behavior Research
- The AIR framework for research transparency: a critical analysis of stage-specific AI disclosure in the context of accessibility and research integrity
- Mapping cosmopolitan nationalism in educational research: a literature review
- In relation, not replacement: a reflexive case for generative AI in qualitative research
- Rethinking Memorization Measures and their Implications in Large Language Models
- Towards semantic versioning of open pre-trained language model releases on hugging face
- Re‐Imagining the Epistemic Possibilities of GPT for Public Administration Research in Competitive Settings
- Hippocampo-neocortical interaction as compressive retrieval-augmented generation
- Principals editors científics als canals de Telegram : una aproximació a la detecció de canals falsos amb ChatGPT i DeepSeek
- Beyond Fairness Metrics: Roadblocks and Challenges for Ethical AI in Practice
- A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI
- Socio-Technological Challenges and Opportunities: Paths Forward
- Strategic Polysemy in AI Discourse: A Philosophical Analysis of Language, Hype, and Power
- Inferring Offensiveness In Images From Natural Language Supervision
- Reinforcing intensive motherhood: A study of gender bias in parental responsibilities allocation by large language models
- The Geometry of Harmfulness in LLMs through Subconcept Probing
- Agent Identity Evals: Measuring Agentic Identity
- How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMs
- How generative AI fixes what higher education broke: reframing the “colonization” of knowing and learning
- Assessing the Reliability of LLMs Annotations in the Context of Demographic Bias and Model Explanation
- GPTFootprint: Increasing Consumer Awareness of the Environmental Impacts of LLMs
- Tackling Climate Change with Machine Learning
- REVA: Supporting LLM-Generated Programming Feedback Validation at Scale Through User Attention-based Adaptation
- Towards Creating Infrastructures for Values and Ethics Work in the Production of Software Technologies
- A study of search result aggregation approaches for the digital humanities
- Assembling platform governance as private ordering in the age of generative AI: platform interdependence in policy evolution
- AI chatbot accountability in the age of algorithmic gatekeeping: Comparing generative search engine political information retrieval across five languages
- The value of books in the age of generative AI training data
- Owning the Simulacrum: Commodification, Intellectual Property, and the Collapse of Epistemic Sovereignty
- The Power of Absence: Thinking with Archival Theory in Algorithmic Design
- A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search
- Societal Biases in Language Generation: Progress and Challenges
- Trustworthy AI for Medicine: Continuous Hallucination Detection and Elimination with CHECK
- Guiding LLM Decision-Making with Fairness Reward Models
- EsBBQ and CaBBQ: The Spanish and Catalan Bias Benchmarks for Question Answering
- Patterns, Models, and Challenges in Online Social Media: A Survey
- PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
- Cultural Bias in Large Language Models: Evaluating AI Agents through Moral Questionnaires
- Comprehension Without Competence: Architectural Limits of LLMs in Symbolic Computation and Reasoning
- An Analysis of Chinese Censorship Bias in LLMs
- LegaLMFiT: Efficient Short Legal Text Classification with LSTM Language Model Pre-Training
- Can Large Language Models Understand As Well As Apply Patent Regulations to Pass a Hands-On Patent Attorney Test?
- Generative AI in Science: Applications, Challenges, and Emerging Questions
- When Large Language Models Meet Law: Dual-Lens Taxonomy, Technical Advances, and Ethical Governance
- AI Should Sense Better, Not Just Scale Bigger: Adaptive Sensing as a Paradigm Shift
- GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing
- Strategyproof Learning: Building Trustworthy User-Generated Datasets
- A Mathematical Theory of Discursive Networks
- Beyond Scale: Small Language Models are Comparable to GPT-4 in Mental Health Understanding
- The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
- FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation
- Large Language Models Predict Human Well-being -- But Not Equally Everywhere
- The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models
- Mechanistic Indicators of Understanding in Large Language Models
- Assessing the Ecological Impact of AI
- On the Semantics of Large Language Models
- Lilith: Developmental Modular LLMs with Chemical Signaling
- Fairness Evaluation of Large Language Models in Academic Library Reference Services
- A Technical Survey of Reinforcement Learning Techniques for Large Language Models
- Losing our Tail, Again: (Un)Natural Selection & Multilingual LLMs
- Roadmap for using large language models (LLMs) to accelerate cross-disciplinary research with an example from computational biology
- Four Shades of Life Sciences: A Dataset for Disinformation Detection in the Life Sciences
- Emergent Inabilities? Inverse Scaling Over the Course of Pretraining
- Considering the ethics of large machine learning models in the chemical sciences
- ELLMA-T: an Embodied LLM-agent for Supporting English Language Learning in Social VR
- From Measurement to Mitigation: Exploring the Transferability of Debiasing Approaches to Gender Bias in Maltese Language Models
- Hi, my name is Martha: Using names to measure and mitigate bias in generative dialogue models
- Can Artificial Intelligence solve the blockchain oracle problem? Unpacking the Challenges and Possibilities
- Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions
- Stereotype Detection as a Catalyst for Enhanced Bias Detection: A Multi-Task Learning Approach
- Epistemic Scarcity: The Economics of Unresolvable Unknowns
- Low-Perplexity LLM-Generated Sequences and Where To Find Them
- Discourse Heuristics For Paradoxically Moral Self-Correction
- The Generalist Brain Module: Module Repetition in Neural Networks in Light of the Minicolumn Hypothesis
- Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models
- Geography for AI sustainability and sustainability for GeoAI
- “Desired behaviors”: alignment and the emergence of a machine learning ethics
- On Recipe Memorization and Creativity in Large Language Models: Is Your Model a Creative Cook, a Bad Cook, or Merely a Plagiator?
- A Distributional Approach to Controlled Text Generation
- Positioning AI Tools to Support Online Harm Reduction Practice: Applications and Design Directions
- BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute
- Against 'softmaxing' culture
- Text Production and Comprehension by Human and Artificial Intelligence: Interdisciplinary Workshop Report
- Temperature Matters: Enhancing Watermark Robustness Against Paraphrasing Attacks
- Bias, Accuracy, and Trust: Gender-Diverse Perspectives on Large Language Models
- A Dual-Layered Evaluation of Geopolitical and Cultural Bias in LLMs
- PromptAug: Fine-grained Conflict Classification Using Data Augmentation
- Challenges in Deploying Machine Learning: a Survey of Case Studies
- Fanfiction in the Age of AI: Community Perspectives on Creativity, Authenticity and Adoption
- Deciphering Emotions in Children Storybooks: A Comparative Analysis of Multimodal LLMs in Educational Applications
- HIDE and Seek: Detecting Hallucinations in Language Models via Decoupled Representations
- Critical Generative AI Literacy for Social Studies Educators: A Typology of GenAI Errors and Their Impacts on Epistemology
- Whither Chat-GPT: Generative AI and Teaching Introduction to Human Geography in Higher Education
- SysTemp: A Multi-Agent System for Template-Based Generation of SysML v2
- Language Bottleneck Models for Qualitative Knowledge State Modeling
- Semantic Outlier Removal with Embedding Models and LLMs
- Alignment between Brains and AI: Evidence for Convergent Evolution across Modalities, Scales and Training Trajectories
- Mapping Caregiver Needs to AI Chatbot Design: Strengths and Gaps in Mental Health Support for Alzheimer's and Dementia Caregivers
- Finance Language Model Evaluation (FLaME)
- SANSKRITI: A Comprehensive Benchmark for Evaluating Language Models' Knowledge of Indian Culture
- Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor
- From General Reasoning to Domain Expertise: Uncovering the Limits of Generalization in Large Language Models
- The poverty of ethical AI: impact sourcing and AI supply chains
- Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
- Attribution-Guided Pruning for Insight and Control: Circuit Discovery and Targeted Correction in Small-scale LLMs
- Rectifying Privacy and Efficacy Measurements in Machine Unlearning: A New Inference Attack Perspective
- Deflating Deflationism: A Critical Perspective on Debunking Arguments Against LLM Mentality
- On the Expressive Power of Self-Attention Matrices
- Designing GenAI Tools for Personalized Learning Implementation: Theoretical Analysis and Prototype of a Multi-Agent System
- Social-Emotional Learning and Generative AI: A Critical Literature Review and Framework for Teacher Education
- Sense and Sensibility: What makes a social robot convincing to high-school students?
- Exploring Cultural Variations in Moral Judgments with Large Language Models
- Advances in LLMs with Focus on Reasoning, Adaptability, Efficiency and Ethics
- Conversational AI as a Catalyst for Informal Learning: An Empirical Large-Scale Study on LLM Use in Everyday Learning
- Large Language Models for History, Philosophy, and Sociology of Science: Interpretive Uses, Methodological Challenges, and Critical Perspectives
- LoRA Users Beware: A Few Spurious Tokens Can Manipulate Your Finetuned Model
- How large language models can reshape collective intelligence
- The voice of artificial intelligence: Philosophical and educational reflections
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- Social Scientists on the Role of AI in Research
- Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
- Not quite Sherlock Holmes: Language model predictions do not reliably differentiate impossible from improbable events
- Epistemic Artificial Intelligence is Essential for Machine Learning Models to Truly 'Know When They Do Not Know'
- Evaluating Prompt-Driven Chinese Large Language Models: The Influence of Persona Assignment on Stereotypes and Safeguards
- Truly Self-Improving Agents Require Intrinsic Metacognitive Learning
- Energentic Intelligence: From Self-Sustaining Systems to Enduring Artificial Life
- LLM-First Search: Self-Guided Exploration of the Solution Space
- Future directions for chatbot research: an interdisciplinary research agenda
- Fine-Grained Interpretation of Political Opinions in Large Language Models
- Metaheuristics and Large Language Models Join Forces: Toward an Integrated Optimization Approach
- Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
- A Tale of Two Identities: An Ethical Audit of Human and AI-Crafted Personas
- CLIPort: What and Where Pathways for Robotic Manipulation
- Measurement as governance in and for responsible AI
- Large Means Left: Political Bias in Large Language Models Increases with Their Number of Parameters
- Is the end of Insight in Sight ?
- The promise and limits of LLMs in constructing proofs and hints for logic problems in intelligent tutoring systems
- Misalignment or misuse? The AGI alignment tradeoff
- ClozeMath: Improving Mathematical Reasoning in Language Models by Learning to Fill Equations
- Linear Spatial World Models Emerge in Large Language Models
- Intersectional Bias in Causal Language Models
- CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge
- A Multi-Agent Framework for Mitigating Dialect Biases in Privacy Policy Question-Answering Systems
- The Future of Continual Learning in the Era of Foundation Models: Three Key Directions
- NetArena: Dynamic Benchmarks for AI Agents in Network Automation
- An Empirical Study of Group Conformity in Multi-Agent Systems
- Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean
- Benford's Curse: Tracing Digit Bias to Numerical Hallucination in LLMs
- Fodor and Pylyshyn's Legacy: Still No Human-like Systematic Compositionality in Neural Networks
- Integrating Neural and Symbolic Components in a Model of Pragmatic Question-Answering
- HADA: Human-AI Agent Decision Alignment Architecture
- Probing the Geometry of Truth: Consistency and Generalization of Truth Directions in LLMs Across Logical Transformations and Question Answering Tasks
- Translate With Care: Addressing Gender Bias, Neutrality, and Reasoning in Large Language Model Translations
- Recover Experimental Data with Selection Bias using Counterfactual Logic
- Position: Olfaction Standardization is Essential for the Advancement of Embodied Artificial Intelligence
- The Impact of Large Language Models on K-12 Education in Rural India: A Thematic Analysis of Student Volunteer's Perspectives
- Soft Best-of-n Sampling for Model Alignment
- The World As Large Language Models See It: Exploring the reliability of LLMs in representing geographical features
- Multiple LLM Agents Debate for Equitable Cultural Alignment
- KI-Tools für die wissenschaftliche Literaturrecherche: Potenziale, Problematiken, Didaktik und Zukunftsperspektiven
- MedArabiQ: Benchmarking Large Language Models on Arabic Medical Tasks
- Neither Stochastic Parroting nor AGI: LLMs Solve Tasks through Context-Directed Extrapolation from Training Data Priors
- A Mathematical Framework for AI-Human Integration in Work
- Talent or Luck? Evaluating Attribution Bias in Large Language Models
- In Dialogue with Intelligence: Rethinking Large Language Models as Collective Knowledge
- Pink for Princesses, Blue for Superheroes: The Need to Examine Gender Stereotypes in Kid's Products in Search and Recommendations
- BiasFilter: An Inference-Time Debiasing Framework for Large Language Models
- From prosthetic memory to prosthetic denial: Auditing whether large language models are prone to mass atrocity denialism
- MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs
- Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)
- Representative Language Generation
- Describe Me Something You Do Not Remember - Challenges and Risks of Exposure Design Using Generative Artificial Intelligence for Therapy of Complex Post-traumatic Disorder
- Explaining Large Language Models with gSMILE
- R3-RAG: Learning Step-by-Step Reasoning and Retrieval for LLMs via Reinforcement Learning
- Large Language Models in Code Co-generation for Safe Autonomous Vehicles
- Fairness-in-the-Workflow: How Machine Learning Practitioners at Big Tech Companies Approach Fairness in Recommender Systems
- The Power of Iterative Filtering for Supervised Learning with (Heavy) Contamination
- Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets
- CoTGuard: Using Chain-of-Thought Triggering for Copyright Protection in Multi-Agent LLM Systems
- A Survey on Progress in LLM Alignment from the Perspective of Reward Design
- Moderating Harm: Benchmarking Large Language Models for Cyberbullying Detection in YouTube Comments
- MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent Systems
- Do Large Language Models (Really) Need Statistical Foundations?
- Ethical behavior in humans and machines -- Evaluating training data quality for beneficial machine learning
- Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects
- Social Good or Scientific Curiosity? Uncovering the Research Framing Behind NLP Artefacts
- Evaluating Intra-firm LLM Alignment Strategies in Business Contexts
- OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models
- Diversity and Inclusion in AI: Insights from a Survey of AI/ML Practitioners
- Language Model Behavior: A Comprehensive Survey
- Towards Anonymous Neural Network Inference
- The Staircase of Ethics: Probing LLM Value Priorities through Multi-Step Induction to Complex Moral Dilemmas
- Does Chain-of-Thought Reasoning Really Reduce Harmfulness from Jailbreaking?
- Surfacing Semantic Orthogonality Across Model Safety Benchmarks: A Multi-Dimensional Analysis
- Meaning Is Not A Metric: Using LLMs to make cultural context legible at scale
- Is It Bad to Work All the Time? Cross-Cultural Evaluation of Social Norm Biases in GPT-4
- The Case for Repeatable, Open, and Expert-Grounded Hallucination Benchmarks in Large Language Models
- Mitigate One, Skew Another? Tackling Intersectional Biases in Text-to-Image Models
- Rethinking Generative AI Literacy: An Integrative, Developmental, and Dialectical Framework for K-12 Teacher Education
- Only Large Weights (And Not Skip Connections) Can Prevent the Perils of Rank Collapse
- "AI just keeps guessing": Using ARC Puzzles to Help Children Identify Reasoning Errors in Generative AI
- Adaptive Plan-Execute Framework for Smart Contract Security Auditing
- Insiders and Outsiders in Research on Machine Learning and Society
- Temperature-driven inversion and nonlinear dynamics in ChatGPT-like AIs
- Attributional Safety Failures in Large Language Models under Code-Mixed Perturbations
- DECASTE: Unveiling Caste Stereotypes in Large Language Models through Multi-Dimensional Bias Analysis
- Chain-of-Thought Driven Adversarial Scenario Extrapolation for Robust Language Models
- Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
- Language and Thought: The View from LLMs
- The Hidden Structure -- Improving Legal Document Understanding Through Explicit Text Formatting
- Do people rely on ChatGPT more than their peers to detect deepfake news?
- How should the advancement of large language models affect the practice of science?
- Risks and Benefits of Large Language Models for the Environment
- Striking the (im)balance – a review of the relative prevalence of meta-ethical models in AI journalism research
- How to Turn Ethical Values Into System Requirements: Lessons Learned from Adopting a New IEEE Standard in the Business World
- The Rise of Artificial Intelligence in Educational Measurement: Opportunities and Ethical Challenges
- Harm to Nonhuman Animals from AI: a Systematic Account and Framework
- Wisdom from Diversity: Bias Mitigation Through Hybrid Human-LLM Crowds
- AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference
- The Effects of Demographic Instructions on LLM Personas
- Fast RoPE Attention: Combining the Polynomial Method and Fast Fourier Transform
- Class Distillation with Mahalanobis Contrast: An Efficient Training Paradigm for Pragmatic Language Understanding Tasks
- The Epistemic Politics of AI Anthropomorphism
- EcoSafeRAG: Efficient Security through Context Analysis in Retrieval-Augmented Generation
- Speciesism in natural language processing research
- Care and Scale: Decorrelative Ethics in Algorithmic Recommendation
- Commemorar, especular: algunes tendències en la poesia editada a les Illes Balears
- “Just a pocket knife, not a machete”: Large language models in TEFL teacher education & digital text sovereignty
- Where did the ambiguity go? Examining how multimodal models interpret polysemous words
- Artificial intelligence and illusions of understanding in scientific research
- AI Empire: Unraveling the interlocking systems of oppression in generative AI's global order
- Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems
- Embodying the Algorithm
- Context parroting: A simple but tough-to-beat baseline for foundation models in scientific machine learning
- Can Large Language Models Correctly Interpret Equations with Errors?
- "There Is No Such Thing as a Dumb Question," But There Are Good Ones
- Interpretable Risk Mitigation in LLM Agent Systems
- Comparing LLM Text Annotation Skills: A Study on Human Rights Violations in Social Media Data
- AI-enhanced semantic feature norms for 786 concepts
- Evaluating GPT- and Reasoning-based Large Language Models on Physics Olympiad Problems: Surpassing Human Performance and Implications for Educational Assessment
- Autofocus Retrieval: An Effective Pipeline for Multi-Hop Question Answering With Semi-Structured Knowledge
- Ethics of Socially Disruptive Technologies
- TESCREAL hallucinations: Psychedelic and AI hype as inequality engines
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishing
- How to cheat on your final paper: Assigning AI for student writing
- We went to look for meaning and all we got were these lousy representations: aspects of meaning representation for computational semantics
- Why we need biased AI -- How including cognitive and ethical machine biases can enhance AI systems
- A Social Robot with Inner Speech for Dietary Guidance
- Efficient Unstructured Pruning of Mamba State-Space Models for Resource-Constrained Environments
- For GPT-4 as with Humans: Information Structure Predicts Acceptability of Long-Distance Dependencies
- Detecting Prefix Bias in LLM-based Reward Models
- Multi-Agent Ethnography: Post-Conventional Anthropological Practice Through Human−AI Collaboration
- A Turing Test for ''Localness'': Conceptualizing, Defining, and Recognizing Localness in People and Machines
- Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
- AI in Money Matters
- Sandcastles in the Storm: Revisiting the (Im)possibility of Strong Watermarking
- Engineering Risk-Aware, Security-by-Design Frameworks for Assurance of Large-Scale Autonomous AI Models
- The fragility of AI companionship: Ontological, structural, and normative uncertainty in human-AI relationships
- Toward Reasonable Parrots: Why Large Language Models Should Argue with Us by Design
- Interaction Configurations and Prompt Guidance in Conversational AI for Question Answering in Human-AI Teams
- One Search Fits All: Pareto-Optimal Eco-Friendly Model Selection
- Don't be lazy: CompleteP enables compute-efficient deep transformers
- Correct codes for the wrong reasons? validating LLMs as measurement instruments for theoretical constructs
- Gender Disambiguation in Machine Translation: Diagnostic Evaluation in Decoder-Only Architectures
- Behavioral Indicators of Overreliance During Interaction with Conversational Language Models
- GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models
- On the missing data layer and a potential solution
- Tournesol: A quest for a large, secure and trustworthy database of reliable human judgments
- Auditing Preferences for Brands and Cultures in LLMs
- AI Virtue: What is "Good" Knowledge in the Age of Artificial Intelligence?
- Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering
- The Neuroscience of Transformers
- LLM BiasScope: A Real-Time Bias Analysis Platform for Comparative LLM Evaluation
- No-Free-Fairness: Fundamental Limits and Trade-offs in Learning Systems
- When does pretraining help?
- Flaws in the LLM Automation Narrative
- The Rising Unsustainability of AI Graphics Cards Production
- The limits of large language models and the necessity of human cognition in K-12 education
- Towards Efficient and Explainable Hate Speech Detection via Model Distillation
- Structural Hallucination in Large Language Models: A Network-Based Evaluation of Knowledge Organization and Citation Integrity
- LLM Ethics Benchmark: A Three-Dimensional Assessment System for Evaluating Moral Reasoning in Large Language Models
- The myth of meaning: generative AI as language-endowed machines and the machinic essence of the human being
- Spilled Energy in Large Language Models
- Arithmetic Pedagogy for Language Models
- The Statistical Signature of LLMs
- Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications
- From Reflection to Repair: A Scoping Review of Dataset Documentation Tools
- Chatbots Output Meaningful (but Problematic) Language
- Causal methods for LLM development and evaluation
- Who Gets to Do Physics? Occupational Stereotypes in AI-Generated Problem Sets
- TabletCraft: Bridging a 4,000-Year Cultural Gap with Bidirectional Akkadian NMT and Cuneiform Rendering
- Fair outputs, Biased Internals: Causal Potency and Asymmetry of Latent Bias in LLMs for High-Stakes Decisions
- Low-Cost Black-Box Detection of LLM Hallucinations via Dynamical System Prediction
- Beyond Detection: Governing GenAI in Academic Peer Review as a Sociotechnical Challenge
- Formal and Empirical Study of Metadata-Based Profiling for Resource Management in the Computing Continuum
- Information Retrieval in the Age of Generative AI: The RGB Model
- SAGE: A Generic Framework for LLM Safety Evaluation
- LLM-Assisted Empirical Software Engineering: Systematic Literature Review and Research Agenda
- Exposing the Unsaid: Visualizing Hidden LLM Bias through Stochastic Path Aggregation
- Locating acts of mechanistic reasoning in student team conversations with mechanistic machine learning
- Perspective on Bias in Biomedical AI: Preventing Downstream Healthcare Disparities
- LLM Consumer Behavior Theory: Foundations of a Novel Research Field
- Computational Hermeneutics: Evaluating generative AI as a cultural technology
- Eroding the Truth-Default: A Causal Analysis of Human Susceptibility to Foundation Model Hallucinations and Disinformation in the Wild
- Why AI Alignment Failure Is Structural: Learned Human Interaction Structures and AGI as an Endogenous Evolutionary Shock
- Regulatory gray areas of LLM Terms
- Keyword search is all you need: Achieving RAG-Level Performance without vector databases using agentic tool use
- Towards Automated Scoping of AI for Social Good Projects
- Can a Large Language Model Assess Urban Design Quality? Evaluating Walkability Metrics Across Expertise Levels
- A Platform for Generating Educational Activities to Teach English as a Second Language
- The State of AI Governance Research: AI Safety and Reliability in Real World Commercial Deployment
- Generative AI Literacy: A Comprehensive Framework for Literacy and Responsible Use
- Where Knowledge Collides: A Mechanistic Study of Intra-Memory Knowledge Conflict in Language Models
- Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text
- Mapping the City Through the Lens of Language Models
- "Are we writing an advice column for Spock here?" Understanding Stereotypes in AI Advice for Autistic Users
- Mapping the Stochastic Penal Colony
- Performance in a dialectal profiling task of LLMs for varieties of Brazilian Portuguese
- Unequal Verdicts: Investigating Gender Bias in LLM-Based Fake News Detection
- On the missing benchmarks layer and a potential solution
- Cultural Encoding in Large Language Models: The Existence Gap in AI-Mediated Brand Discovery
- Evolution of AI in Education: Agentic Workflows
- Evaluating Machine Expertise: How Graduate Students Develop Frameworks for Assessing GenAI Content
- Auditing the Ethical Logic of Generative AI Models
- Generating Synthetic Text Data to Evaluate Causal Inference Methods
- Unpacking the Expressed Consequences of AI Research in Broader Impact Statements
- What We Do Not Know: GPT Use in Business and Management
- Cognitive Silicon: An Architectural Blueprint for Post-Industrial Computing Systems
- LLMCode: Evaluating and Enhancing Researcher-AI Alignment in Qualitative Analysis
- Data and its (dis)contents: A survey of dataset development and use in machine learning research
- Sourcing behavior and the role of news media in AI-powered search engines in the digital media ecosystem: Comparing political news retrieval across five languages
- An Empirical Comparison of Text Summarization: A Multi-Dimensional Evaluation of Large Language Models
- Balancing Complexity and Informativeness in LLM-Based Clustering: Finding the Goldilocks Zone
- The implications of generative artificial intelligence for mathematics education
- Position: Bayesian Statistics Facilitates Stakeholder Participation in Evaluation of Generative AI
- Combating Toxic Language: A Review of LLM-Based Strategies for Software Engineering
- Aspirational Affordances of AI
- Contemplative Agent
- Knowledge Distillation and Dataset Distillation of Large Language Models: Emerging Trends, Challenges, and Future Directions
- Reliable Confidence Intervals for Information Retrieval Evaluation Using Generative A.I.
- Biased by Design: Leveraging AI Biases to Enhance Critical Thinking of News Readers
- Mind the Language Gap: Automated and Augmented Evaluation of Bias in LLMs for High- and Low-Resource Languages
- Beyond Misinformation: A Conceptual Framework for Studying AI Hallucinations in (Science) Communication
- Data-Efficient Pretraining via Contrastive Self-Supervision
- DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification
- Large language model-supported companion robots for loneliness in older people: A UK–Japan qualitative study integrating focus groups and in-home deployment
- Generate, but Verify: Reducing Hallucination in Vision-Language Models with Retrospective Resampling
- Limitations of Autoregressive Models and Their Alternatives
- Is Trust Correlated With Explainability in AI? A Meta-Analysis
- AI Safety Should Prioritize the Future of Work
- Large Language Models as Quasi-crystals: Coherence Without Repetition in Generative Text
- From job titles to jawlines: Using context voids to study generative AI systems
- Bridging the Semantic Gaps: Improving Medical VQA Consistency with LLM-Augmented Question Sets
- Mitigating LLM Hallucinations with Knowledge Graphs: A Case Study
- Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift
- Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic Data
- Bias Beyond English: Evaluating Social Bias and Debiasing Methods in a Low-Resource Setting
- DataPuzzle: Breaking Free from the Hallucinated Promise of LLMs in Data Analysis
- Weight-of-Thought Reasoning: Exploring Neural Network Weights for Enhanced LLM Reasoning
- The Human Visual System Can Inspire New Interaction Paradigms for LLMs
- The Code Barrier: What LLMs Actually Understand?
- Managing the Twin Faces of AI: A Commentary on “Is AI Changing the World for Better or Worse?”
- An Evaluation of Cultural Value Alignment in LLM
- Research as Resistance: Recognizing and Reconsidering HCI's Role in Technology Hype Cycles
- Knowledge Graph-extended Retrieval Augmented Generation for Question Answering
- "i am a stochastic parrot, and so r u": Is AI-based framing of human behaviour and cognition a conceptual metaphor or conceptual engineering?
- A taxonomy of epistemic injustice in the context of AI and the case for generative hermeneutical erasure
- Tacit knowledge in large language models
- Linguistic Interpretability of Transformer-based Language Models: a systematic review
- Navigating the Rabbit Hole: Emergent Biases in LLM-Generated Attack Narratives Targeting Mental Health Groups
- RETROcode: Leveraging a Code Database for Improved Natural Language to Code Generation
- From PMI to Bots
- Connecting Feedback to Choice: Understanding Educator Preferences in GenAI vs. Human-Created Lesson Plans in K-12 Education -- A Comparative Analysis
- Digital rhetoric [wikipedia]
- GPT-3 [wikipedia]
- A Benchmark for Data Imputation Methods. [europepmc]
- Operationalizing and Implementing Pretrained, Large Artificial Intelligence Linguistic Models in the US Health Care System: Outlook of Generative Pretrained Transformer 3 (GPT-3) as a Service Model. [europepmc]
- AI recognition of patient race in medical imaging: a modelling study. [europepmc]
- Implicit-Bias Remedies: Treating Discriminatory Bias as a Public-Health Problem. [europepmc]
- Exclusion of the non-English-speaking world from the scientific literature: Recommendations for change for addiction journals and publishers. [europepmc]
- Artificial Intelligence in Emergency Medicine: Viewpoint of Current Applications and Foreseeable Opportunities and Challenges. [europepmc]
- ChatGPT's inconsistent moral advice influences users' judgment. [europepmc]
- ChatGPT goes to the operating room: evaluating GPT-4 performance and its potential in surgical education and training in the era of large language models. [europepmc]
- A Medical Ethics Framework for Conversational Artificial Intelligence. [europepmc]
- Putting ChatGPT's Medical Advice to the (Turing) Test: Survey Study. [europepmc]
- Assessing the Utility of ChatGPT Throughout the Entire Clinical Workflow: Development and Usability Study. [europepmc]
- Large language models propagate race-based medicine. [europepmc]
- Black Box Warning: Large Language Models and the Future of Infectious Diseases Consultation. [europepmc]
- The promises of large language models for protein design and modeling. [europepmc]
- Ethical Considerations of Artificial Intelligence in Health Care: Examining the Role of Generative Pretrained Transformer-4. [europepmc]
- Understanding the Benefits and Challenges of Using Large Language Model-based Conversational Agents for Mental Well-being Support. [europepmc]
- Large Language Models: A Guide for Radiologists. [europepmc]
- Academic publisher guidelines on AI usage: A ChatGPT supported thematic analysis. [europepmc]
- Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES): a method for populating knowledge bases using zero-shot learning. [europepmc]
- Large language models could change the future of behavioral healthcare: a proposal for responsible development and evaluation. [europepmc]
- Applied Artificial Intelligence in Healthcare: A Review of Computer Vision Technology Application in Hospital Settings. [europepmc]
- The Artificial Third: A Broad View of the Effects of Introducing Generative Artificial Intelligence on Psychotherapy. [europepmc]
- The Role of Humanization and Robustness of Large Language Models in Conversational Artificial Intelligence for Individuals With Depression: A Critical Analysis. [europepmc]
- Why we need to be careful with LLMs in medicine. [europepmc]
- A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. [europepmc]
Discussions
- Ça colle parfaitement à ce que disent des critiques des LLM : ces modèles sont faits pour générer du texte plausible, ni vrai ni pensé. Ils manipulent la forme du langage sans accès au sens comme acti [bsky, 15 points, 3 comments]
- Inte bara i podden. Det är ett papper hon skrev ihop med Timnit Gebru och några andra, som fått stort genomslag: doi.org/10.1145/3442... [bsky, 2 points, 0 comments]
- The model still took a of GPU time, energy/water, and millions to train…and those data used. Not saying some models can’t be useful, but let’s not lose the plot of the problematic nature of these mode [bsky, 0 points, 0 comments]
- doi.org/10.1145/3442... [bsky, 0 points, 1 comments]
- Ups, no sabía que había que traer referencias. 😅 Loros estocasticos, me declaro fan, ahí argumentan que el significado esta en la mirada del que observa puesto que los textos sé generar solo en funci [bsky, 0 points, 1 comments]
- In other words, a stochastic parrot: a system for haphazardly stitching together sequences of linguistic forms it has observed in its vast training data, according to probabilistic information about h [bsky, 0 points, 1 comments]
Related