Sparks of Artificial General Intelligence: Early experiments with GPT-4
2023/03/22 by Sébastien Bubeck, Bubeck, Sébastien, Varun Chandrasekaran +26 · 21 voices · 1,564 citations
Computer Science · Medicine · Psychology · #Artificial Intelligence in Healthcare and Education #Artificial general intelligence #Artificial intelligence #Coding (social sciences) #Cognition #Cognitive psychology #Cognitive science #Computer science #Machine Learning in Healthcare #Machine learning #Predictive coding #Psychology #Social science #Sociology #Topic Modeling #Variety (cybernetics)
paper · pdf · doi:10.48550/arxiv.2303.12712
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/03/22 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Abstract
Artificial intelligence (AI) researchers have been developing and refining large language models (LLMs) that exhibit remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition. The latest model developed by OpenAI, GPT-4, was trained using an unprecedented scale of compute and data. In this paper, we report on our investigation of an early version of GPT-4, when it was still in active development by OpenAI. We contend that (this early version of) GPT-4 is part of a new cohort of LLMs (along with ChatGPT and Google's PaLM for example) that exhibit more general intelligence than previous AI models. We discuss the rising capabilities and implications of these models. We demonstrate that, beyond its mastery of language, GPT-4 can solve novel and difficult tasks that span mathematics, coding, vision, medicine, law, psychology and more, without needing any special prompting. Moreover, in all of these tasks, GPT-4's performance is strikingly close to human-level performance, and often vastly surpasses prior models such as ChatGPT. Given the breadth and depth of GPT-4's capabilities, we believe that it could reasonably be viewed as an early (yet still incomplete) version of an artificial general intelligence (AGI) system. In our exploration of GPT-4, we put special emphasis on discovering its limitations, and we discuss the challenges ahead for advancing towards deeper and more comprehensive versions of AGI, including the possible need for pursuing a new paradigm that moves beyond next-word prediction. We conclude with reflections on societal influences of the recent technological leap and future research directions.
Cited by
- Efficient Online LLM Watermark Detection via Rao-Blackwellized E-Processes
- UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs
- A Unified Moral-Value Dataset for Instruction Tuning
- From Assistance to Autonomy -- A Researcher Study on the Potential of AI Support for Qualitative Data Analysis
- Breaking the Block: Preserving Data Continuity to Train Superior SAEs for Instruct Models
- GhostShell: Streaming LLM Function Calls for Concurrent Embodied Programming
- HEPTAPOD: Orchestrating High Energy Physics Workflows Towards Autonomous Agency
- Cost-efficient generative AI summarization for scalable automated essay scoring in educational assessment
- When Models Meet Users: An Empirical Study of Perceptions of General LLMs and Multimodal LLMs on Hugging Face
- The Riddle Riddle: Testing Flexible Reasoning in Large Language Models and Humans
- Faster Completion, Less Learning: Generative AI Reduced Study Time on Math Problems and the Knowledge They Build
- The Storyteller in the Model: Narrative Pattern Inheritance, Escalation Dynamics, and Alignment Governance in LLMs
- Shared sensitivity to data distribution during learning in humans and transformer networks
- What does it mean to understand language?
- Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
- Large Language Models Report Subjective Experience Under Self-Referential Processing
- Attention to Non-Adopters
- Everyone prefers human writers, including AI
- Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired Sketching
- Large Language Models Do Not Simulate Human Psychology
- Whither symbols in the era of advanced neural networks?
- Learning without training: The implicit dynamics of in-context learning
- Large language models for scholarly ontology generation: An extensive analysis in the engineering field
- Research Community Perspectives on "Intelligence" and Large Language Models
- Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis
- Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
- The Art of Audience Engagement: LLM-Based Thin-Slicing of Scientific Talks
- Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
- LLM Social Simulations Are a Promising Research Method
- How Deep Do Large Language Models Internalize Scientific Literature and Citation Practices?
- A Taxonomy of Linguistic Expressions That Contribute To Anthropomorphism of Language Technologies
- AI Sees Your Location, But With A Bias Toward The Wealthy World
- Gender disparities in the impact of generative artificial intelligence: Evidence from academia.
- Playing With AI: How Do State-Of-The-Art Large Language Models Perform in the 1977 Text-Based Adventure Game Zork?
- Towards High-Level Semantic Intelligence
- Artificial General Intelligence (AGI)-Native Wireless Systems: A Journey Beyond 6G
- A Conjecture on a Fundamental Trade-Off between Certainty and Scope in Symbolic and Generative AI
- Unifying Learning Dynamics and Generalization in Transformers Scaling Law
- PTStore (Prefix Tensor Store): Distributed Prefix Caching and Replication for High Throughput Inference Serving
- The Cartesian Cut in Agentic AI
- Coherent without Grounding, Grounded without Success: The Bidirectional Coherence Paradox in Artificial Epistemic Agents
- AegisAgent: An Autonomous Defense Agent Against Prompt Injection Attacks in LLM-HARs
- How Tech Workers Contend with Hazards of Humanlikeness in Generative AI
- Humanlike AI Design Increases Anthropomorphism but Yields Divergent Outcomes on Engagement and Trust Globally
- Can abstract concepts from LLM improve SLM performance?
- CienaLLM: Generative Climate-Impact Extraction from News Articles with Autoregressive LLMs
- External Hippocampus: Topological Cognitive Maps for Guiding Large Language Model Reasoning
- From Priors to Predictions: Explaining and Visualizing Human Reasoning in a Graph Neural Network Framework
- Plausibility as Failure: How LLMs and Humans Co-Construct Epistemic Error
- Quantifying Return on Security Controls in LLM Systems
- Large Language Newsvendor: Decision Biases and Cognitive Mechanisms
- One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs
- Large Language Models have Chain-of-Affect
- Evolutionary Reinforcement Learning based AI tutor for Socratic Interdisciplinary Instruction
- On the Dynamics of Multi-Agent LLM Communities Driven by Value Diversity
- When Medical AI Explanations Help and When They Harm
- Procrustean Bed for AI-Driven Retrosynthesis: A Unified Framework for Reproducible Evaluation
- SJD++: Improved Speculative Jacobi Decoding for Training-free Acceleration of Discrete Auto-regressive Text-to-Image Generation
- On the Computability of Artificial General Intelligence
- MCP-AI: Protocol-Driven Intelligence Framework for Autonomous Reasoning in Healthcare
- Nex-N1: Agentic Models Trained via a Unified Ecosystem for Large-Scale Environment Construction
- Catching UX Flaws in Code: Leveraging LLMs to Identify Usability Flaws at the Development Stage
- AsymPuzl: An Asymmetric Puzzle for multi-agent cooperation
- LLM-Generated Ads: From Personalization Parity to Persuasion Superiority
- ASCIIBench: Evaluating Language-Model-Based Understanding of Visually-Oriented Text
- A Human-centric Framework for Debating the Ethics of AI Consciousness Under Uncertainty
- WISE: Weighted Iterative Society-of-Experts for Robust Multimodal Multi-Agent Debate
- LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess
- Towards Active Synthetic Data Generation for Finetuning Language Models
- A Comparison of Human and ChatGPT Classification Performance on Complex Social Media Data
- Memory-Amortized Inference: A Topological Unification of Search, Closure, and Structure
- The Geometry of Certainty: Recursive Topological Condensation and the Limits of Inference
- TrafficLens: Multi-Camera Traffic Video Analysis Using LLMs
- FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
- MoodBench 1.0: An Evaluation Benchmark for Emotional Companionship Dialogue Systems
- Identifying Quantum Structure in AI Language: Evidence for Evolutionary Convergence of Human and Artificial Cognition
- Bridging Symbolic Control and Neural Reasoning in LLM Agents -- The Structured Cognitive Loop
- The Impact of Quantization on Large Reasoning Model Reinforcement Learning
- Automatic Pruning Discovery for Large Language Models
- Realist and Pluralist Conceptions of Intelligence and Their Implications on AI Research
- Agent READMEs: An Empirical Study of Context Files for Agentic Coding
- Structured Decomposition for LLM Reasoning: Cross-Domain Validation and Semantic Web Integration
- LAET: A Layer-wise Adaptive Ensemble Tuning Framework for Pretrained Language Models
- On the Measure of a Model: From Intelligence to Generality
- Does Scientific Writing Converge to U.S. English? Evidence from Generative AI-Assisted Publications
- Spontaneous eye movements reflect the representational geometries of conceptual spaces
- Place Matters: Comparing LLM Hallucination Rates for Place-Based Legal Queries
- Integrating large language models into EFL writing instruction: effects on performance, self-regulated learning strategies, and motivation
- A Collaborative Reasoning Framework for Anomaly Diagnostics in Underwater Robotics
- AthenaBench: A Dynamic Benchmark for Evaluating LLMs in Cyber Threat Intelligence
- DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning
- The Pervasive Blind Spot: Benchmarking VLM Inference Risks on Everyday Personal Videos
- FP8-Flow-MoE: A Casting-Free FP8 Recipe without Double Quantization Error
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
- TSVer: A Benchmark for Fact Verification Against Time-Series Evidence
- TempoBench: Evaluating Temporal Causal Reasoning in Large Language Models
- Budgeted Multiple-Expert Deferral
- 1+1>2: A Synergistic Sparse and Low-Rank Compression Method for Large Language Models
- Artificial Intelligence in Elementary STEM Education: A Systematic Review of Current Applications and Future Challenges
- TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
- TextualVerifier: Verify TextGrad Step-by-Step
- Testing theory of mind in large language models and humans
- Reclaiming AI as a theoretical tool for cognitive science
- MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning
- Large models of what? Mistaking engineering achievements for human linguistic agency
- Artificial Intelligence in Health Professions Education assessment: AMEE Guide No. 178
- Large Language Models and the Future of Organization Theory
- The TESCREAL bundle: Eugenics and the promise of utopia through artificial general intelligence
- Which Humans?
- Pixels and Predictions: Potential of GPT-4V in Meteorological Imagery Analysis and Forecast Communication
- How close is AI to human-level intelligence?
- Mapping the Mind With Free Associations: A Tutorial Using the R Package associatoR
- The cognitive biases that may exacerbate inflationary and deflationary positions about large language models
- Can we Trust Chatbots for now? Accuracy, reproducibility, traceability; a Case Study on Leonardo da Vinci's Contribution to Astronomy
- The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science
- The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem
- Bridging the Divide: End-to-End Sequence-Graph Learning
- Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants
- Education Paradigm Shift To Maintain Human Competitive Advantage Over AI
- Will Humanity Be Rendered Obsolete by AI?
- RaCoT: Plug-and-Play Contrastive Example Generation Mechanism for Enhanced LLM Reasoning Reliability
- Frustratingly Easy Task-aware Pruning for Large Language Models
- Learning "Partner-Aware" Collaborators in Multi-Party Collaboration
- Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents: Pathways and Paradigms
- When and Why Does Multi-Agent Debate Fail and Does It Really Underperform?
- Think Parallax: Solving Multi-Hop Problems via Multi-View Knowledge-Graph-Based Retrieval-Augmented Generation
- LLMartini: Seamless and Interactive Leveraging of Multiple LLMs through Comparison and Composition
- A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
- Learning Efficient and Generalizable Graph Retriever for Knowledge-Graph Question Answering
- Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge
- Presenting Large Language Models as Companions Affects What Mental Capacities People Attribute to Them
- AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM
- LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet Arena
- Pedoman Peri-operatif Renal Assessment
- The Spark Effect: On Engineering Creative Diversity in Multi-Agent AI Systems
- Assessing Coherency and Consistency of Code Execution Reasoning by Large Language Models
- Demystifying the Mechanisms Behind Emergent Exploration in Goal-conditioned RL
- Static Sandboxes Are Inadequate: Modeling Societal Complexity Requires Open-Ended Co-Evolution in LLM-Based Multi-Agent Simulations
- Doing Things with Words: Rethinking Theory of Mind Simulation in Large Language Models
- Interpreting the Latent Structure of Operator Precedence in Language Models
- A Survey on Evaluation of Large Language Models
- PADME: Procedure Aware DynaMic Execution
- Automating Structural Engineering Workflows with Large Language Model Agents
- BanglaMATH : A Bangla benchmark dataset for testing LLM mathematical reasoning at grades 6, 7, and 8
- HyperAgent: Leveraging Hypergraphs for Topology Optimization in Multi-Agent Communication
- MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation
- A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis
- Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey
- On the Role of Domain Experts in Creating Effective Tutoring Systems
- Hidden Secrets in the arXiv: Discovering, Analyzing, and Preventing Unintentional Information Disclosure in Source Files of Scientific Preprints
- CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMs
- Improving AGI Evaluation: A Data Science Perspective
- DICE: Structured Reasoning in LLMs through SLM-Guided Chain-of-Thought Correction
- TinyGraphEstimator: Adapting Lightweight Language Models for Graph Structure Inference
- Opponent Shaping in LLM Agents
- Assessing the nature of large language models: A caution against anthropocentrism
- In-Context Clustering with Large Language Models
- Lemma Dilemma: On Lemma Generation Without Domain- or Language-Specific Training Data
- Large language models show human-like content biases in transmission chain experiments
- Exploring regional vulnerability to the Fourth Industrial Revolution: a European perspective
- Foundations of LLM Knowledge Materialization: Termination, Reproducibility, Robustness
- Aligning Large Language Models via Fully Self-Synthetic Data
- Iterative LLM-Based Generation and Refinement of Distracting Conditions in Math Word Problems
- On the Role of Difficult Prompts in Self-Play Preference Optimization
- More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models
- GenQuest: An LLM-based Text Adventure Game for Language Learners
- MetaMuse: Algorithm Generation via Creative Ideation
- AgenticRAG: Tool-Augmented Foundation Models for Zero-Shot Explainable Recommender Systems
- Homophily-induced Emergence of Biased Structures in LLM-based Multi-Agent AI Systems
- Batch-CAM: Introduction to better reasoning in convolutional deep learning models
- An empirical investigation of the impact of ChatGPT on creativity
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- UniAPL: A Unified Adversarial Preference Learning Framework for Instruct-Following
- Hallucination is Inevitable for LLMs with the Open World Assumption
- LatentEvolve: Self-Evolving Test-Time Scaling in Latent Space
- Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMs
- Enabling Physical AI through Biological Principles
- Effectiveness of Large Language Models in Simulating Regional Psychological Structures: An Empirical Examination of Personality and Subjective Well-being
- Watermarking Diffusion Language Models
- The future of academic publishing
- A Theoretical Computer Science Perspective on Consciousness and Artificial General Intelligence
- Towards Human-interpretable Explanation in Code Clone Detection using LLM-based Post Hoc Explainer
- GPT (Generative Pre-Trained Transformer)— A Comprehensive Review on Enabling Technologies, Potential Applications, Emerging Challenges, and Future Directions
- Linear Causal Representation Learning by Topological Ordering, Pruning, and Disentanglement
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- “It happened to be the perfect thing”: experiences of generative AI chatbots for mental health
- Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect Personas
- Can LLMs Forecast Internet Traffic from Social Media?
- The future of machine learning for small-molecule drug discovery will be driven by data
- When and how to disclose AI use in academic publishing: AMEE Guide No.192
- Generative AI for Economic Research: Use Cases and Implications for Economists
- AI Gamestore: Scalable, Open-Ended Evaluation of Machine General Intelligence with Human Games
- A large-scale evaluation of commonsense knowledge in humans and large language models
- Patterns Lead the Way to Far-from-Equilibrium Materials
- Generative artificial intelligence and engineering education
- The Brain Abstracted
- Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models
- Through the Lens of Human-Human Collaboration: A Configurable Research Platform for Exploring Human-Agent Collaboration
- nDNA -- the Semantic Helix of Artificial Cognition
- Small LLMs with Expert Blocks Are Good Enough for Hyperparamter Tuning
- Learning in Context: Personalizing Educational Content with Large Language Models to Enhance Student Learning
- Charting trajectories of human thought using large language models
- Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
- AI For Privacy in Smart Homes: Exploring How Leveraging AI-Powered Smart Devices Enhances Privacy Protection
- CultranAI at PalmX 2025: Data Augmentation for Cultural Knowledge Representation
- Artificial Intelligence and Entrepreneurship: A Call for Research to Prospect and Establish the Scholarly AI Frontiers
- Genome-wide prediction of disease variant effects with a deep protein language model
- Human resource management in the age of generative artificial intelligence: Perspectives and research directions on ChatGPT
- Cycle is All You Need: More Is Different
- Co-Alignment: Rethinking Alignment as Bidirectional Human-AI Cognitive Adaptation
- Data-Driven Analysis of Text-Conditioned AI-Generated Music: A Case Study with Suno and Udio
- MAPGD: Multi-Agent Prompt Gradient Descent for Collaborative Prompt Optimization
- Compartmentalised Agentic Reasoning for Clinical NLI
- Towards Fully Automated Molecular Simulations: Multi-Agent Framework for Simulation Setup and Force Field Extraction
- AI Wellbeing
- Exploring the Impact of Generative Artificial Intelligence on Software Development in the IT Sector: Preliminary Findings on Productivity, Efficiency and Job Security
- Investigating Student Interaction Patterns with Large Language Model-Powered Course Assistants in Computer Science Courses
- Mitigating Catastrophic Forgetting in Large Language Models with Forgetting-aware Pruning
- Language Self-Play For Data-Free Training
- Disentangling Interaction and Bias Effects in Opinion Dynamics of Large Language Models
- Code2MCP: Transforming Code Repositories into MCP Services
- The human biological advantage over AI
- The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
- DiaCBT: A Long-Periodic Dialogue Corpus Guided by Cognitive Conceptualization Diagram for CBT-based Psychological Counseling
- Assessing Consciousness-Related Behaviors in Large Language Models Using the Maze Test
- On the Alignment of Large Language Models with Global Human Opinion
- Analysis of Error Sources in LLM-based Hypothesis Search for Few-Shot Rule Induction
- Transforming Agency. On the mode of existence of Large Language Models
- CVPD at QIAS 2025 Shared Task: An Efficient Encoder-Based Approach for Islamic Inheritance Reasoning
- Going over Fine Web with a Fine-Tooth Comb: Technical Report of Indexing Fine Web for Problematic Content Search and Retrieval
- Personality Matters: User Traits Predict LLM Preferences in Multi-Turn Collaborative Tasks
- Integrating Large Language Models with Network Optimization for Interactive and Explainable Supply Chain Planning: A Real-World Case Study
- Pruning Strategies for Backdoor Defense in LLMs
- Secure Multi-LLM Agentic AI and Agentification for Edge General Intelligence by Zero-Trust: A Survey
- Emotion Transfer with Enhanced Prototype for Unseen Emotion Recognition in Conversation
- MQAD: A Large-Scale Question Answering Dataset for Training Music Large Language Models
- Grounding the Ungrounded: A Spectral-Graph Framework for Quantifying Hallucinations in Multimodal LLMs
- APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration
- A Concurrent Modular Agent: Framework for Autonomous LLM Agents
- S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner
- ELATE: Evolutionary Language model for Automated Time-series Engineering
- Cohort-Aware Agents for Individualized Lung Cancer Risk Prediction Using a Retrieval-Augmented Model Selection Framework
- Leveraging Large Language Models for Predictive Analysis of Human Misery
- Uncovering Spontaneous Physics Representations in In-Context Learning
- Structuring the Unstructured: A Systematic Review of Text-to-Structure Generation for Agentic AI with a Universal Evaluation Framework
- SupraTok: Cross-Boundary Tokenization for Enhanced Language Model Performance
- Applied causality to infer protein dynamics and kinetics
- AI That Helps Us Help Each Other: A Proactive System for Scaffolding Mentor-Novice Collaboration in Entrepreneurship Coaching
- What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles
- EvoCurr: Self-evolving Curriculum with Behavior Code Generation for Complex Decision-making
- Large Language Models Show Signs of Alignment with Human Neurocognition During Abstract Reasoning
- Dynamic Uncertainty-aware Multimodal Fusion for Outdoor Health Monitoring
- Compass-Thinker-7B Technical Report
- "Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas
- Multimodal learning with next-token prediction for large multimodal models
- Intuition emerges in Maximum Caliber models at criticality
- Between Tool and Trouble: Student Attitudes Toward AI in Programming Education
- Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support
- Panel-Scale Reconfigurable Photonic Interconnects for Scalable AI Computation
- LLMs for Resource Allocation: A Participatory Budgeting Approach to Inferring Preferences
- AGI for the Earth, the path, possibilities and how to evaluate intelligence of models that work with Earth Observation Data?
- P-CoT: A Pedagogically-motivated Participatory Chain-of-Thought Prompting for Phonological Reasoning in LLMs
- FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and Challenging
- Why are LLMs' abilities emergent?
- The Science Fiction Science Method
- A Pragmatist Robot: Learning to Plan Tasks by Experiencing the Real World
- VLMQ: Efficient Post-Training Quantization for Large Vision-Language Models via Hessian Augmentation
- Autonomous Inorganic Materials Discovery via Multi-Agent Physics-Aware Scientific Reasoning
- Authorship Attribution in Multilingual Machine-Generated Texts
- How Does Controllability Emerge In Language Models During Pretraining?
- Autonomous Penetration Testing: Solving Capture-the-Flag Challenges with LLMs
- Distributed AI Agents for Cognitive Underwater Robot Autonomy
- Improving Generative Ad Text on Facebook using Reinforcement Learning
- LLM world models are mental: Output layer evidence of brittle world model use in LLM mechanical reasoning
- Semantic Convergence: Investigating Shared Representations Across Scaled LLMs
- Rep-MTL: Unleashing the Power of Representation-level Task Saliency for Multi-Task Learning
- Length Representations in Large Language Models
- PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
- Policy-Driven AI in Dataspaces: Taxonomy, Explainability, and Pathways for Compliant Innovation
- Inducing Causal World Models in LLMs for Zero-Shot Physical Reasoning
- DeltaLLM: A Training-Free Framework Exploiting Temporal Sparsity for Efficient Edge LLM Inference
- Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes
- Objectifying the Subjective: Cognitive Biases in Topic Interpretations
- Large language models provide unsafe answers to patient-posed medical questions
- SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models
- VeriMinder: Mitigating Analytical Vulnerabilities in NL2SQL
- Rethinking Memorization Measures and their Implications in Large Language Models
- Harnessing LLMs for Document-Guided Fuzzing of OpenCV Library
- Making Abstraction Concrete: A Design Space and Interaction Model of Abstraction in Interactive Systems
- Strategic Polysemy in AI Discourse: A Philosophical Analysis of Language, Hype, and Power
- Real Money, Fake Models: Deceptive Model Claims in Shadow APIs
- Theory of Mind Abilities of Large Language Models in Human-Robot Interaction: An Illusion?
- Can LLMs Infer Personality from Real World Conversations?
- Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks
- Augmented Question-guided Retrieval (AQgR) of Indian Case Law with LLM, RAG, and Structured Summaries
- From Text to Actionable Intelligence: Automating STIX Entity and Relationship Extraction
- Playing repeated games with large language models
- Aligning Knowledge Graphs and Language Models for Factual Accuracy
- RETRACTED: EFL university students’ emotional engagement in AI ‐mediated learning contexts: A sentiment analysis
- MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions
- From Semantic Web and MAS to Agentic AI: A Unified Narrative of the Web of Agents
- Cultural Bias in Large Language Models: Evaluating AI Agents through Moral Questionnaires
- Stabilizing Black-Box Prompt Optimization with Textual Regularization and Signal Aggregation
- Pluri-perspectivism in Human-robot Co-creativity with Older Adults
- Voltage Regulation in Distribution Systems with Data Center Loads
- Generative AI for Software Practitioners
- Video Event Reasoning and Prediction by Fusing World Knowledge from LLMs with Vision Foundation Models
- Development and Evaluation of HopeBot: an LLM-based chatbot for structured and interactive PHQ-9 depression screening
- The Case for Instance-Optimized LLMs in OLAP Databases
- On the Semantics of Large Language Models
- Towards Machine Theory of Mind with Large Language Model-Augmented Inverse Planning
- Unanticipated Effects of Generative AI on Expertise Pathways and Performance Perception in System Administration
- Strategic Intelligence in Large Language Models: Evidence from evolutionary Game Theory
- Position: A Theory of Deep Learning Must Include Compositional Sparsity
- MedVAL: Toward Expert-Level Medical Text Validation with Language Models
- Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models
- Towards the Holographic Characteristic of LLMs for Efficient Short-text Generation
- Evaluation of GPT-4o and GPT-4o-Mini’s Vision Capabilities for Compositional Analysis from Dried Solution Drops
- Spatial transcriptomics AI agent charts hPSC-pancreas maturation in vivo
- Holistic Artificial Intelligence in Medicine; improved performance and explainability
- Prompting as Scientific Inquiry
- Can "consciousness" be observed from large language model (LLM) internal states? Dissecting LLM representations obtained from Theory of Mind test with Integrated Information Theory and Span Representation analysis
- From 2D to 3D Cognition: A Brief Survey of General World Models
- Mastering Multiple-Expert Routing: Realizable H-Consistency and Strong Guarantees for Learning to Defer
- Narrative Shift Detection: A Hybrid Approach of Dynamic Topic Models and Large Language Models
- Can One Safety Loop Guard Them All? Agentic Guard Rails for Federated Computing
- Distillation-Enabled Knowledge Alignment for Generative Semantic Communications in AIGC Provisioning Tasks
- GraspMAS: Zero-Shot Language-driven Grasp Detection with Multi-Agent System
- Programming Quantum Computers with Large Language Models
- A Comparative Study of Open-Source Libraries for Synthetic Tabular Data Generation: SDV vs. SynthCity
- Online Multi-LLM Selection via Contextual Bandits under Unstructured Context Evolution
- Safe Pruning LoRA: Robust Distance-Guided Pruning for Safety Alignment in Adaptation of LLMs
- Co-VisiON: Co-Visibility ReasONing on Sparse Image Sets of Indoor Scenes
- Predicting New Research Directions in Materials Science using Large Language Models and Concept Graphs
- Explainable Rule Application via Structured Prompting: A Neural-Symbolic Approach
- Bridging Brain with Foundation Models through Self-Supervised Learning
- Quantile Regression with Large Language Models for Price Prediction
- High computational density nanophotonic media for machine learning inference
- From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models
- Causes in neuron diagrams, and testing causal reasoning in Large Language Models. A glimpse of the future of philosophy?
- Translating Federated Learning Algorithms in Python into CSP Processes Using ChatGPT
- Constitutive Components for Human-Like Autonomous Artificial Intelligence
- Assessing the Role of Data Quality in Training Bilingual Language Models
- OneEval: Benchmarking LLM Knowledge-intensive Reasoning over Diverse Knowledge Bases
- Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory
- Are Multimodal Large Language Models Pragmatically Competent Listeners in Simple Reference Resolution Tasks?
- How large language models can reshape collective intelligence
- Transforming Physiology and Healthcare through Foundation Models
- Formalizing Learning from Language Feedback with Provable Guarantees
- Superstudent intelligence in thermodynamics
- Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives
- Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models
- Sample Complexity and Representation Ability of Test-time Scaling Paradigms
- Sensory-Motor Control with Large Language Models via Iterative Policy Refinement
- Agents of Change: Self-Evolving LLM Agents for Strategic Planning
- Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
- LLM Code Customization with Visual Results: A Benchmark on TikZ
- Lossless Compression of Large Language Model-Generated Text via Next-Token Prediction
- R3-VQA: "Read the Room" by Video Social Reasoning
- Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving
- Delta-KNN: Improving Demonstration Selection in In-Context Learning for Alzheimer's Disease Detection
- "Don't Do That!": Guiding Embodied Systems through Large Language Model-based Constraint Generation
- Autonomous chemical research with large language models
- Linear Spatial World Models Emerge in Large Language Models
- XToM: Exploring the Multilingual Theory of Mind for Large Language Models
- Decompose, Plan in Parallel, and Merge: A Novel Paradigm for Large Language Models based Planning with Multiple Constraints
- StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs
- Natural, Artificial, and Human Intelligences
- RoboEgo System Card: An Omnimodal Model with Native Full Duplexity
- RAISE: Reasoning Agent for Interactive SQL Exploration
- Fodor and Pylyshyn's Legacy: Still No Human-like Systematic Compositionality in Neural Networks
- Artificial creativity: can there be creativity without cognition?
- zip2zip: Inference-Time Adaptive Tokenization via Online Compression
- Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know?
- Bridging Expertise Gaps: The Role of LLMs in Human-AI Collaboration for Cybersecurity
- From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models
- From Macro to Micro: Probing Dataset Diversity in Language Model Fine-Tuning
- Adaptable Cardiovascular Disease Risk Prediction from Heterogeneous Data using Large Language Models
- Beyond Semantic Entropy: Boosting LLM Uncertainty Quantification with Pairwise Semantic Similarity
- ChARM: Character-based Act-adaptive Reward Modeling for Advanced Role-Playing Language Agents
- Cross-Task Experiential Learning on LLM-based Multi-Agent Collaboration
- Be.FM: Open Foundation Models for Human Behavior
- Neither Stochastic Parroting nor AGI: LLMs Solve Tasks through Context-Directed Extrapolation from Training Data Priors
- Enhancing Large Language Models'Machine Translation via Dynamic Focus Anchoring
- A Mathematical Framework for AI-Human Integration in Work
- The Turing Test Is More Relevant Than Ever
- Contextual Memory Intelligence -- A Foundational Paradigm for Human-AI Collaboration and Reflective Generative AI Systems
- Co-Saving: Resource Aware Multi-Agent Collaboration for Software Development
- Distinguishing Fact from Fiction: Student Traits, Attitudes, and AI Hallucination Detection in Business School Assessment
- Divide-Then-Align: Honest Alignment based on the Knowledge Boundary of RAG
- SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis
- Fundamental Limits of Game-Theoretic LLM Alignment: Smith Consistency and Preference Matching
- Is Your LLM Overcharging You? Tokenization, Transparency, and Incentives
- Self-Reflective Planning with Knowledge Graphs: Enhancing LLM Reasoning Reliability for Question Answering
- Multi-Agent Collaboration via Evolving Orchestration
- Large Language Models' Reasoning Stalls: An Investigation into the Capabilities of Frontier Models
- Minimalist Softmax Attention Provably Learns Constrained Boolean Functions
- Do Large Language Models (Really) Need Statistical Foundations?
- SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs
- AI-Driven Climate Policy Scenario Generation for Sub-Saharan Africa
- BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook
- μ-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
- LatentLLM: Attention-Aware Joint Tensor Compression
- AI-Augmented LLMs Achieve Therapist-Level Responses in Motivational Interviewing
- Next-token pretraining implies in-context learning
- Sudoku-Bench: Evaluating creative reasoning with Sudoku variants
- Shadows in the Attention: Contextual Perturbation and Representation Drift in the Dynamics of Hallucination in LLMs
- Human-like Semantic Navigation for Autonomous Driving using Knowledge Representation and Large Language Models
- OpenEthics: A Comprehensive Ethical Evaluation of Open-Source Generative Large Language Models
- Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge
- Multilingual Test-Time Scaling via Initial Thought Transfer
- What Does Success Look Like? Catalyzing Meeting Intentionality with AI-Assisted Prospective Reflection
- Interpretable Traces, Unexpected Outcomes: Investigating the Disconnect in Trace-Based Knowledge Distillation
- Creative Preference Optimization
- Toward Embodied AGI: A Review of Embodied AI and the Road Ahead
- ABBA-Adapters: Efficient and Expressive Fine-Tuning of Foundation Models
- Gauging Growth: AGI Mathematical Metrics for Economic Progress
- Understanding Task Representations in Neural Networks via Bayesian Ablation
- Attention-based clustering
- Language and Thought: The View from LLMs
- On the Thinking-Language Modeling Gap in Large Language Models
- The Epistemic Politics of AI Anthropomorphism
- Social preferences with unstable interactive reasoning: Large language models in economic trust games
- LD-Scene: LLM-Guided Diffusion for Controllable Generation of Adversarial Safety-Critical Driving Scenarios
- DiffuseAgent-MI: Distributionally-Grounded,Tool-Integrated Self-Evolving Agents for Faithful Visual Reasoning
- Commemorar, especular: algunes tendències en la poesia editada a les Illes Balears
- WorldView-Bench: A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models
- Decomposed Inductive Procedure Learning: Learning Academic Tasks with Human-Like Data Efficiency
- AI-enhanced semantic feature norms for 786 concepts
- Educational impacts of generative artificial intelligence on learning and performance of engineering students in China
- 14 examples of how LLMs can transform materials science and chemistry: a reflection on a large language model hackathon
- Language Agents Mirror Human Causal Reasoning Biases. How Can We Help Them Think Like Scientists?
- Towards Contamination Resistant Benchmarks
- DSADF: Thinking Fast and Slow for Decision Making
- Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training
- WaLLM -- Insights from an LLM-Powered Chatbot deployment via WhatsApp
- The Wicked Nature of AGI
- Towards Requirements Engineering for RAG Systems
- Control Plane as a Tool: A Scalable Design Pattern for Agentic AI Systems
- Building a Human-Verified Clinical Reasoning Dataset via a Human LLM Hybrid Pipeline for Trustworthy Medical AI
- Bridging AI and Carbon Capture: A Dataset for LLMs in Ionic Liquids and CBE Research
- Applying Cognitive Design Patterns to General LLM Agents
- Insertion Language Models: Sequence Generation with Arbitrary-Position Insertions
- ComPO: Preference Alignment via Comparison Oracles
- Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
- A Survey on Privacy Risks and Protection in Large Language Models
- Leveraging LLM Agents and Digital Twins for Fault Handling in Process Plants
- Restoring Calibration for Aligned Large Language Models: A Calibration-Aware Fine-Tuning Approach
- QiMeng-Xpiler: Transcompiling Tensor Programs for Deep Learning Systems with a Neural-Symbolic Approach
- Beyond Recognition: Evaluating Visual Perspective Taking in Vision Language Models
- Passing the Buck to AI: How Individuals' Decision-Making Patterns Affect Reliance on AI
- Value Portrait: Assessing Language Models' Values through Psychometrically and Ecologically Valid Items
- Should AI Mimic People? Understanding AI-Supported Writing Technology Among Black Users
- An Onto-Relational-Sophic Framework for Governing Synthetic Minds
- MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs
- Alignment among Language, Vision and Action Representations
- How Small Can 6G Reason? Scaling Tiny-to-Small Language Models for AI-Native Networks
- Demystifying the oracle: A "20 Questions" game to promote AI ethics and literacy
- Reasoning Capabilities and Invariability of Large Language Models
- From Texts to Shields: Convergence of Large Language Models and Cybersecurity
- The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning
- HUOZIIME: An On-Device LLM-enhanced Input Method for Deep Personalization
- A New Strategy for Artificial Intelligence: Training Foundation Models Directly on Human Brain Data
- Engineering-Oriented Symbolic Regression: LLMs as Physics Agents for Discovery of Simulation-Ready Constitutive Laws
- Skill Discovery for Software Scripting Automation via Offline Simulations with LLMs
- A Systematic Literature Review of Parameter-Efficient Fine-Tuning for Large Code Models
- Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action
- Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents
- PolyBench: Benchmarking LLM Forecasting and Trading Capabilities on Live Prediction Market Data
- HoneyTrap: Deceiving Large Language Model Attackers to Honeypot Traps with Resilient Multi-Agent Defense
- semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
- LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects
- River-LLM: Large Language Model Seamless Exit Based on KV Share
- Dynamics of Cognitive Heterogeneity: Investigating Behavioral Biases in Multi-Stage Supply Chains with LLM-Based Simulation
- Theory of Mind in Large Language Models: Assessment and Enhancement
- FOCUS: DLLMs Know How to Tame Their Compute Bound
- Autonomous Market Intelligence: Agentic AI Nowcasting Predicts Stock Returns
- VerifAI: A Verifiable Open-Source Search Engine for Biomedical Question Answering
- ReCellTy: Domain-specific knowledge graph retrieval-augmented LLMs workflow for single-cell annotation
- PACE: Adaptive Budget Allocation for Time-Efficient Embodied Planning
- An Empirical Study of Evaluating Long-form Question Answering
- Optimism, Expectation, or Sarcasm? Multi-Class Hope Speech Detection in Spanish and English
- AI Awareness
- Private Direct Preference Optimization for LLM Alignment
- Speech-to-Trajectory: Learning Human-Like Verbal Guidance for Robot Motion
- Rethinking Reflection in Pre-Training
- ProbRes: Probabilistic Jump Diffusion for Open-World Egocentric Activity Recognition
- How Private is Your Attention? Bridging Privacy with In-Context Learning
- Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression
- Research on Navigation Methods Based on LLMs
- RainbowPlus: Enhancing Adversarial Prompt Generation via Evolutionary Quality-Diversity Search
- Object-Level Verbalized Confidence Calibration in Vision-Language Models via Semantic Perturbation
- Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes
- Contemplative Agent
- Evaluating large language models on a highly-specialized topic, radiation oncology physics
- Enhancing systematic reviews in orthodontics: a comparative examination of GPT-3.5 and GPT-4 for generating PICO-based queries with tailored prompts and configurations
- Identifying and Mitigating the Influence of the Prior Distribution in Large Language Models
- You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models
- Agentic Nesting: A New Methodology for Existing Enterprise Application Integration and Services
- ARise: Towards Knowledge-Augmented Reasoning via Risk-Adaptive Search
- Towards A Universal Graph Structural Encoder
- StruPhantom: Evolutionary Injection Attacks on Black-Box Tabular Agents Powered by Large Language Models
- You've Changed: Detecting Modification of Black-Box Large Language Models
- Two Heads are Better Than One: Test-time Scaling of Multi-agent Collaborative Reasoning
- Evaluating the Quality of Benchmark Datasets for Low-Resource Languages: A Case Study on Turkish
- Large Language Models integration in Smart Grids
- An Empirical Study of Production Incidents in Generative AI Cloud Services
- A Survey of Reasoning with Foundation Models: Concepts, Methodologies, and Outlook
- Generating Fine Details of Entity Interactions
- Hallucination, reliability, and the role of generative AI in science
- Open Problems and a Hypothetical Path Forward in LLM Knowledge Paradigms
- Artificial intelligence in medicine: How it works, how it fails
- Graph-based Approaches and Functionalities in Retrieval-Augmented Generation: A Comprehensive Survey
- TRATSS: Transformer-Based Task Scheduling System for Autonomous Vehicles
- BRIDGES: Bridging Graph Modality and Large Language Models within EDA Tasks
- GPT-4 [wikipedia]
- History of artificial intelligence [wikipedia]
- January–March 2023 in science [wikipedia]
- Large language model [wikipedia]
- Pause Giant AI Experiments: An Open Letter [wikipedia]
- Products and applications of OpenAI [wikipedia]
- Sébastien Bubeck [wikipedia]
- Sally–Anne test [wikipedia]
- Superintelligence [wikipedia]
- Timeline of computing 2020–present [wikipedia]
Discussions
- Sparks of Artificial General Intelligence: Early Experiments with GPT-4 [hn, 180 points, 236 comments]
- Sparks of AGI: early experiments with GPT-4 (6.04.2023 yt vid) [lemmy, 7 points, 5 comments]
- Es ist schwierig, Literatur zu AGI zu empfehlen, da bislang keine allgemein akzeptierte Definition zu AGI existiert. Einigkeit besteht nur darüber, dass rein künstliche neuronale Netze in ihrer heutig [bsky, 6 points, 1 comments]
- 2024-07-01 [lemmy, 6 points, 0 comments]
- Não há produção científica que justifique essa tese. Uns pesquisadores da Microsoft chegaram a publicar um artigo nesse sentido (arxiv.org/abs/2303.12712), mas não é convincente. [bsky, 5 points, 1 comments]
- Sparks of AGI: early experiments with GPT-4 (2 month old yt vid) [lemmy, 4 points, 0 comments]
- Sparks of Artificial General Intelligence: Early experiments with GPT-4 [lobsters, 3 points, 0 comments]
- A paper from researchers at Microsoft claims AI shows the ability to understand the way people do. This is a report from the New York Times. The paper: “Sparks of Artificial General Intelligence.” a [bsky, 3 points, 1 comments]
- It was, version one of the document also had the teeny tiny problem of using a study that was basically demographical phrenology to 'measure intelligence' (reference from the first line of the conclus [bsky, 2 points, 1 comments]
- Sparks of Artificial General Intelligence: Early Experiments with GPT-4 (2023) [hn, 2 points, 0 comments]
- I'm trying to understand the very impressive GPT-4 example from this paper: https://arxiv.org/pdf/2303.12712.pdf titled "Sparks of Artificial General Intelligence: Early experiments with GPT-4" [bsky, 1 points, 1 comments]
- It's not AGI. It has weaknesses (which are different from human weaknesses - the weaknesses of humans always get left out of this conversation, IMHO). But it also has great strengths. Good research [bsky, 1 points, 1 comments]
- Very nice documentation 👍 Do you know this extensive comparison of a few bots from 2023? Just skimming through its pictures still leads me to interpret "AI" as "Acquired Incompetence", reflecting bo [bsky, 1 points, 1 comments]
- Sparks of Artificial General Intelligence: Early Experiments with GPT-4 [bsky, 1 points, 0 comments]
- Paper -> arxiv.org/abs/2303.12712 [bsky, 1 points, 1 comments]
- KI lernen nicht, sie werden trainiert. Entweder durch Menschen oder durch "Pretrained ANN" (die wiederum von Menschen trainiert wurden). Google "GAN" & "diskriminatorische Modelle" & "generative Model [bsky, 1 points, 4 comments]
- No prob! If you're still on the fence, you can always check their arxiv preprint: https://arxiv.org/pdf/2303.12712 Thought it was pretty impressive. Also Poe also gives you 1 free query a day, iirc. [bsky, 1 points, 1 comments]
- Tests by MSR on pre-RLHF GPT4, very famous paper: arxiv.org/abs/2303.12712 Pre-RLHF is importent because RLHF dumbs down models (imo) and forces them to output a certain way. Which is why we saw goo [bsky, 0 points, 1 comments]
- AGI isn't a well-defined concept but similarly a matter of degree and what skills are being measured. "Given the breadth and depth of GPT-4’s capabilities, we believe that it could reasonably be view [bsky, 0 points, 1 comments]
- Of course this was the subject of the [Sparks of AGI](arxiv.org/pdf/2303.12712) paper. Same thing on [Claude 3.5 Sonnet](claude.site/artifacts/3f...) [bsky, 0 points, 1 comments]
- GPT-4 might be showing “sparks” of AGI according to this paper. For example it can and knows when to use tools… arxiv.org/abs/2303.12712 [bsky, 0 points, 0 comments]
Related