Sparks of Artificial General Intelligence: Early experiments with GPT-4
2023/03/22 by Sébastien Bubeck, Bubeck, Sébastien, Varun Chandrasekaran +26 · 21 voices · 276 citations
Medicine · Computer Science · #Artificial Intelligence in Healthcare and Education #Topic Modeling #Machine Learning in Healthcare
paper · pdf · doi:10.48550/arxiv.2303.12712
Abstract
Artificial intelligence (AI) researchers have been developing and refining large language models (LLMs) that exhibit remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition. The latest model developed by OpenAI, GPT-4, was trained using an unprecedented scale of compute and data. In this paper, we report on our investigation of an early version of GPT-4, when it was still in active development by OpenAI. We contend that (this early version of) GPT-4 is part of a new cohort of LLMs (along with ChatGPT and Google's PaLM for example) that exhibit more general intelligence than previous AI models. We discuss the rising capabilities and implications of these models. We demonstrate that, beyond its mastery of language, GPT-4 can solve novel and difficult tasks that span mathematics, coding, vision, medicine, law, psychology and more, without needing any special prompting. Moreover, in all of these tasks, GPT-4's performance is strikingly close to human-level performance, and often vastly surpasses prior models such as ChatGPT. Given the breadth and depth of GPT-4's capabilities, we believe that it could reasonably be viewed as an early (yet still incomplete) version of an artificial general intelligence (AGI) system. In our exploration of GPT-4, we put special emphasis on discovering its limitations, and we discuss the challenges ahead for advancing towards deeper and more comprehensive versions of AGI, including the possible need for pursuing a new paradigm that moves beyond next-word prediction. We conclude with reflections on societal influences of the recent technological leap and future research directions.
Cited by
- Efficient Online LLM Watermark Detection via Rao-Blackwellized E-Processes
- A Unified Moral-Value Dataset for Instruction Tuning
- From Assistance to Autonomy -- A Researcher Study on the Potential of AI Support for Qualitative Data Analysis
- Breaking the Block: Preserving Data Continuity to Train Superior SAEs for Instruct Models
- GhostShell: Streaming LLM Function Calls for Concurrent Embodied Programming
- HEPTAPOD: Orchestrating High Energy Physics Workflows Towards Autonomous Agency
- Cost-efficient generative AI summarization for scalable automated essay scoring in educational assessment
- When Models Meet Users: An Empirical Study of Perceptions of General LLMs and Multimodal LLMs on Hugging Face
- The Riddle Riddle: Testing Flexible Reasoning in Large Language Models and Humans
- Faster Completion, Less Learning: Generative AI Reduced Study Time on Math Problems and the Knowledge They Build
- The Storyteller in the Model: Narrative Pattern Inheritance, Escalation Dynamics, and Alignment Governance in LLMs
- Shared sensitivity to data distribution during learning in humans and transformer networks
- What does it mean to understand language?
- Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
- Large Language Models Report Subjective Experience Under Self-Referential Processing
- Attention to Non-Adopters
- Everyone prefers human writers, including AI
- Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired Sketching
- Large Language Models Do Not Simulate Human Psychology
- Whither symbols in the era of advanced neural networks?
- Learning without training: The implicit dynamics of in-context learning
- Large language models for scholarly ontology generation: An extensive analysis in the engineering field
- Research Community Perspectives on "Intelligence" and Large Language Models
- Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis
- Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
- The Art of Audience Engagement: LLM-Based Thin-Slicing of Scientific Talks
- Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
- LLM Social Simulations Are a Promising Research Method
- How Deep Do Large Language Models Internalize Scientific Literature and Citation Practices?
- A Taxonomy of Linguistic Expressions That Contribute To Anthropomorphism of Language Technologies
- AI Sees Your Location, But With A Bias Toward The Wealthy World
- Gender disparities in the impact of generative artificial intelligence: Evidence from academia.
- Playing With AI: How Do State-Of-The-Art Large Language Models Perform in the 1977 Text-Based Adventure Game Zork?
- Towards High-Level Semantic Intelligence
- Artificial General Intelligence (AGI)-Native Wireless Systems: A Journey Beyond 6G
- A Conjecture on a Fundamental Trade-Off between Certainty and Scope in Symbolic and Generative AI
- Unifying Learning Dynamics and Generalization in Transformers Scaling Law
- PTStore (Prefix Tensor Store): Distributed Prefix Caching and Replication for High Throughput Inference Serving
- The Cartesian Cut in Agentic AI
- Coherent without Grounding, Grounded without Success: The Bidirectional Coherence Paradox in Artificial Epistemic Agents
- AegisAgent: An Autonomous Defense Agent Against Prompt Injection Attacks in LLM-HARs
- How Tech Workers Contend with Hazards of Humanlikeness in Generative AI
- Humanlike AI Design Increases Anthropomorphism but Yields Divergent Outcomes on Engagement and Trust Globally
- Can abstract concepts from LLM improve SLM performance?
- CienaLLM: Generative Climate-Impact Extraction from News Articles with Autoregressive LLMs
- External Hippocampus: Topological Cognitive Maps for Guiding Large Language Model Reasoning
- From Priors to Predictions: Explaining and Visualizing Human Reasoning in a Graph Neural Network Framework
- Plausibility as Failure: How LLMs and Humans Co-Construct Epistemic Error
- Quantifying Return on Security Controls in LLM Systems
- Large Language Newsvendor: Decision Biases and Cognitive Mechanisms
- One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs
- Large Language Models have Chain-of-Affect
- Evolutionary Reinforcement Learning based AI tutor for Socratic Interdisciplinary Instruction
- On the Dynamics of Multi-Agent LLM Communities Driven by Value Diversity
- When Medical AI Explanations Help and When They Harm
- Procrustean Bed for AI-Driven Retrosynthesis: A Unified Framework for Reproducible Evaluation
- SJD++: Improved Speculative Jacobi Decoding for Training-free Acceleration of Discrete Auto-regressive Text-to-Image Generation
- On the Computability of Artificial General Intelligence
- MCP-AI: Protocol-Driven Intelligence Framework for Autonomous Reasoning in Healthcare
- Nex-N1: Agentic Models Trained via a Unified Ecosystem for Large-Scale Environment Construction
- Catching UX Flaws in Code: Leveraging LLMs to Identify Usability Flaws at the Development Stage
- AsymPuzl: An Asymmetric Puzzle for multi-agent cooperation
- LLM-Generated Ads: From Personalization Parity to Persuasion Superiority
- ASCIIBench: Evaluating Language-Model-Based Understanding of Visually-Oriented Text
- A Human-centric Framework for Debating the Ethics of AI Consciousness Under Uncertainty
- WISE: Weighted Iterative Society-of-Experts for Robust Multimodal Multi-Agent Debate
- LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess
- Towards Active Synthetic Data Generation for Finetuning Language Models
- A Comparison of Human and ChatGPT Classification Performance on Complex Social Media Data
- Memory-Amortized Inference: A Topological Unification of Search, Closure, and Structure
- The Geometry of Certainty: Recursive Topological Condensation and the Limits of Inference
- TrafficLens: Multi-Camera Traffic Video Analysis Using LLMs
- FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
- MoodBench 1.0: An Evaluation Benchmark for Emotional Companionship Dialogue Systems
- Identifying Quantum Structure in AI Language: Evidence for Evolutionary Convergence of Human and Artificial Cognition
- Bridging Symbolic Control and Neural Reasoning in LLM Agents: The Structured Cognitive Loop
- The Impact of Quantization on Large Reasoning Model Reinforcement Learning
- Automatic Pruning Discovery for Large Language Models
- Realist and Pluralist Conceptions of Intelligence and Their Implications on AI Research
- Agent READMEs: An Empirical Study of Context Files for Agentic Coding
- Structured Decomposition for LLM Reasoning: Cross-Domain Validation and Semantic Web Integration
- LAET: A Layer-wise Adaptive Ensemble Tuning Framework for Pretrained Language Models
- On the Measure of a Model: From Intelligence to Generality
- Generative AI as a Linguistic Equalizer in Global Science
- Spontaneous eye movements reflect the representational geometries of conceptual spaces
- Place Matters: Comparing LLM Hallucination Rates for Place-Based Legal Queries
- Integrating large language models into EFL writing instruction: effects on performance, self-regulated learning strategies, and motivation
- A Collaborative Reasoning Framework for Anomaly Diagnostics in Underwater Robotics
- AthenaBench: A Dynamic Benchmark for Evaluating LLMs in Cyber Threat Intelligence
- DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning
- The Pervasive Blind Spot: Benchmarking VLM Inference Risks on Everyday Personal Videos
- FP8-Flow-MoE: A Casting-Free FP8 Recipe without Double Quantization Error
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
- TSVer: A Benchmark for Fact Verification Against Time-Series Evidence
- TempoBench: Evaluating Temporal Causal Reasoning in Large Language Models
- Budgeted Multiple-Expert Deferral
- 1+1>2: A Synergistic Sparse and Low-Rank Compression Method for Large Language Models
- Artificial Intelligence in Elementary STEM Education: A Systematic Review of Current Applications and Future Challenges
- TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
- TextualVerifier: Verify TextGrad Step-by-Step
- Testing theory of mind in large language models and humans
- Reclaiming AI as a theoretical tool for cognitive science
- MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning
- Large models of what? Mistaking engineering achievements for human linguistic agency
- Artificial Intelligence in Health Professions Education assessment: AMEE Guide No. 178
- Large Language Models and the Future of Organization Theory
- The TESCREAL bundle: Eugenics and the promise of utopia through artificial general intelligence
- Which Humans?
- Pixels and Predictions: Potential of GPT-4V in Meteorological Imagery Analysis and Forecast Communication
- How close is AI to human-level intelligence?
- Mapping the Mind With Free Associations: A Tutorial Using the R Package associatoR
- The cognitive biases that may exacerbate inflationary and deflationary positions about large language models
- Can we Trust Chatbots for now? Accuracy, reproducibility, traceability; a Case Study on Leonardo da Vinci's Contribution to Astronomy
- The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science
- The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem
- Bridging the Divide: End-to-End Sequence-Graph Learning
- Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants
- Education Paradigm Shift To Maintain Human Competitive Advantage Over AI
- Will Humanity Be Rendered Obsolete by AI?
- RaCoT: Plug-and-Play Contrastive Example Generation Mechanism for Enhanced LLM Reasoning Reliability
- Frustratingly Easy Task-aware Pruning for Large Language Models
- Learning "Partner-Aware" Collaborators in Multi-Party Collaboration
- Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents: Pathways and Paradigms
- Towards Scalable Oversight with Collaborative Multi-Agent Debate in Error Detection
- Think Parallax: Solving Multi-Hop Problems via Multi-View Knowledge-Graph-Based Retrieval-Augmented Generation
- LLMartini: Seamless and Interactive Leveraging of Multiple LLMs through Comparison and Composition
- A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
- Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge
- Presenting Large Language Models as Companions Affects What Mental Capacities People Attribute to Them
- AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM
- LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet Arena
- Pedoman Peri-operatif Renal Assessment
- The Spark Effect: On Engineering Creative Diversity in Multi-Agent AI Systems
- Assessing Coherency and Consistency of Code Execution Reasoning by Large Language Models
- Demystifying the Mechanisms Behind Emergent Exploration in Goal-conditioned RL
- Static Sandboxes Are Inadequate: Modeling Societal Complexity Requires Open-Ended Co-Evolution in LLM-Based Multi-Agent Simulations
- Doing Things with Words: Rethinking Theory of Mind Simulation in Large Language Models
- Interpreting the Latent Structure of Operator Precedence in Language Models
- A Survey on Evaluation of Large Language Models
- PADME: Procedure Aware DynaMic Execution
- Automating Structural Engineering Workflows with Large Language Model Agents
- BanglaMATH : A Bangla benchmark dataset for testing LLM mathematical reasoning at grades 6, 7, and 8
- HyperAgent: Leveraging Hypergraphs for Topology Optimization in Multi-Agent Communication
- MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation
- A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis
- Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey
- On the Role of Domain Experts in Creating Effective Tutoring Systems
- Hidden Secrets in the arXiv: Discovering, Analyzing, and Preventing Unintentional Information Disclosure in Source Files of Scientific Preprints
- CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMs
- Improving AGI Evaluation: A Data Science Perspective
- DICE: Structured Reasoning in LLMs through SLM-Guided Chain-of-Thought Correction
- TinyGraphEstimator: Adapting Lightweight Language Models for Graph Structure Inference
- Opponent Shaping in LLM Agents
- Assessing the nature of large language models: A caution against anthropocentrism
- In-Context Clustering with Large Language Models
- Lemma Dilemma: On Lemma Generation Without Domain- or Language-Specific Training Data
- Large language models show human-like content biases in transmission chain experiments
- Exploring regional vulnerability to the Fourth Industrial Revolution: a European perspective
- Foundations of LLM Knowledge Materialization: Termination, Reproducibility, Robustness
- Aligning Large Language Models via Fully Self-Synthetic Data
- Iterative LLM-Based Generation and Refinement of Distracting Conditions in Math Word Problems
- On the Role of Difficult Prompts in Self-Play Preference Optimization
- More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models
- GenQuest: An LLM-based Text Adventure Game for Language Learners
- MetaMuse: Algorithm Generation via Creative Ideation
- AgenticRAG: Tool-Augmented Foundation Models for Zero-Shot Explainable Recommender Systems
- Homophily-induced Emergence of Biased Structures in LLM-based Multi-Agent AI Systems
- Batch-CAM: Introduction to better reasoning in convolutional deep learning models
- An empirical investigation of the impact of ChatGPT on creativity
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- UniAPL: A Unified Adversarial Preference Learning Framework for Instruct-Following
- Hallucination is Inevitable for LLMs with the Open World Assumption
- LatentEvolve: Self-Evolving Test-Time Scaling in Latent Space
- Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMs
- Enabling Physical AI through Biological Principles
- Effectiveness of Large Language Models in Simulating Regional Psychological Structures: An Empirical Examination of Personality and Subjective Well-being
- Watermarking Diffusion Language Models
- The future of academic publishing
- A Theoretical Computer Science Perspective on Consciousness and Artificial General Intelligence
- Towards Human-interpretable Explanation in Code Clone Detection using LLM-based Post Hoc Explainer
- GPT (Generative Pre-Trained Transformer)— A Comprehensive Review on Enabling Technologies, Potential Applications, Emerging Challenges, and Future Directions
- Linear Causal Representation Learning by Topological Ordering, Pruning, and Disentanglement
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- “It happened to be the perfect thing”: experiences of generative AI chatbots for mental health
- Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect Personas
- Can LLMs Forecast Internet Traffic from Social Media?
- The future of machine learning for small-molecule drug discovery will be driven by data
- When and how to disclose AI use in academic publishing: AMEE Guide No.192
- Generative AI for Economic Research: Use Cases and Implications for Economists
- AI Gamestore: Scalable, Open-Ended Evaluation of Machine General Intelligence with Human Games
- A large-scale evaluation of commonsense knowledge in humans and large language models
- Patterns Lead the Way to Far-from-Equilibrium Materials
- Generative artificial intelligence and engineering education
- The Brain Abstracted
- Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models
- Through the Lens of Human-Human Collaboration: A Configurable Research Platform for Exploring Human-Agent Collaboration
- nDNA -- the Semantic Helix of Artificial Cognition
- Small LLMs with Expert Blocks Are Good Enough for Hyperparamter Tuning
- Learning in Context: Personalizing Educational Content with Large Language Models to Enhance Student Learning
- Charting trajectories of human thought using large language models
- Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
- AI For Privacy in Smart Homes: Exploring How Leveraging AI-Powered Smart Devices Enhances Privacy Protection
- CultranAI at PalmX 2025: Data Augmentation for Cultural Knowledge Representation
- Artificial Intelligence and Entrepreneurship: A Call for Research to Prospect and Establish the Scholarly AI Frontiers
- Genome-wide prediction of disease variant effects with a deep protein language model
- Human resource management in the age of generative artificial intelligence: Perspectives and research directions on ChatGPT
- Cycle is All You Need: More Is Different
- Co-Alignment: Rethinking Alignment as Bidirectional Human-AI Cognitive Adaptation
- Data-Driven Analysis of Text-Conditioned AI-Generated Music: A Case Study with Suno and Udio
- MAPGD: Multi-Agent Prompt Gradient Descent for Collaborative Prompt Optimization
- Compartmentalised Agentic Reasoning for Clinical NLI
- Towards Fully Automated Molecular Simulations: Multi-Agent Framework for Simulation Setup and Force Field Extraction
- AI Wellbeing
- Exploring the Impact of Generative Artificial Intelligence on Software Development in the IT Sector: Preliminary Findings on Productivity, Efficiency and Job Security
- Investigating Student Interaction Patterns with Large Language Model-Powered Course Assistants in Computer Science Courses
- Mitigating Catastrophic Forgetting in Large Language Models with Forgetting-aware Pruning
- Language Self-Play For Data-Free Training
- Disentangling Interaction and Bias Effects in Opinion Dynamics of Large Language Models
- Code2MCP: Transforming Code Repositories into MCP Services
- The human biological advantage over AI
- The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
- DiaCBT: A Long-Periodic Dialogue Corpus Guided by Cognitive Conceptualization Diagram for CBT-based Psychological Counseling
- Assessing Consciousness-Related Behaviors in Large Language Models Using the Maze Test
- On the Alignment of Large Language Models with Global Human Opinion
- Analysis of Error Sources in LLM-based Hypothesis Search for Few-Shot Rule Induction
- Transforming Agency. On the mode of existence of Large Language Models
- CVPD at QIAS 2025 Shared Task: An Efficient Encoder-Based Approach for Islamic Inheritance Reasoning
- Going over Fine Web with a Fine-Tooth Comb: Technical Report of Indexing Fine Web for Problematic Content Search and Retrieval
- Personality Matters: User Traits Predict LLM Preferences in Multi-Turn Collaborative Tasks
- Integrating Large Language Models with Network Optimization for Interactive and Explainable Supply Chain Planning: A Real-World Case Study
- Pruning Strategies for Backdoor Defense in LLMs
- Secure Multi-LLM Agentic AI and Agentification for Edge General Intelligence by Zero-Trust: A Survey
- Emotion Transfer with Enhanced Prototype for Unseen Emotion Recognition in Conversation
- MQAD: A Large-Scale Question Answering Dataset for Training Music Large Language Models
- Grounding the Ungrounded: A Spectral-Graph Framework for Quantifying Hallucinations in Multimodal LLMs
- APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration
- A Concurrent Modular Agent: Framework for Autonomous LLM Agents
- S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner
- ELATE: Evolutionary Language model for Automated Time-series Engineering
- Cohort-Aware Agents for Individualized Lung Cancer Risk Prediction Using a Retrieval-Augmented Model Selection Framework
- Leveraging Large Language Models for Predictive Analysis of Human Misery
- Uncovering Emergent Physics Representations Learned In-Context by Large Language Models
- Structuring the Unstructured: A Systematic Review of Text-to-Structure Generation for Agentic AI with a Universal Evaluation Framework
- SupraTok: Cross-Boundary Tokenization for Enhanced Language Model Performance
- Applied causality to infer protein dynamics and kinetics
- AI That Helps Us Help Each Other: A Proactive System for Scaffolding Mentor-Novice Collaboration in Entrepreneurship Coaching
- What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles
- EvoCurr: Self-evolving Curriculum with Behavior Code Generation for Complex Decision-making
- Large Language Models Show Signs of Alignment with Human Neurocognition During Abstract Reasoning
- Dynamic Uncertainty-aware Multimodal Fusion for Outdoor Health Monitoring
- Compass-Thinker-7B Technical Report
- "Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas
- Multimodal learning with next-token prediction for large multimodal models
- Intuition emerges in Maximum Caliber models at criticality
- Between Tool and Trouble: Student Attitudes Toward AI in Programming Education
- Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support
- Panel-Scale Reconfigurable Photonic Interconnects for Scalable AI Computation
- LLMs for Resource Allocation: A Participatory Budgeting Approach to Inferring Preferences
- AGI for the Earth, the path, possibilities and how to evaluate intelligence of models that work with Earth Observation Data?
- FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and Challenging
- Why are LLMs' abilities emergent?
- The Science Fiction Science Method
- VLMQ: Efficient Post-Training Quantization for Large Vision-Language Models via Hessian Augmentation
- Autonomous Inorganic Materials Discovery via Multi-Agent Physics-Aware Scientific Reasoning
- Authorship Attribution in Multilingual Machine-Generated Texts
- How Does Controllability Emerge In Language Models During Pretraining?
- Autonomous Penetration Testing: Solving Capture-the-Flag Challenges with LLMs
- Distributed AI Agents for Cognitive Underwater Robot Autonomy
- GPT-4 [wikipedia]
- History of artificial intelligence [wikipedia]
- January–March 2023 in science [wikipedia]
- Large language model [wikipedia]
- Pause Giant AI Experiments: An Open Letter [wikipedia]
- Products and applications of OpenAI [wikipedia]
- Sébastien Bubeck [wikipedia]
- Sally–Anne test [wikipedia]
- Superintelligence [wikipedia]
- Timeline of computing 2020–present [wikipedia]
Discussions
- Sparks of Artificial General Intelligence: Early Experiments with GPT-4 [hn, 180 points, 236 comments]
- Sparks of AGI: early experiments with GPT-4 (6.04.2023 yt vid) [lemmy, 7 points, 5 comments]
- Es ist schwierig, Literatur zu AGI zu empfehlen, da bislang keine allgemein akzeptierte Definition zu AGI existiert. Einigkeit besteht nur darüber, dass rein künstliche neuronale Netze in ihrer heutig [bsky, 6 points, 1 comments]
- 2024-07-01 [lemmy, 6 points, 0 comments]
- Não há produção científica que justifique essa tese. Uns pesquisadores da Microsoft chegaram a publicar um artigo nesse sentido (arxiv.org/abs/2303.12712), mas não é convincente. [bsky, 5 points, 1 comments]
- Sparks of AGI: early experiments with GPT-4 (2 month old yt vid) [lemmy, 4 points, 0 comments]
- Sparks of Artificial General Intelligence: Early experiments with GPT-4 [lobsters, 3 points, 0 comments]
- A paper from researchers at Microsoft claims AI shows the ability to understand the way people do. This is a report from the New York Times. The paper: “Sparks of Artificial General Intelligence.” a [bsky, 3 points, 1 comments]
- It was, version one of the document also had the teeny tiny problem of using a study that was basically demographical phrenology to 'measure intelligence' (reference from the first line of the conclus [bsky, 2 points, 1 comments]
- Sparks of Artificial General Intelligence: Early Experiments with GPT-4 (2023) [hn, 2 points, 0 comments]
- I'm trying to understand the very impressive GPT-4 example from this paper: https://arxiv.org/pdf/2303.12712.pdf titled "Sparks of Artificial General Intelligence: Early experiments with GPT-4" [bsky, 1 points, 1 comments]
- It's not AGI. It has weaknesses (which are different from human weaknesses - the weaknesses of humans always get left out of this conversation, IMHO). But it also has great strengths. Good research [bsky, 1 points, 1 comments]
- Very nice documentation 👍 Do you know this extensive comparison of a few bots from 2023? Just skimming through its pictures still leads me to interpret "AI" as "Acquired Incompetence", reflecting bo [bsky, 1 points, 1 comments]
- Sparks of Artificial General Intelligence: Early Experiments with GPT-4 [bsky, 1 points, 0 comments]
- Paper -> arxiv.org/abs/2303.12712 [bsky, 1 points, 1 comments]
- KI lernen nicht, sie werden trainiert. Entweder durch Menschen oder durch "Pretrained ANN" (die wiederum von Menschen trainiert wurden). Google "GAN" & "diskriminatorische Modelle" & "generative Model [bsky, 1 points, 4 comments]
- No prob! If you're still on the fence, you can always check their arxiv preprint: https://arxiv.org/pdf/2303.12712 Thought it was pretty impressive. Also Poe also gives you 1 free query a day, iirc. [bsky, 1 points, 1 comments]
- Tests by MSR on pre-RLHF GPT4, very famous paper: arxiv.org/abs/2303.12712 Pre-RLHF is importent because RLHF dumbs down models (imo) and forces them to output a certain way. Which is why we saw goo [bsky, 0 points, 1 comments]
- AGI isn't a well-defined concept but similarly a matter of degree and what skills are being measured. "Given the breadth and depth of GPT-4’s capabilities, we believe that it could reasonably be view [bsky, 0 points, 1 comments]
- Of course this was the subject of the [Sparks of AGI](arxiv.org/pdf/2303.12712) paper. Same thing on [Claude 3.5 Sonnet](claude.site/artifacts/3f...) [bsky, 0 points, 1 comments]
- GPT-4 might be showing “sparks” of AGI according to this paper. For example it can and knows when to use tools… arxiv.org/abs/2303.12712 [bsky, 0 points, 0 comments]
Related