Computational Economics in Large Language Models: Exploring Model Behavior and Incentive Design under Resource Constraints
2025/08/14 by Reddy, Sandeep, Khan, Kabir, Patil, Rohit +5
#Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.6 #I.2.7 #I.5.1
paper · doi:10.48550/arxiv.2508.10426
Abstract
Large language models (LLMs) are limited by substantial computational cost. We introduce a "computational economics" framework that treats an LLM as an internal economy of resource-constrained agents (attention heads and neuron blocks) that must allocate scarce computation to maximize task utility. First, we show empirically that when computation is scarce, standard LLMs reallocate attention toward high-value tokens while preserving accuracy. Building on this observation, we propose an incentive-driven training paradigm that augments the task loss with a differentiable computation cost term, encouraging sparse and efficient activations. On GLUE (MNLI, STS-B, CoLA) and WikiText-103, the method yields a family of models that trace a Pareto frontier and consistently dominate post-hoc pruning; for a similar accuracy we obtain roughly a forty percent reduction in FLOPS and lower latency, together with more interpretable attention patterns. These results indicate that economic principles offer a principled route to designing efficient, adaptive, and more transparent LLMs under strict resource constraints.
Citations
- Learning from Heterogeneity: Generalizing Dynamic Facial Expression Recognition via Distributionally Robust Optimization
- Unified Static and Dynamic Network: Efficient Temporal Filtering for Video Grounding
- WiOpen: A Robust Wi-Fi-based Open-set Gesture Recognition Framework
- Mixtral of Experts
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- LongNet: Scaling Transformers to 1,000,000,000 Tokens
- ReSup: Reliable Label Noise Suppression for Facial Expression Recognition
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Self-collaboration Code Generation via ChatGPT
- Self-Collaboration Code Generation via ChatGPT
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Distributing Accountability, Not Capability: Phase Separation and the LLM Workflow Quadrant in Autonomous AI Agent Architectures
- Emergent Abilities of Large Language Models
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Do Vision Transformers See Like Convolutional Neural Networks?
- On the Dangers of Stochastic Parrots
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- Transformer Feed-Forward Layers Are Key-Value Memories
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- The Hardware Lottery
- Language Models are Few-Shot Learners
- DeeBERT: Dynamic Early Exiting for Accelerating BERT Inference
- Longformer: The Long-Document Transformer
- BERT-of-Theseus: Compressing BERT by Progressive Module Replacing
- Scaling Laws for Neural Language Models
- Reformer: The Efficient Transformer
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- TinyBERT: Distilling BERT for Natural Language Understanding
- Probing Neural Network Comprehension of Natural Language Arguments
- Attention is not Explanation
- GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Wide Residual Networks
- Adaptive Computation Time for Recurrent Neural Networks
- The information bottleneck method
Cited by
Related