GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models
2023/03/17 by Tyna Eloundou, Sam Manning, Eloundou, Tyna +6 · 18 voices · 545 citations
Business, Management and Accounting · Computer Science · Economics, Econometrics and Finance · Medicine · #Artificial Intelligence in Healthcare and Education #Development economics #Economic growth #Economics #FinTech, Crowdfunding, Digital Finance #Geography #Labour economics #Rubric #Sociology #Timeline #Topic Modeling #Wage #Workforce #cs.AI #cs.CY #econ.GN
paper · pdf · doi:10.48550/arxiv.2303.10130
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/03/17 · openalex created_date 2023/03/22 · openalex updated_date 2026/08/02
Abstract
Public estimates of AI’s labor-market exposure typically come from one of two sources: a theoretical judgment about what a model could do, or a record of what people have actually asked it to do. This paper tests whether the second source has a specific, measurable blind spot – that usage-based exposure measures understate AI’s capability for occupations whose tasks have not yet entered common chat usage – using an independently constructed, task-decomposition- based capability judgment, the TRIPS Framework by Trust Insights (which scores individual job tasks for AI suitability), compared directly against Anthropic’s published Economic Index data. Across 434 O*NET-SOC occupations, TRIPS’s coverage share – the estimated share of a job’s tasks current AI can complete independently – correlates strongly with Eloundou et al.’s (2023) capability rating (Spearman’s rho = 0.747), a concordance an independent human- rated column from the same source closely reproduces. Restricted to the 286 occupations where Anthropic’s usage-volume data exists, TRIPS still correlates strongly with Eloundou et al.’s rating (rho = 0.718) but only weakly with Anthropic’s actual usage volume (rho = 0.254): two independently built capability judgments track each other far more closely than either tracks real usage, and both correlations survive a Monte Carlo check simulating realistic classifier label noise. This is directionally consistent with a usage-gating mechanism, corroborated by Anthropic’s own report naming occupations that register zero measured exposure purely for lack of chat traffic – though it cannot, on cross-sectional evidence alone, be distinguished from a slower adoption-lag explanation. Category-level coverage varies substantially (9.4 to 81.7 percent across 22 SOC major groups), supporting neither the claim that any category is immune to AI nor that any nears full automation. We report these findings as directionally consistent and noise-surviving, not confirmed, with construct, sampling, and classification-accuracy limitations detailed throughout.
Cited by
- Anarchist Automation: A Sociotechnical Framework for Decentralization and Universal Care
- LLMs Corrupt Your Documents When You Delegate
- Deep Hype in Artificial General Intelligence: Uncertainty, Sociotechnical Fictions and the Governance of AI Futures
- Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the U.S. Workforce
- Large Language Models Pass the Turing Test
- A matter of principle? AI alignment as the fair treatment of claims
- Who Prices Cognitive Labor in the Age of Agents? Compute-Anchored Wages
- TrajSyn: Privacy-Preserving Dataset Distillation from Federated Model Trajectories for Server-Side Adversarial Training
- How Well Does Agent Development Reflect Real-World Work?
- How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
- A theory-based AI automation exposure index: Applying Moravec's Paradox to the US labor market
- Reproducibility: The New Frontier in AI Governance
- Humanoid Artificial Consciousness Designed with Large Language Model Based on Psychoanalysis and Personality Theory
- Exploring regional vulnerability to the Fourth Industrial Revolution: a European perspective
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
- PromptPilot: Improving Human-AI Collaboration Through LLM-Enhanced Prompt Engineering
- Artificial intelligence and work design: implications for frontline service employees and future research
- The Impact of AI Adoption on Retail Across Countries and Industries
- Incentives for Digital Twins: Task-Based Productivity Enhancements with Generative AI
- Making AI Inevitable: Historical Perspective and the Problems of Predicting Long-Term Technological Change
- The Quasi-Creature and the Uncanny Valley of Agency: A Synthesis of Theory and Evidence on User Interaction with Inconsistent Generative AI
- PB-IAD: Utilizing multimodal foundation models for semantic industrial anomaly detection in dynamic manufacturing environments
- Ask ChatGPT: Caveats and Mitigations for Individual Users of AI Chatbots
- Idempotent Equilibrium Analysis of Hybrid Workflow Allocation: A Mathematical Schema for Future Work
- Agentic AI and Occupational Displacement: A Multi-Regional Task Exposure Analysis of Emerging Labor Market Disruption
- VoyagerVision: Investigating the Role of Multi-modal Information for Open-ended Learning Systems
- NLPnorth @ TalentCLEF 2025: Comparing Discriminative, Contrastive, and Prompt-Based Methods for Job Title and Skill Matching
- Superstudent intelligence in thermodynamics
- Diverging paths: AI exposure and employment across European regions
- A Mathematical Framework for AI-Human Integration in Work
- Competition between AI foundation models: dynamics and policy recommendations
- Can AI Freelancers Compete? Benchmarking Earnings, Reliability, and Task Success at Scale
- 14 examples of how LLMs can transform materials science and chemistry: a reflection on a large language model hackathon
- KI og effektivisering: Hvordan GPT-4 fant 155.000 overflødige årsverk i norsk offentlig sektor
- Vibe Researching as Wolf Coming: Can AI Agents with Skills Replace or Augment Social Scientists?
- From Reflection to Repair: A Scoping Review of Dataset Documentation Tools
- Boom, Bubble, or Buildout? A Multi-Method Evaluation of Whether Artificial Intelligence Is in an Ongoing Financial Bubble
- The AI Skills Shift: Mapping Skill Obsolescence, Emergence, and Transition Pathways in the LLM Era
- IberBench: LLM Evaluation on Iberian Languages
- AI-Augmented Design Thinking: Potentials, Challenges, and Mitigation Strategies of Integrating Artificial Intelligence in Human-Centered Innovation Processes
- AI Safety Should Prioritize the Future of Work
- AI as a resource for the clarification of medical terminology
Discussions
- GPTs Are GPTs: An Early Look at the Labor Market Impact Potential of LLMs [hn, 190 points, 230 comments]
- I would also like to remind folks that OpenAI wrote a paper in which they prompted GPT-4 on which jobs they thought would be most exposed to automation. They validated it by comparing it to responses [bsky, 126 points, 6 comments]
- The influence of AI on professions. GPTs is a play on words. GPT means 'Generative Pre-trained Transformers' or 'General-Purpose Technologies', i.e. cross-sectional technologies. "GPTs are GPTs: An E [bsky, 3 points, 0 comments]
- An early look at the labor market impact potential of LLMs (2023) [hn, 2 points, 0 comments]
- já tem paper bom sobre isso https://arxiv.org/pdf/2303.10130.pdf as profissões NÃO expostas a impactos de LLM como o chatgpt seriam essas aqui @lapapaespop.bsky.social [bsky, 2 points, 0 comments]
- A new LLM is like a new source of steel: Fascinating scientific discovery. But nothing that really affects anyone not using raw materials. arxiv.org/abs/2303.10130 predicts 80% of jobs will have 10% [bsky, 2 points, 0 comments]
- This paper is a little old 😅 but its methodology is still relevant for assessing AI's labor market impact. A must-read for understanding how LLMs intersect with the workforce. arxiv.org/pdf/2303.1013 [bsky, 1 points, 0 comments]
- Sorry I was not clear. This paper is by Anthropic but the calculation of potential job market use values (blue area) comes from this paper published by open AI, as cited in the appendix. arxiv.org/pdf [bsky, 1 points, 1 comments]
- I believe they used the ideas from this paper arxiv.org/pdf/2303.10130 [bsky, 1 points, 0 comments]
- Ready to see how GPTs are shaking up the labor market? Check out this paper for an early peek at the potential of large language models. Link: https://arxiv.org/abs/2303.10130 #LLM #ChatGPT #Labor [bsky, 1 points, 0 comments]
- 例えば AI による翻訳はまともに使えるようになってきたが、まだ翻訳者が求めるレベルではなく、翻訳の手直しが必要で人間の介在が残る。しかし需給の総量は変わらないので、初期的な翻訳が自動化され、翻訳者は膨大に残る最後の手直しに追われ、結果的に賃金が押し下げられる構造。翻訳は GPTs are GPTs (arxiv.org/abs/2303.10130) で言うところの exposed な仕事であり [bsky, 0 points, 1 comments]
- https://arxiv.org/pdf/2303.10130.pdf?fbclid=PAAabftYt1RJZOilZWZSZU1obm6rBe6FVsoNv8V2v615H119RTYuNt3wF_LLk #generativeAI #GPT #AI #FutureofWork [bsky, 0 points, 0 comments]
- Interesting paper looking into the effect of LLMs on the US labor market 15% of all worker tasks in the US could be completed significantly faster at the same level of quality. 80% could see 10% of [bsky, 0 points, 1 comments]
- Daniel has made a career studying your first statement [bsky, 0 points, 1 comments]
- According to this research paper by @OpenAI () the following jobs have the least risk of being affected or replaced by LLMs like GPT-4. What are you going to be doing in 5 years time? I’m considerin [bsky, 0 points, 0 comments]
- GPTs Are GPTs: An Early Look at the Labor Market Impact Potential of LLMs [bsky, 0 points, 0 comments]
- #GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models #LLM #GenAI #TechPakistan #AI #ML arxiv.org/pdf/2303.10130 [bsky, 0 points, 0 comments]
- “findings indicate that approximately 80% of the U.S. workforce could have at least 10% of their work tasks affected by the introduction of GPTs, while around 19% of workers may see at least 50% of th [bsky, 0 points, 0 comments]
Related