vix.ing · top · new · best · stats

GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models

2023/03/17 by Tyna Eloundou, Sam Manning, Eloundou, Tyna +6 · 18 voices · 545 citations
Business, Management and Accounting · Computer Science · Economics, Econometrics and Finance · Medicine · #Artificial Intelligence in Healthcare and Education #Development economics #Economic growth #Economics #FinTech, Crowdfunding, Digital Finance #Geography #Labour economics #Rubric #Sociology #Timeline #Topic Modeling #Wage #Workforce #cs.AI #cs.CY #econ.GN

paper · pdf · doi:10.48550/arxiv.2303.10130

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2023/03/17 · openalex created_date 2023/03/22 · openalex updated_date 2026/08/02

Abstract

Public estimates of AI’s labor-market exposure typically come from one of two sources: a theoretical judgment about what a model could do, or a record of what people have actually asked it to do. This paper tests whether the second source has a specific, measurable blind spot – that usage-based exposure measures understate AI’s capability for occupations whose tasks have not yet entered common chat usage – using an independently constructed, task-decomposition- based capability judgment, the TRIPS Framework by Trust Insights (which scores individual job tasks for AI suitability), compared directly against Anthropic’s published Economic Index data. Across 434 O*NET-SOC occupations, TRIPS’s coverage share – the estimated share of a job’s tasks current AI can complete independently – correlates strongly with Eloundou et al.’s (2023) capability rating (Spearman’s rho = 0.747), a concordance an independent human- rated column from the same source closely reproduces. Restricted to the 286 occupations where Anthropic’s usage-volume data exists, TRIPS still correlates strongly with Eloundou et al.’s rating (rho = 0.718) but only weakly with Anthropic’s actual usage volume (rho = 0.254): two independently built capability judgments track each other far more closely than either tracks real usage, and both correlations survive a Monte Carlo check simulating realistic classifier label noise. This is directionally consistent with a usage-gating mechanism, corroborated by Anthropic’s own report naming occupations that register zero measured exposure purely for lack of chat traffic – though it cannot, on cross-sectional evidence alone, be distinguished from a slower adoption-lag explanation. Category-level coverage varies substantially (9.4 to 81.7 percent across 22 SOC major groups), supporting neither the claim that any category is immune to AI nor that any nears full automation. We report these findings as directionally consistent and noise-surviving, not confirmed, with construct, sampling, and classification-accuracy limitations detailed throughout.

Cited by

Discussions

Related