vix.ing · top · new · best · stats · spec

Working with AI: Measuring the Applicability of Generative AI to Occupations

Analyzing 200k real Bing Copilot chats, researchers found AI is used most for information work, ranking occupations by AI applicability.

2025/07/10 by Kiran Tomlinson, Sonia Jaffe, Tomlinson, Kiran +8 · 110 voices · 10 citations
Social Sciences · Computer Science · Medicine · #Ethics and Social Impacts of AI #AI in Service Interactions #Artificial Intelligence in Healthcare and Education

paper · pdf · doi:10.48550/arxiv.2507.07935

Abstract

With generative AI emerging as a general-purpose technology, understanding its economic effects is among society's most pressing questions. Existing studies of AI impact have largely relied on predictions of AI capabilities or focused narrowly on individual firms. Drawing instead on real-world AI usage, we analyze a dataset of 200k anonymized conversations with Microsoft Bing Copilot to measure AI applicability to occupations. We use an LLM-based pipeline to classify the O*NET work activities assisted or performed by AI in each conversation. We find that the most common and successful AI-assisted work activities involve information work--the creation, processing, and communication of information. At the occupation level, we find widespread AI applicability cutting across sectors, as most occupations have information work components. Our methodology also allows us to predict which occupations are more likely to delegate tasks to AI and which are more likely to use AI to assist existing workflows.

Summary

The authors analyzed 200,000 anonymized Microsoft Bing Copilot conversations, using an LLM pipeline to classify what the user wanted and what the AI actually did in O*NET work-activity terms. Most usage and success clusters around information work such as writing, explaining, and answering questions, and this translates into an occupation-level 'AI applicability score' that ranks media, sales, and clerical jobs highest and manual, physical jobs lowest.

machine-generated · claude-sonnet-5

Outline

machine-generated · claude-sonnet-5

Claims

machine-generated · claude-sonnet-5

Key figure

Figure 1 — Four charts showing which work activities came up most in Bing Copilot chats: the 15 most common things users wanted help with, the 15 most common things the AI actually did, how these compare to how often those activities happen across the whole U.S. workforce, and which activities the AI handled most and least successfully.

machine-generated · claude-sonnet-5

Glossary

O*NET
A U.S. Department of Labor database that breaks occupations down into their component tasks and work activities.
Intermediate Work Activity (IWA)
A mid-level, cross-occupation description of a job task from O*NET, e.g. 'Edit written materials or documents.'
Generalized Work Activity (GWA)
A broader O*NET grouping of related IWAs (37 total) used to compare Copilot usage against overall workforce activity.
User goal
What the human in a Copilot conversation was trying to accomplish, as classified by the paper's LLM pipeline.
AI action
What the AI itself did in the conversation, classified separately from the user's stated goal.
Scope (of impact)
A six-point rating (none to complete) of how much of a work activity's total work Copilot demonstrated it could handle.
Completion rate
The share of conversations in which an LLM classifier judged that Copilot successfully completed the user's request.
AI applicability score
The paper's composite per-occupation metric combining activity coverage, completion rate, and scope, weighted by how important each activity is to that occupation.
Standard Occupational Classification (SOC)
The U.S. government's standard system for coding and grouping detailed occupations, used here to report results by job and job group.
Activity share
The fraction of all conversations attributed to a given work activity, splitting credit evenly when a conversation matches multiple activities.

machine-generated · claude-sonnet-5

Audience

Labor economists, workforce and AI policy researchers, and product or policy teams at AI companies who need evidence-based (not forecast-based) estimates of which jobs and tasks generative AI already affects.

prerequisites: Basic familiarity with occupational task taxonomies like O*NET or SOC, Comfort reading correlation coefficients and PCA-based scatterplots, General understanding of how an LLM classification pipeline works

machine-generated · claude-sonnet-5

Open questions

machine-generated · claude-sonnet-5

Supplementary links

machine-generated · claude-sonnet-5

Citations

Cited by

Discussions

Related