2026/06/06 by Shuyao Gao, Minghao Huang
#cs.CY
Research on artificial intelligence and work assigns each occupation a single exposure score. We build an instrument to see what those scores average over: a decomposition of 1,961 O*NET work activities into 15,817 atomic micro-actions by a consensus multi-agent LLM pipeline, clustered from text alone into seven semantic classes. Projecting exposure indicators onto these classes reveals two extreme poles, tool-mediated physical execution and planning-and-design, separated by a gap far larger than random partitions of the same data produce (permutation P < 10-4; Cliff's δ= 0.80 under our tech-risk index and 0.90 under GPT-4 task ratings). The poles flank a broad central band that carries most work and is only weakly more compressed than chance. The poles are stable across clustering resolution, sentence encoder (under a common partition), and indicator, yet which pole is most exposed has inverted since 2013: the two extremes swap identity between the Frey-Osborne computerisation era and the LLM era, and at the occupation level an occupation's 2013 automatability declines as its linguistic content rises (ρ= -0.40, n = 618). We release the instrument and its outputs. The durable object for forecasting is the structure of work itself, not any era's exposure ranking.