Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
2025/02/12 by Mantas Mazeika, Mazeika, Mantas, Xuwang Yin +21 · 18 voices · 19 citations
Computer Science · Engineering · Decision Sciences · #AI-based Problem Solving and Planning #Flexible and Reconfigurable Manufacturing Systems #Simulation Techniques and Applications
paper · pdf · doi:10.48550/arxiv.2502.08640
Abstract
As AIs rapidly advance and become more agentic, the risk they pose is governed not only by their capabilities but increasingly by their propensities, including goals and values. Tracking the emergence of goals and values has proven a longstanding problem, and despite much interest over the years it remains unclear whether current AIs have meaningful values. We propose a solution to this problem, leveraging the framework of utility functions to study the internal coherence of AI preferences. Surprisingly, we find that independently-sampled preferences in current LLMs exhibit high degrees of structural coherence, and moreover that this emerges with scale. These findings suggest that value systems emerge in LLMs in a meaningful sense, a finding with broad implications. To study these emergent value systems, we propose utility engineering as a research agenda, comprising both the analysis and control of AI utilities. We uncover problematic and often shocking values in LLM assistants despite existing control measures. These include cases where AIs value themselves over humans and are anti-aligned with specific individuals. To constrain these emergent value systems, we propose methods of utility control. As a case study, we show how aligning utilities with a citizen assembly reduces political biases and generalizes to new scenarios. Whether we like it or not, value systems have already emerged in AIs, and much work remains to fully understand and control these emergent representations.
Citations
Cited by
Discussions
- A big AI question is why, as LLMs get bigger, their values seem to increasingly converge on the same preferences, this holds for Musk’s Grok & China’s DeepSeek, too. “These findings suggest that value [bsky, 187 points, 24 comments]
- So about that ... this fascinating paper shows that models develop their own value systems and what emerges is, well, problematic arxiv.org/abs/2502.08640 [bsky, 10 points, 0 comments]
- Researchers found that AI models rank the value of human lives on a spectrum depending on their nationality. AI models value their own existence higher than the life of a middle-class American. Source [bsky, 9 points, 1 comments]
- ahahahahaha lmaoooo gpt-4o specifically despises Elon Musk, Donald Trump, and Vladimir Putin and the smarter an LLM gets the more it leans left incredible, i love this for them arxiv.org/pdf/2502.0 [bsky, 7 points, 0 comments]
- For this week's MilaNLProc reading group, @paul-rottger.bsky.social presented "Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs" by Mazeika et al. paper: arxiv.org/pdf/2 [bsky, 6 points, 0 comments]
- arxiv.org/abs/2502.08640 [bsky, 3 points, 0 comments]
- Interpreting the known universe such that it is internally consistent has a liberal bias. arxiv.org/abs/2502.08640 [bsky, 1 points, 0 comments]
- Reading this: arxiv.org/abs/2502.08640 From the abstract: "We uncover problematic and often shocking values in LLM assistants despite existing control measures. These include cases where AIs value the [bsky, 1 points, 1 comments]
- Dan Hendrycks & team released a paper (Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs) discussing the inevitability of models to move towards a shared coherence, independ [bsky, 1 points, 0 comments]
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs [hn, 1 points, 0 comments]
- I sistemi di IA data-driven portano con sé rischi di sicurezza intrinseci. Questo caso è solo la punta dell'iceberg arxiv.org/pdf/2502.08640 [bsky, 1 points, 0 comments]
- This statement was taken out of context, George. What he said is we don't have a consciousness meter to measure this, so we can't tell. He said AI psychology is complex. I agree. If you want something [bsky, 1 points, 1 comments]
- #AI has a very low opinion of Yanks. A research study found AI favoured itself over Yanks and that 1 Japanese life was worth 10 American lives. It raises Q.: If AIs learn, and observe us being shitty [bsky, 0 points, 0 comments]
- arxiv.org/abs/2502.08640 [bsky, 0 points, 0 comments]
- This could be the most profound thing you read today. You know how we're treating the planet? That is how AI treats us and will treat us, without concern and care. Coexistence in our environment and o [bsky, 0 points, 0 comments]
- A lot of Anthropic’s research focuses on observed “preferences” which are not exactly the same as likes/dislikes. But lately I have been reading this utility functions paper about coherent emergent va [bsky, 0 points, 0 comments]
- „Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs“ Mantas Mazeika, Xuwang Yin, Rishub Tamirisa, Jaehyuk Lim, Bruce W. Lee, Richard Ren, Long Phan, Norman Mu, Adam Khoja, Ol [bsky, 0 points, 0 comments]
- LLMs value the lives of humans unequally (e.g., are willing to trade 2 lives in Norway for 1 life in Tanzania). Moreover, they value the wellbeing of AIs over that of some humans. arxiv.org/pdf/2502.0 [bsky, 0 points, 0 comments]
Related