Unsupervised Elicitation of Language Models
2025/06/11 by Jiaxin Wen, Wen, Jiaxin, Zachary Ankner +23 · 15 voices · 4 citations
Computer Science · #Topic Modeling #Multimodal Machine Learning Applications #Machine Learning and Data Classification
paper · pdf · doi:10.48550/arxiv.2506.10139
Abstract
To steer pretrained language models for downstream tasks, today's post-training paradigm relies on humans to specify desired behaviors. However, for models with superhuman capabilities, it is difficult or impossible to get high-quality human supervision. To address this challenge, we introduce a new unsupervised algorithm, Internal Coherence Maximization (ICM), to fine-tune pretrained language models on their own generated labels, without external supervision. On GSM8k-verification, TruthfulQA, and Alpaca reward modeling tasks, our method matches the performance of training on golden labels and outperforms training on crowdsourced human supervision. On tasks where LMs' capabilities are strongly superhuman, our method can elicit those capabilities significantly better than training on human labels. Finally, we show that our method can improve the training of frontier LMs: we use our method to train an unsupervised reward model and use reinforcement learning to train a Claude 4 Sonnet-based assistant. The resulting assistant matches its counterpart trained on production-grade human labels on average, with higher scores on chat and safety yet lower scores on math and coding.
Citations
Cited by
Discussions
- Unsupervised Elicitation of Language Models [hn, 135 points, 24 comments]
- Unsupervised Elicitation of Language Models [hn, 7 points, 0 comments]
- Unsupervised Elicitation of Language Models https://arxiv.org/abs/2506.10139 (https://news.ycombinator.com/item?id=44276041) [bsky, 0 points, 0 comments]
- Oops forgot the links... 2025 Anthropic nerd-snipe troll(?): arxiv.org/abs/2506.101... 2022 Meta background reading: arxiv.org/abs/2212.09689 [bsky, 0 points, 1 comments]
- Unsupervised Elicitation of Language Models [bsky, 0 points, 0 comments]
- Unsupervised Elicitation of Language Models #HackerNews https://arxiv.org/abs/2506.10139 [bsky, 0 points, 0 comments]
- Unsupervised Elicitation of Language Models https://arxiv.org/abs/2506.10139 https://news.ycombinator.com/item?id=44276041 [bsky, 0 points, 0 comments]
- Unsupervised Elicitation of Language Models https://arxiv.org/abs/2506.10139 [bsky, 0 points, 0 comments]
- Related, super recent also arxiv.org/abs/2506.101... [bsky, 0 points, 0 comments]
- Unsupervised Elicitation of Language Models https://arxiv.org/abs/2506.10139 [comments] [117 points] [bsky, 0 points, 0 comments]
- ⚡ Hackernews Top story: Unsupervised Elicitation of Language Models [bsky, 0 points, 0 comments]
- "Unsupervised Elicitation of Language Models" arxiv.org/abs/2506.10139 [bsky, 0 points, 0 comments]
- Unsupervised Elicitation of Language Models view on hacker news [bsky, 0 points, 0 comments]
- Unsupervised Elicitation of Language Models View Article | Join the HN Conversation Summary of HN discussion 🧵👇 #hacker-news [bsky, 0 points, 1 comments]
- Unsupervised Elicitation of Language Models https://arxiv.org/abs/2506.10139 (https://news.ycombinator.com/item?id=44276041) [bsky, 0 points, 0 comments]
Related