Knowledge Distillation of Black-Box Large Language Models
2024/01/13 by Hongzhan Chen, Ruijun Chen, Chen, Hongzhan +11 · 11 voices
#cs.CL
paper · pdf · doi:10.48550/arxiv.2401.07013
Abstract
Given the exceptional performance of proprietary large language models (LLMs) like GPT-4, recent research has increasingly focused on boosting the capabilities of smaller models through knowledge distillation (KD) from these powerful yet black-box teachers. While leveraging the high-quality outputs of these teachers is advantageous, the inaccessibility of their internal states often limits effective knowledge transfer. To overcome this limitation, we introduce Proxy-KD, a novel method that uses a proxy model to facilitate the efficient transfer of knowledge from black-box LLMs to smaller models. Our experiments show that Proxy-KD not only enhances the performance of KD from black-box teacher models but also surpasses traditional white-box KD techniques.~This approach presents a compelling new avenue for distilling knowledge from advanced LLMs.
Discussions
- Knowledge Distillation of Black-Box Large Language Models (2024) [hn, 123 points, 23 comments]
- Knowledge Distillation of Black-Box Large Language Models https://arxiv.org/abs/2401.07013 (https://news.ycombinator.com/item?id=48712420) [bsky, 0 points, 0 comments]
- Knowledge Distillation of Black-Box Large Language Models (2024) https://arxiv.org/abs/2401.07013 (https://news.ycombinator.com/item?id=48712420) [bsky, 0 points, 0 comments]
- Knowledge Distillation of Black-Box Large Language Models (2024) [bsky, 0 points, 0 comments]
- Knowledge Distillation of Black-Box Large Language Models #HackerNews https://arxiv.org/abs/2401.07013 [bsky, 0 points, 0 comments]
- ブラックボックス化された巨大LLMの知識を、より小さなモデルへ効率的に抽出(蒸留)する新手法が登場。推論コスト削減と軽量化を両立し、実用的なAI導入を加速させる重要なブレイクスルーです。技術的課題の克服により、AI活用がさらに身近になります。 #AI #TechNews https://arxiv.org/abs/2401.07013 [bsky, 0 points, 0 comments]
- Knowledge Distillation of Black-Box Large Language Models https://arxiv.org/abs/2401.07013 [bsky, 0 points, 0 comments]
- 📰 Knowledge Distillation of Black-Box Large Language Models (2024) 🔗 https://arxiv.org/abs/2401.07013 💬 Discuss on HN [bsky, 0 points, 0 comments]
- Student models learning from teacher outputs without seeing the internal weights – but are we just copying surface patterns or actually transferring reasoning? https://arxiv.org/abs/2401.07013 [bsky, 0 points, 1 comments]
- Knowledge Distillation of Black-Box Large Language Models (2024) https:// arxiv.org/abs/2401.07013 # arxiv [mastodon, 0 points, 0 comments]
- Knowledge Distillation of Black-Box Large Language Models (2024) https://arxiv.org/abs/2401.07013 https://news.ycombinator.com/item?id=48712420 [bsky, 0 points, 0 comments]
Related