vix.ing · top · new · best · stats · spec

Scavenging Hyena: Distilling Transformers into Long Convolution Models

2024/01/31 by Tokiniaina Raharison Ralambomihanta, Shahrad Mohammadzadeh, Ralambomihanta, Tokiniaina Raharison +7 · 2 citations
Arts and Humanities · Earth and Planetary Sciences · Physics and Astronomy · #Archaeological and Geological Studies #Astro and Planetary Science #Computation and Language (cs.CL) #FOS: Computer and information sciences #Geological formations and processes #Machine Learning (cs.LG)

paper · pdf · doi:10.48550/arxiv.2401.17574

openalex publication_date 2024/01/31 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

The rapid evolution of Large Language Models (LLMs), epitomized by architectures like GPT-4, has reshaped the landscape of natural language processing. This paper introduces a pioneering approach to address the efficiency concerns associated with LLM pre-training, proposing the use of knowledge distillation for cross-architecture transfer. Leveraging insights from the efficient Hyena mechanism, our method replaces attention heads in transformer models by Hyena, offering a cost-effective alternative to traditional pre-training while confronting the challenge of processing long contextual information, inherent in quadratic attention mechanisms. Unlike conventional compression-focused methods, our technique not only enhances inference speed but also surpasses pre-training in terms of both accuracy and efficiency. In the era of evolving LLMs, our work contributes to the pursuit of sustainable AI solutions, striking a balance between computational power and environmental impact.

Cited by

Related