vix.ing · top · new · best · stats

Towards Structured Dynamic Sparse Pre-Training of BERT

2021/08/13 by Anastasia Dietrich, Frithjof Gressmann, Dietrich, Anastasia +9 · 2 citations
Computer Science · #Advanced Neural Network Applications #Artificial intelligence #Computation and Language (cs.CL) #Computer science #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Physics #Topic Modeling #Training (meteorology) #cs.CL #cs.LG

paper · pdf · doi:10.48550/arxiv.2108.06277

published in arXiv (Cornell University) (Cornell University)

arxiv created 2021/08/13 · openalex publication_date 2021/08/13 · arxiv updated 2021/08/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Identifying algorithms for computational efficient unsupervised training of large language models is an important and active area of research. In this work, we develop and study a straightforward, dynamic always-sparse pre-training approach for BERT language modeling task, which leverages periodic compression steps based on magnitude pruning followed by random parameter re-allocation. This approach enables us to achieve Pareto improvements in terms of the number of floating-point operations (FLOPs) over statically sparse and dense models across a broad spectrum of network sizes. Furthermore, we demonstrate that training remains FLOP-efficient when using coarse-grained block sparsity, making it particularly promising for efficient execution on modern hardware accelerators.

Citations

Related