vix.ing · top · new · best · stats

Reweighted Proximal Pruning for Large-Scale Language Representation

2019/09/27 by Fu-Ming Guo, Sijia Liu, Guo, Fu-Ming +7 · 2 citations
Computer Science · Mathematics · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Neural and Evolutionary Computing (cs.NE) #Topic Modeling #cs.CL #cs.LG #cs.NE #stat.ML

paper · pdf · doi:10.48550/arxiv.1909.12486

openalex publication_date 2019/09/27 · arxiv created 2019/12/23 · arxiv updated 2019/12/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Recently, pre-trained language representation flourishes as the mainstay of the natural language understanding community, e.g., BERT. These pre-trained language representations can create state-of-the-art results on a wide range of downstream tasks. Along with continuous significant performance improvement, the size and complexity of these pre-trained neural models continue to increase rapidly. Is it possible to compress these large-scale language representation models? How will the pruned language representation affect the downstream multi-task transfer learning objectives? In this paper, we propose Reweighted Proximal Pruning (RPP), a new pruning method specifically designed for a large-scale language representation model. Through experiments on SQuAD and the GLUE benchmark suite, we show that proximal pruned BERT keeps high accuracy for both the pre-training task and the downstream multiple fine-tuning tasks at high prune ratio. RPP provides a new perspective to help us analyze what large-scale language representation might learn. Additionally, RPP makes it possible to deploy a large state-of-the-art language representation model such as BERT on a series of distinct devices (e.g., online servers, mobile phones, and edge devices).

Citations

Cited by

Related