2020/12/30 by Sheng Shen, Shen, Sheng, Alexei Baevski +9 · 5 citations
Computer Science · #Machine Learning and ELM #Neural Networks and Applications #Neural Networks and Reservoir Computing #cs.CL
paper · pdf · doi:10.48550/arxiv.2012.15045
ACL 2021
arxiv created 2021/06/01 · arxiv updated 2021/06/03
We demonstrate that transformers obtain impressive performance even when some of the layers are randomly initialized and never updated. Inspired by old and well-established ideas in machine learning, we explore a variety of non-linear "reservoir" layers interspersed with regular transformer layers, and show improvements in wall-clock compute time until convergence, as well as overall performance, on various machine translation and (masked) language modelling tasks.