2021/10/07 by Dhruv Guliani, Guliani, Dhruv, Lillian Zhou +13
Computer Science · Engineering · #68T10 #Distributed #FOS: Computer and information sciences #I.2.7 #Internet Traffic Analysis and Secure E-voting #Machine Learning (cs.LG) #Parallel #Privacy-Preserving Technologies in Data #Traffic Prediction and Management Techniques #and Cluster Computing (cs.DC)
paper · pdf · doi:10.48550/arxiv.2110.03634
openalex publication_date 2021/10/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Federated learning can be used to train machine learning models on the edge on local data that never leave devices, providing privacy by default. This presents a challenge pertaining to the communication and computation costs associated with clients' devices. These costs are strongly correlated with the size of the model being trained, and are significant for state-of-the-art automatic speech recognition models. We propose using federated dropout to reduce the size of client models while training a full-size model server-side. We provide empirical evidence of the effectiveness of federated dropout, and propose a novel approach to vary the dropout rate applied at each layer. Furthermore, we find that federated dropout enables a set of smaller sub-models within the larger model to independently have low word error rates, making it easier to dynamically adjust the size of the model deployed for inference.