2024/01/23 by Tamir David Hay, Hay, Tamir David, Lior Wolf +1 · 5 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Ferroelectric and Negative Capacitance Devices #Machine Learning (cs.LG) #Neural Networks and Applications
paper · pdf · doi:10.48550/arxiv.2401.12819
openalex publication_date 2024/01/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In the pursuit of reducing the number of trainable parameters in deep transformer networks, we employ Reinforcement Learning to dynamically select layers during training and tie them together. Every few iterations, the RL agent is asked whether to train each layer i independently or to copy the weights of a previous layer j<i></i>