2019/07/15 by Alejandro Molina, Patrick Schramowski, Molina, Alejandro +3 · 1 voice · 20 citations
Chemistry · Computer Science · Mathematics · #Activation function #Advanced Neural Network Applications #Artificial intelligence #Artificial neural network #Chemistry #Computer science #Deep learning #Domain Adaptation and Few-Shot Learning #End-to-end principle #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification #Mathematics #Neural and Evolutionary Computing (cs.NE) #Parametric statistics #Robustness (evolution) #cs.LG #cs.NE
paper · pdf · doi:10.48550/arxiv.1907.06732
published in arXiv (Cornell University) (Cornell University) · 17 Pages, 8 Figures
openalex publication_date 2019/07/15 · arxiv published 2019/07/15 · arxiv created 2020/02/04 · arxiv updated 2020/02/05 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
The performance of deep network learning strongly depends on the choice of the non-linear activation function associated with each neuron. However, deciding on the best activation is non-trivial, and the choice depends on the architecture, hyper-parameters, and even on the dataset. Typically these activations are fixed by hand before training. Here, we demonstrate how to eliminate the reliance on first picking fixed activation functions by using flexible parametric rational functions instead. The resulting Padé Activation Units (PAUs) can both approximate common activation functions and also learn new ones while providing compact representations. Our empirical evidence shows that end-to-end learning deep networks with PAUs can increase the predictive performance. Moreover, PAUs pave the way to approximations with provable robustness. https://github.com/ml-research/pau