2025/09/14 by Jiyong Ma, Ma, Jiyong
Computer Science · #Artificial Intelligence (cs.AI) #Digital Filter Design and Implementation #FOS: Computer and information sciences #Image and Signal Denoising Methods #Machine Learning (cs.LG) #Neural Networks and Applications
paper · pdf · doi:10.48550/arxiv.2509.12285
openalex publication_date 2025/09/14 · openalex created_date 2025/10/18 · openalex updated_date 2026/07/28
In this paper, we present a maximum likelihood estimation approach to determine the value vector in transformer models. We model the sequence of value vectors, key vectors, and the query vector as a sequence of Gaussian distributions. The variance in each Gaussian distribution depends on the time step, the corresponding key vector, and the query vector. The mean value in each Gaussian distribution depends on the time step, and the corresponding value vector. This analysis may offer a new explanation of the scaled-dot-product function or softmax function used in transformer architectures [1]. Another explanation, inspired by [4], is based on the maximum entropy approach in natural language processing [5]. In this approach, a query vector and key vectors are used to derive the feature functions for the maximum entropy model.