2025/02/19 by Burc Gokden, Gokden, Burc · 1 voice · 1 citation
Chemistry · Computer Science · Mathematics · #Artificial Intelligence (cs.AI) #Artificial intelligence #Artificial neural network #Chemistry #Computation and Language (cs.CL) #Computational Physics and Python Applications #Computer science #FOS: Computer and information sciences #Inference #Machine Learning (cs.LG) #Machine learning #Mathematics #Net (polyhedron) #Operator (biology) #Tensor (intrinsic definition) #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2502.13502
openalex publication_date 2025/02/19 · arxiv published 2025/02/19 · arxiv updated 2025/02/22 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
We show that Large Language Model from Power Law Decoder Representations (PLDR-LLM) is a foundational model whose deductive outputs are invariant tensors up to a small perturbation. PLDR-LLM learns a singularity condition for the deductive outputs that enable the once-inferred energy-curvature tensor GLM to replace the deep neural network of power law graph attention (PLGA) generating the deductive outputs at inference. We demonstrate that a cache for GLM (G-cache) and KV-cache can be implemented in a straightforward manner to improve the inference time. The invariance and generalizable nature of deductive outputs is at a very high fidelity where deductive outputs have same RMSE and determinant values up to 15 decimal places after caching, and zero-shot benchmark scores remain unchanged. Ablation studies show that learned deductive outputs have distinct loss and accuracy characteristics from models pretrained with transferred, randomly initialized or identity tensors as a constant tensor operator and an LLM with scaled-dot product attention (SDPA) is a special case of PLDR-LLM where GLM is predefined as identity. The observed invariance characteristic introduces a novel asymmetry between training and inference phases with caching. We outline observed common characteristics of the deductive outputs for the learned singularity condition. We provide an implementation of a training and inference framework for PLDR-LLM with KV-cache and G-cache.