2026/04/16 by Ye Su, Yong Liu
#cs.LG
To quantify the geometric capacity of transformers, we develop a tropical-geometric framework for analyzing the spatial partitions induced by conditioned self-attention. In the zero-temperature limit, we show that fixed-key top-1 routing is exactly represented by a power diagram in query space, while an auxiliary log-lifted value parameterization yields a vector-valued tropical rational representation. For Multi-Head Self-Attention (MHSA) with sequence length N and H attention heads, the joint routing geometry is encoded by Minkowski sums of headwise Newton polytopes, giving an O(NH) universal bound that sharpens to O((HN)^dmodel-1) once the number of heads reaches the intrinsic dimension dmodel. Extending this analysis across depth L, we derive the first tight asymptotic bounds on the number of linear regions in transformers (Θ (N^min\H,dmodel-1\L)). We further show that finite-temperature softmax preserves the top-1 routing structure and admits exponentially decaying local approximation and differential bounds away from routing boundaries.