2025/04/03 by Christopher Wolfram, Aaron Schein, Wolfram, Christopher +1 · 1 voice · 6 citations
Computer Science · Engineering · #VLSI and Analog Circuit Testing #Advancements in Photolithography Techniques #Advancements in Semiconductor Devices and Circuit Design
paper · pdf · doi:10.48550/arxiv.2504.08775
How do the latent spaces used by independently-trained LLMs relate to one another? We study the nearest neighbor relationships induced by activations at different layers of 24 open-weight LLMs, and find that they 1) tend to vary from layer to layer within a model, and 2) are approximately shared between corresponding layers of different models. Claim 2 shows that these nearest neighbor relationships are not arbitrary, as they are shared across models, but Claim 1 shows that they are not "obvious" either, as there is no single set of nearest neighbor relationships that is universally shared. Together, these suggest that LLMs generate a progression of activation geometries from layer to layer, but that this entire progression is largely shared between models, stretched and squeezed to fit into different architectures.