Deep sequence models tend to memorize geometrically; it is unclear why
2025/10/30 by Shahriar Noroozizadeh, Vaishnavh Nagarajan, Elan Rosenfeld +1 · 11 voices · 1 citation
#cs.LG #cs.AI #cs.CL #stat.ML
paper · pdf
Abstract
Deep sequence models are said to store atomic facts predominantly in the form of associative memory: a brute-force lookup of co-occurring entities. We identify a dramatically different form of storage of atomic facts that we term as geometric memory. Here, the model has synthesized embeddings encoding novel global relationships between all entities, including ones that do not co-occur in training. Such storage is powerful: for instance, we show how it transforms a hard reasoning task involving an ℓ-fold composition into an easy-to-learn 1-step navigation task. From this phenomenon, we extract fundamental aspects of neural embedding geometries that are hard to explain. We argue that the rise of such a geometry, as against a lookup of local associations, cannot be straightforwardly attributed to typical supervisory, architectural, or optimizational pressures. Counterintuitively, a geometry is learned even when it is more complex than the brute-force lookup. Then, by analyzing a connection to Node2Vec, we demonstrate how the geometry stems from a spectral bias that -- in contrast to prevailing theories -- indeed arises naturally despite the lack of various pressures. This analysis also points out to practitioners a visible headroom to make Transformer memory more strongly geometric. We hope the geometric view of parametric memory encourages revisiting the default intuitions that guide researchers in areas like knowledge acquisition, capacity, discovery, and unlearning.
Cited by
Discussions
- Fascinating piece of work arxiv.org/abs/2510.26745 [bsky, 22 points, 0 comments]
- 2/ Our v2 paper has more experiments, hopefully clearer discussions & more related work. arxiv.org/abs/2510.26745 Work led by amazing intern @shnoroozi.bsky.social during an eventful summer w @ElanRos [bsky, 19 points, 1 comments]
- Deep sequence models tend to memorize geometrically; it is unclear why [hn, 4 points, 0 comments]
- Deep sequence models tend to memorize geometrically [hn, 3 points, 0 comments]
- Deep sequence models tend to memorize geometrically; it is unclear why [hn, 3 points, 0 comments]
- I love this piece and agree fully with the cultural artifact framing. I’m curious if you’ve been following any of these work that’s an emerging around geometric learning and what we’re discovering for [bsky, 3 points, 1 comments]
- arxiv.org/abs/2510.26745 [bsky, 2 points, 0 comments]
- Deep sequence models tend to memorize geometrically; it is unclear why [hn, 1 points, 0 comments]
- available when, what countries existed, what political systems were in place, etc. that makes dates far from independent and identically distributed conditioned on the event (or vice versa). And we kn [bsky, 1 points, 1 comments]
- “Deep sequence models tend to memorize geometrically; it is unclear why” arxiv.org/abs/2510.26745 [bsky, 0 points, 0 comments]
- Every author writing like this should be required to rewrite abstracts in plain English and read it aloud to an audience of their peers, before they can publish it. Summary: Conjectural with nice diag [bsky, 0 points, 0 comments]
Related