Harnessing the Universal Geometry of Embeddings
2025/05/18 by Rishi Jha, Jha, Rishi, Collin Zhang +5 · 40 voices · 36 citations
Computer Science · #Advanced Graph Neural Networks #Adversarial Robustness in Machine Learning #Cosine similarity #Embedding #Feature learning #Representation (politics) #Similarity (geometry) #Space (punctuation) #Topic Modeling #Vector space
paper · pdf · doi:10.48550/arxiv.2505.12540
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/05/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Abstract
We introduce the first method for translating text embeddings from one vector space to another without any paired data, encoders, or predefined sets of matches. Our unsupervised approach translates any embedding to and from a universal latent representation (i.e., a universal semantic structure conjectured by the Platonic Representation Hypothesis). Our translations achieve high cosine similarity across model pairs with different architectures, parameter counts, and training datasets. The ability to translate unknown embeddings into a different space while preserving their geometry has serious implications for the security of vector databases. An adversary with access only to embedding vectors can extract sensitive information about the underlying documents, sufficient for classification and attribute inference.
Citations
Cited by
Discussions
- Huh. Looks like Plato was right. A new paper shows all language models converge on the same "universal geometry" of meaning. Researchers can translate between ANY model's embeddings without seeing the [bsky, 255 points, 9 comments]
- anyway turns out platonism is true arxiv.org/abs/2505.12540 [bsky, 188 points, 17 comments]
- Harnessing the Universal Geometry of Embeddings [hn, 123 points, 40 comments]
- Strong Platonic Representation Hypothesis All embedding models, given large enough scale, can be translated between them without paired data Security implication: Embeddings aren’t encryption, they’re [bsky, 49 points, 6 comments]
- arxiv.org/abs/2505.12540 [bsky, 25 points, 1 comments]
- they are saying this is a win for platonism, but we used to say features persistent in data and relationships between different systems were simply physical reality arxiv.org/abs/2505.12540 [bsky, 18 points, 1 comments]
- That is absolutely fascinating. LLMs converge on the same structures. [bsky, 8 points, 1 comments]
- "The Platonic Representation Hypothesis [17] conjectures that all image models of sufficient size have the same latent representation." arxiv.org/pdf/2505.12540 [bsky, 7 points, 0 comments]
- arxiv.org/abs/2505.12540 - it is based around this paper which can be read as giving a lot of evidence to the notion that differently trained LLM models have shared ‘geometric’ features; that the lang [bsky, 4 points, 0 comments]
- arxiv.org/abs/2505.12540 [bsky, 3 points, 1 comments]
- oh I see. these guys are just amazingly philosophically naive. yes, different algorithms used to produce a concordance will create similar concordances, and those concordances should have similar stru [bsky, 3 points, 1 comments]
- Strong Platonic Hypothesis ftw arxiv.org/abs/2505.12540 [bsky, 3 points, 1 comments]
- Harnessing the Universal Geometry of Embeddings https://arxiv.org/abs/2505.12540 (https://news.ycombinator.com/item?id=44054425) [bsky, 2 points, 0 comments]
- we gotta stop letting computer scientists do linguistics arxiv.org/abs/2505.12540 [bsky, 2 points, 0 comments]
- This research introduces a method to translate (without training) text embeddings between models This may have significant implications, including the potential to extract sensitive information from e [bsky, 2 points, 0 comments]
- Harnessing the Universal Geometry of Embeddings “Our unsupervised approach translates any embedding to and from a universal latent representation (i.e., a universal semantic structure conjectured by t [bsky, 2 points, 1 comments]
- everyone is being all "uhh platonism is true" over this arxiv.org/abs/2505.12540 [bsky, 2 points, 1 comments]
- Like one paper on universal language embeddings name-drops Plato in a "Plato was right" context: [bsky, 2 points, 1 comments]
- My priors completely align with this humanist approach, but new results which point toward an ur-geometry of meaning have me questioning everything 🤯 arxiv.org/pdf/2505.12540 [bsky, 2 points, 1 comments]
- Rishi Jha et. al. show that the "Strong Platonic Representation Hypothesis" holds in practice: that two models trained in the same modality on the same objective converge on the same embedding space. [bsky, 2 points, 1 comments]
- Why is there so little discussion about this paper? Seems like a big deal: universal representation of thingness. arxiv.org/abs/2505.12540 [bsky, 2 points, 0 comments]
- Wow! #Embeddings are similar across models and text can be recovered with only the embedding. This is huge! arxiv.org/abs/2505.12540 [bsky, 1 points, 0 comments]
- We are turbocooked - arxiv.org/pdf/2505.125... [bsky, 1 points, 0 comments]
- I came across the paper Harnessing the Universal Geometry of Embeddings (arxiv.org/abs/2505.12540) and The Platonic Representation Hypothesis (arxiv.org/abs/2405.07987). It's pretty wild that differen [bsky, 1 points, 1 comments]
- What could possibly go wrong? [bsky, 1 points, 0 comments]
- I think the question is to what extent the thing learned is representational versus an alien agent that incidentally implements a compatible API to concepts implied by human language. Representation c [bsky, 1 points, 1 comments]
- This is an extremely powerful method! In essence, they can re-create the raw data from embedding space across models. arxiv.org/abs/2505.12540 [bsky, 1 points, 0 comments]
- Skipping over the bit in here about a "universal geometry of embeddings"* the news that researchers think they can extract original text from embedded vectors is 🤔 arxiv.org/pdf/2505.12540 * speaking [bsky, 1 points, 0 comments]
- Impressive work, but representational alignment is tricky. Sometimes preserving global geometry is ideal, other times, distinctions matter more. Philosophically, computational theories demand a more p [bsky, 0 points, 0 comments]
- What if we measured the cosine similarity of our similarity spaces? 👉👈 arxiv.org/abs/2505.12540 [bsky, 0 points, 0 comments]
- arxiv.org/abs/2505.12540 [bsky, 0 points, 0 comments]
- Harnessing the Universal Geometry of Embeddings https://arxiv.org/abs/2505.12540 (https://news.ycombinator.com/item?id=44054425) [bsky, 0 points, 0 comments]
- Harnessing the Universal Geometry of Embeddings https://arxiv.org/abs/2505.12540 [bsky, 0 points, 0 comments]
- Harnessing the Universal Geometry of Embeddings https://arxiv.org/abs/2505.12540 https://news.ycombinator.com/item?id=44054425 [bsky, 0 points, 0 comments]
- Harnessing the Universal Geometry of Embeddings [bsky, 0 points, 0 comments]
- Essential to understand the concept of universal semantic structure, a conjecture that there is a universal way that humans conceptualize and organize meaning. Starting to be evidence that LLMs conver [bsky, 0 points, 1 comments]
- Sobre esse paper arxiv.org/pdf/2505.12540 [bsky, 0 points, 0 comments]
- OK, this is truly cool: arxiv.org/abs/2505.12540 [bsky, 0 points, 0 comments]
- This might be a very useful paper in this context arxiv.org/abs/2505.12540 [bsky, 0 points, 0 comments]
- Harnessing the Universal Geometry of Embeddings arxiv.org/pdf/2505.12540 [bsky, 0 points, 0 comments]
Related