Harnessing the Universal Geometry of Embeddings
- Published
- 09/06, 20:31
- Ingested
- 09/06, 23:00
- Source type
- Research
- Tier
- Community
- Source status
- Healthy
Researchers have introduced the first method for translating text embeddings between different vector spaces without paired data, encoders, or predefined matches. This unsupervised approach translates embeddings to and from a universal latent representation, achieving high cosine similarity across models with varying architectures, parameter counts, and training datasets. This capability has significant implications for vector database security, as adversaries could extract sensitive information from embedding vectors, enabling classification and attribute inference.
View PDF HTML (experimental)
Abstract: We introduce the first method for translating text embeddings from one vector space to another without any paired data, encoders, or predefined sets of matches. Our unsupervised approach translates any embedding to and from a universal latent representation (i.e., a universal semantic structure conjectured by the Platonic Representation Hypothesis). Our translations achieve high cosine similarity across model pairs with different architectures, parameter counts, and training datasets.
The ability to translate unknown embeddings into a different space while preserving their geometry has serious implications for the security of vector databases. An adversary with access only to embedding vectors can extract sensitive information about the underlying documents, sufficient for classification and attribute inference.
Subjects:
Machine Learning (cs.LG)
Cite as: arXiv:2505.12540 [cs.LG]
(or arXiv:2505.12540v4 [cs.LG] for this version)
https://doi.org/10.48550/arXiv.2505.12540
arXiv-issued DOI via DataCite
Submission history
From: Rishi Jha [view email] [v1] Sun, 18 May 2025 20:37:07 UTC (3,179 KB)
[v2] Tue, 20 May 2025 15:38:41 UTC (3,180 KB)
[v3] Wed, 25 Jun 2025 21:04:02 UTC (2,407 KB)
[v4] Mon, 26 Jan 2026 14:47:13 UTC (2,424 KB)