Harnessing the Universal Geometry of Embeddings
研究人员提出了一种无需配对数据、编码器或预定义匹配即可在不同向量空间之间转换文本嵌入的方法。这种无监督方法将嵌入转换为通用潜在表示,并在具有不同架构、参数数量和训练数据集的模型之间实现了高余弦相似度。这项技术对向量数据库的安全性具有重要影响,因为攻击者可能仅通过嵌入向量提取敏感信息,从而进行分类和属性推断。
- 发布
- 2026年9月6日 20:31
- 来源类型
- 研究
- 档位
- 社区
- 信源状态
- 正常
时间以 UTC 显示
更多信息
View PDF HTML (experimental)
Abstract: We introduce the first method for translating text embeddings from one vector space to another without any paired data, encoders, or predefined sets of matches. Our unsupervised approach translates any embedding to and from a universal latent representation (i.e., a universal semantic structure conjectured by the Platonic Representation Hypothesis). Our translations achieve high cosine similarity across model pairs with different architectures, parameter counts, and training datasets.
The ability to translate unknown embeddings into a different space while preserving their geometry has serious implications for the security of vector databases. An adversary with access only to embedding vectors can extract sensitive information about the underlying documents, sufficient for classification and attribute inference.
Subjects:
Machine Learning (cs.LG)
Cite as: arXiv:2505.12540 [cs.LG]
(or arXiv:2505.12540v4 [cs.LG] for this version)
https://doi.org/10.48550/arXiv.2505.12540
arXiv-issued DOI via DataCite
Submission history
From: Rishi Jha [view email] [v1] Sun, 18 May 2025 20:37:07 UTC (3,179 KB)
[v2] Tue, 20 May 2025 15:38:41 UTC (3,180 KB)
[v3] Wed, 25 Jun 2025 21:04:02 UTC (2,407 KB)
[v4] Mon, 26 Jan 2026 14:47:13 UTC (2,424 KB)