Cache-to-Cache: Direct Semantic Communication Between LLMs (2025)
A new paradigm called Cache-to-Cache (C2C) enables direct semantic communication between Large Language Models (LLMs), addressing limitations of text-based communication. C2C projects and fuses the KV-cache of source and target models using a neural network, allowing direct semantic transfer and avoiding explicit intermediate text generation. Experiments show C2C achieves 6.4-14.2% higher average accuracy than individual models and outperforms text communication by 3.1-5.4%, with a 2.5x speedup in latency. This method leverages rich semantic information for improved performance and efficiency.
This paper introduces Cache-to-Cache, a novel paradigm for direct LLM semantic communication, unlike previous methods that rely on text, achieving 2.5x speedup and higher accuracy.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 18, 2026, 21:00 UTC
- Ingested
- Sep 18, 2026, 21:00
- Source type
- Research
Discussion trend
The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.
Full text isn't available here.
Read at source →