| |
Cache-to-Cache: Direct Semantic Communication Between Large Language Models
Researchers propose Cache-to-Cache (C2C), a new communication method that allows large language models to exchange information directly through their internal KV-cache representations rather than through text, enabling richer semantic transfer between models. The approach uses neural networks to project and fuse cache data between models with a learnable gating mechanism, achieving 3.1-5.4% higher accuracy than text-based communication while delivering 2.5x speedup in latency. This paradigm enables multi-LLM systems to leverage complementary model strengths more efficiently without the information loss and latency costs of generating intermediate text.
Read Full Article →
← More Tech news