Kimi Linear: An Expressive, Efficient Attention Architecture - AllTheNews.today
Kimi Linear: An Expressive, Efficient Attention Architecture

Kimi Linear: An Expressive, Efficient Attention Architecture

The Kimi team has developed Kimi Linear, a hybrid linear attention architecture that outperforms traditional full attention mechanisms across various scenarios including short-context, long-context, and reinforcement learning tasks. The architecture features Kimi Delta Attention (KDA), an improved linear attention module with fine-grained gating that reduces KV cache usage by up to 75% and achieves up to 6 times faster decoding throughput for long contexts while maintaining superior performance. The team has open-sourced the implementation and released model checkpoints, positioning Kimi Linear as a viable replacement for conventional attention architectures.
Read Full Article →
arxiv.org
← Back to Latest