| |
vLLM v0.28.0
vLLM v0.28.0 was released with 584 commits from 270 contributors, featuring major performance optimizations for Kimi-K3 and DeepSeek V4 models, including decode context parallel support, fused kernels, and sparse MLA end-to-end functionality. The release also includes advances in speculative decoding, Model Runner V2 maturation with weight offloading capabilities, tiered KV cache offloading with disk support, a new Rust frontend with gRPC support, and increased default batching parameters for improved throughput.
Read Full Article →
← More Tech news