| |
Orthrus-Qwen3: up to 7.8×tokens/forward on Qwen3, identical output distribution
Orthrus-Qwen3 is a new framework that accelerates Large Language Model inference by up to 7.8× through parallel token generation while maintaining identical output distribution to the original Qwen3 models. The system combines autoregressive and diffusion architectures with a shared KV cache, achieving lossless generation with minimal memory overhead and outperforming existing methods like speculative decoding and other diffusion-based language models.
Read Full Article →
← More Tech news