| |
Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs
Researchers have successfully optimized Kimi K3, a 2.8-trillion-parameter language model, to run on a MacBook Pro at 1 token per second by streaming the model from four SSDs, achieving significant performance improvements over previous iterations. The Deltafin project maintains the full quality of the original model—including all 16 expert routers—without any pruning or shortcuts, demonstrating how far consumer hardware can be pushed for running extremely large models.
Read Full Article →
← More Tech news