| |
Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
A developer created SlotStream, a tool that enables running the 104GB Qwen3.8-Flash-Next language model on a 48GB Mac by streaming model weights from disk storage at approximately 12 tokens per second. The project requires at least 512GB total disk space and provides a single Swift binary compatible with Ollama and OpenAI chat endpoints, with performance scaling based on available RAM (from an 8GB minimum to 48GB for optimal speed).
Read Full Article →
← More Tech news