| |
Run Qwen3.8 27B locally: real numbers from my Mac Studio
Qwen3.8 27B, a new 27-billion-parameter AI model, generates text at approximately 14 tokens per second on a Mac Studio M3 Ultra, making it practical for local use despite being slower than its predecessor Qwen3.6:27B. The model has impressed the AI community with real-world performance in coding and OCR tasks, though its practical speed and token efficiency mean wall-clock time per completed answer is comparable to the older version. For users wanting to run it locally, 32GB of RAM comfortably handles the Q4 quantization version (17GB), while 16GB can run the Q2 variant.
Read Full Article →
← More Tech news