| |
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
A benchmark study of Qwen 3.8 27B quantizations found that 4-bit Q4_K_M quantization (17 GB) maintains full model performance on popular coding and reasoning benchmarks while fitting on consumer GPUs like the RTX 4090, but 1-bit compression severely degrades performance to near random chance. The research tested multiple quantization levels across GPQA Diamond, instruction-following, and coding benchmarks, with results showing that compression quality is heavily dependent on model size and reasoning effort settings.
Read Full Article →
← More Tech news