| |
Cerebras has made the Qwen 3.8 27B model available on its public endpoints, capable of delivering approximately 1,500 tokens per second with a 128k context window on paid tiers. The platform offers unpruned, original versions of open-source models with selective weight-only quantization for storage while maintaining full precision for sensitive operations, and clarifies that any future compression techniques will be offered as separate, transparently named endpoints.
Read Full Article →
← More Tech news