| |
VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO
VibeThinker-3B, a 3-billion parameter language model, achieves competitive reasoning performance comparable to much larger models like DeepSeek V3.2 and Gemini 3 Pro through optimized supervised fine-tuning and reinforcement learning techniques. The model scores 94.3 on AIME26 mathematical reasoning and 80.2 on LiveCodeBench coding tasks while maintaining strong instruction-following capabilities, suggesting that verifiable reasoning can be effectively compressed into compact models without sacrificing performance.
Read Full Article →
← More Tech news