| |
Mithil Vakde trained a small transformer model that achieved 44% accuracy on the ARC-AGI-1 benchmark in just 67 cents of compute time, matching the performance of larger language models while being significantly faster and cheaper. The model uses test-time training on individual puzzles with architectural improvements like SwiGlu activations, RMSnorm, and 3D RoPE embeddings, along with optimizations like reduced augmentations and flash attention to improve sample efficiency. This represents a substantial upgrade to his previous work on the ARC-AGI benchmark, demonstrating that sample efficiency rather than model scale is key to solving complex reasoning tasks.
Read Full Article →
← More Tech news