| |
DAPO: An Open-source RL System from ByteDance Seed and Tsinghua AIR
ByteDance Seed and Tsinghua AIR have released DAPO (Decoupled Clip and Dynamic sAmpling Policy Optimization), an open-source reinforcement learning system for large-scale language model training that achieves state-of-the-art performance with 50% accuracy on AIME 2024 using Qwen2.5-32B. The system includes the novel DAPO algorithm, complete code infrastructure, and training datasets, making scalable LLM reinforcement learning accessible to the broader research community. The approach demonstrates improved training stability through careful management of response length growth, reward signal consistency, and balanced exploration-exploitation dynamics.
Read Full Article →
← More Tech news