XiaomiMiMo/MiMo-V2.6-Flash-MOPD
Text Generation • 311B • Updated • 7.67k • 62
None defined yet.
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL