β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation Paper • 2607.28582 • Published Jul 30 • 24
Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs Paper • 2605.12460 • Published May 12 • 18
Open-Reasoner-Zero/Open-Reasoner-Zero-7B Reinforcement Learning • 8B • Updated Apr 7, 2025 • 1.85k • 34