The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 170
Optimizing Visual Generative Models via Distribution-wise Rewards Paper • 2607.02291 • Published Jul 2 • 18
Native Active Perception as Reasoning for Omni-Modal Understanding Paper • 2606.19341 • Published Jun 17 • 19
Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models Paper • 2606.03988 • Published Jun 3 • 126
AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions Paper • 2605.25707 • Published May 25 • 6
SafeDiffusion-R1: Online Reward Steering for Safe Diffusion Post-Training Paper • 2605.18719 • Published May 18 • 7
π-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows Paper • 2605.14678 • Published May 19 • 108