WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents Paper • 2609.40325 • Published 12 days ago • 106
Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight Paper • 2610.08077 • Published 6 days ago • 159
DepthBench: Measuring How Residual Connections Enable More Computational Depth Paper • 2609.32534 • Published 16 days ago • 32
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents Paper • 2609.27334 • Published 19 days ago • 56
RISE: Recursive Improvement via Self-Extrapolating Policy Distillation Paper • 2609.05295 • Published Sep 4 • 19
EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness? Paper • 2609.04280 • Published Sep 3 • 32
DarwinX: Evolving Agent Harnesses Through Natural Selection Paper • 2608.07545 • Published Jul 31 • 118
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses Paper • 2608.12307 • Published Aug 12 • 118