FrontierChallenge: Evaluating Scientific Workflow Completion Paper • 2608.24979 • Published 10 days ago • 148
JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution Paper • 2608.25593 • Published 9 days ago • 117
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 11 days ago • 205
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 11 days ago • 205
OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation Paper • 2606.17628 • Published Jun 16 • 29
Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates Paper • 2601.18510 • Published Jan 26 • 1
Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies Paper • 2604.00830 • Published Apr 2 • 15
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections Paper • 2605.15030 • Published May 14
GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, and Video Paper • 2602.03328 • Published Feb 3
Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates Paper • 2601.18510 • Published Jan 26 • 1
Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents Paper • 2606.06036 • Published Jun 4 • 78
Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents Paper • 2606.06036 • Published Jun 4 • 78
Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents Paper • 2606.06036 • Published Jun 4 • 78
EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments Paper • 2606.13681 • Published Jun 11 • 144
Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Paper • 2604.28185 • Published Apr 30 • 92
MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome Paper • 2603.28407 • Published Mar 30 • 72
MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome Paper • 2603.28407 • Published Mar 30 • 72
DLLM-Searcher: Adapting Diffusion Large Language Model for Search Agents Paper • 2602.07035 • Published Feb 3 • 31
Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs Paper • 2601.08763 • Published Jan 13 • 150