HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 9 days ago • 261
S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? Paper • 2608.31100 • Published 10 days ago • 38
CogEvol: Towards Efficient and Reliable Learning Environment Generation Paper • 2608.30968 • Published 10 days ago • 30
EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery Paper • 2606.13662 • Published Jun 11 • 33
EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery Paper • 2606.13662 • Published Jun 11 • 33
EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery Paper • 2606.13662 • Published Jun 11 • 33
Diagnosing Harmful Continuation in Answer-Correct Long-CoT Training Traces Paper • 2605.29288 • Published May 28 • 9
LongTraceRL Collection LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards • 5 items • Updated Jun 1 • 1
LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards Paper • 2605.31584 • Published May 29 • 43
Benchmarking Foundation Models with Language-Model-as-an-Examiner Paper • 2306.04181 • Published Jun 7, 2023