AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Paper • 2608.23041 • Published 2 days ago • 44
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 2 days ago • 178
SPADE: Self-Play in Adaptive Synthetic Executable Environments Paper • 2608.19197 • Published 7 days ago • 51
Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL Paper • 2608.17253 • Published 7 days ago • 93
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling Paper • 2608.15089 • Published 11 days ago • 441
SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning Paper • 2608.14277 • Published 12 days ago • 34
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning Paper • 2608.09888 • Published 16 days ago • 754
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published 16 days ago • 340
SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Paper • 2608.09802 • Published 16 days ago • 133