MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis Paper • 2607.27146 • Published 4 days ago • 22
Flux-OPD: On-Policy Distillation with Evolving Contexts Paper • 2607.28022 • Published 3 days ago • 39
Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory Paper • 2607.27919 • Published 3 days ago • 48
DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space Paper • 2607.25675 • Published 5 days ago • 60
CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation Paper • 2607.16955 • Published 15 days ago • 4
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Paper • 2607.25659 • Published 5 days ago • 80
CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents Paper • 2607.25431 • Published 5 days ago • 99
Pass the Baton: Trajectory-Relayed On-Policy Distillation Paper • 2607.26057 • Published 5 days ago • 29
Towards Robust Reinforcement Learning for Small-Scale Language Model Agents Paper • 2607.25091 • Published 6 days ago • 6
Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning Models Paper • 2607.22098 • Published 9 days ago • 8
Codifying the Judge: Scalable Evaluation via Program Distillation Paper • 2607.22561 • Published May 29 • 8
Sample-Efficient Learning from Agent Experience Paper • 2607.21051 • Published 10 days ago • 19
Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking Paper • 2607.19747 • Published 11 days ago • 31
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs Paper • 2605.09635 • Published 10 days ago • 63
Multi-Turn On-Policy Distillation with Prefix Replay Paper • 2607.04763 • Published 17 days ago • 12
AutoIndex: Learning Representation Programs for Retrieval Paper • 2607.18603 • Published 12 days ago • 10