Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement Paper • 2609.01481 • Published 2 days ago • 9
CAST: Critique-Aware Supervision for Training Reliable Long-Horizon Tool-Calling Agents Paper • 2608.30147 • Published 3 days ago • 6
Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents Paper • 2608.31076 • Published 3 days ago • 15
Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence Paper • 2608.31075 • Published 3 days ago • 23
Rubric-to-Code Credit Assignment for Reinforcement Learning Paper • 2608.27906 • Published 6 days ago • 6
Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities Paper • 2608.28122 • Published 6 days ago • 59
Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO Paper • 2608.27351 • Published 7 days ago • 20
What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents Paper • 2608.27260 • Published 7 days ago • 69
Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs Paper • 2608.20953 • Published 13 days ago • 11
Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection Paper • 2608.20169 • Published 10 days ago • 11
One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows Paper • 2608.19741 • Published 14 days ago • 12
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 10 days ago • 205
Meta^n: Recursive Self-Improvement through Emergent Depth Paper • 2608.24735 • Published 9 days ago • 15
CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild Paper • 2608.23181 • Published 10 days ago • 34
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Paper • 2608.23041 • Published 10 days ago • 64