Rethinking OPD Collection This collection includes the models used in paper "Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe" • 5 items • Updated 2 days ago • 2
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 3 days ago • 75
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 6 days ago • 141
Safin-1: Safety from Within through Memory-Native State Evolution Paper • 2609.00092 • Published 6 days ago • 20
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning Paper • 2608.14290 • Published 23 days ago • 34
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering Paper • 2607.28568 • Published Jul 30 • 186
Weak-to-Strong Generalization via Direct On-Policy Distillation Paper • 2607.05394 • Published Jul 8 • 150
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published Jul 10 • 89
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 161
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? Paper • 2606.24530 • Published Jun 23 • 67
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 161 • 4
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 161
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? Paper • 2606.24530 • Published Jun 23 • 67