QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents Paper • 2609.33848 • Published 8 days ago • 43
Marathoner: Ultra-Long-Horizon Autonomous Intelligence Paper • 2609.34378 • Published 7 days ago • 41
What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling Paper • 2609.34981 • Published 6 days ago • 131
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy Paper • 2609.28660 • Published 12 days ago • 16
VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models Paper • 2609.04355 • Published 17 days ago • 9
SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL Paper • 2609.29050 • Published 11 days ago • 13
RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling Paper • 2609.22947 • Published 16 days ago • 43
Fingers as Legs: Learning Self-Supported Locomotion and Manipulation with an Anthropomorphic Hand Paper • 2609.17172 • Published 20 days ago • 5
Towards Full Pipeline FP8 Reinforcement Learning for LLMs Paper • 2609.22870 • Published 16 days ago • 17
CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents Paper • 2609.18779 • Published 19 days ago • 16
RULER: Instance-aware Rubric Rewards for SVG Generation Paper • 2609.25270 • Published 14 days ago • 102
WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory Paper • 2609.24984 • Published 14 days ago • 157