-
Cosmos 3: Omnimodal World Models for Physical AI
Paper • 2606.02800 • Published • 142 -
Robots Need More than VLA and World Models
Paper • 2606.06556 • Published • 31 -
VLANeXt: Recipes for Building Strong VLA Models
Paper • 2602.18532 • Published • 52 -
Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models
Paper • 2606.11025 • Published • 41
Collections
Discover the best community collections!
Collections including paper arxiv:2606.02800
-
Towards Scalable Pre-training of Visual Tokenizers for Generation
Paper • 2512.13687 • Published • 108 -
MMGR: Multi-Modal Generative Reasoning
Paper • 2512.14691 • Published • 121 -
Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
Paper • 2512.23447 • Published • 100 -
LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation
Paper • 2512.23576 • Published • 66
-
LLM Pruning and Distillation in Practice: The Minitron Approach
Paper • 2408.11796 • Published • 61 -
TableBench: A Comprehensive and Complex Benchmark for Table Question Answering
Paper • 2408.09174 • Published • 53 -
To Code, or Not To Code? Exploring Impact of Code in Pre-training
Paper • 2408.10914 • Published • 45 -
Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications
Paper • 2408.11878 • Published • 64
-
Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models
Paper • 2606.03988 • Published • 126 -
Cosmos 3: Omnimodal World Models for Physical AI
Paper • 2606.02800 • Published • 142 -
Agents' Last Exam
Paper • 2606.05405 • Published • 387 -
ABot-Earth 0.5: Generative 3D Earth Model
Paper • 2606.09967 • Published • 488
-
TradingAgents: Multi-Agents LLM Financial Trading Framework
Paper • 2412.20138 • Published • 119 -
PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training
Paper • 2606.03264 • Published • 27 -
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
Paper • 2605.23904 • Published • 264 -
VibeVoice Technical Report
Paper • 2508.19205 • Published • 177
-
Code as Agent Harness
Paper • 2605.18747 • Published • 225 -
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper • 2605.12500 • Published • 195 -
From Context to Skills: Can Language Models Learn from Context Skillfully?
Paper • 2604.27660 • Published • 171 -
PhysBrain 1.0 Technical Report
Paper • 2605.15298 • Published • 145
-
Cosmos World Foundation Model Platform for Physical AI
Paper • 2501.03575 • Published • 84 -
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
Paper • 2502.11831 • Published • 20 -
PhysicsGen: Can Generative Models Learn from Images to Predict Complex Physical Relations?
Paper • 2503.05333 • Published • 8 -
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
Paper • 2503.15558 • Published • 52
-
Foundation Models in Robotics: Applications, Challenges, and the Future
Paper • 2312.07843 • Published • 16 -
Neural Fields in Robotics: A Survey
Paper • 2410.20220 • Published • 5 -
Robots Pre-train Robots: Manipulation-Centric Robotic Representation from Large-Scale Robot Dataset
Paper • 2410.22325 • Published • 10 -
Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning
Paper • 2410.21845 • Published • 16
-
Cosmos 3: Omnimodal World Models for Physical AI
Paper • 2606.02800 • Published • 142 -
Robots Need More than VLA and World Models
Paper • 2606.06556 • Published • 31 -
VLANeXt: Recipes for Building Strong VLA Models
Paper • 2602.18532 • Published • 52 -
Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models
Paper • 2606.11025 • Published • 41
-
Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models
Paper • 2606.03988 • Published • 126 -
Cosmos 3: Omnimodal World Models for Physical AI
Paper • 2606.02800 • Published • 142 -
Agents' Last Exam
Paper • 2606.05405 • Published • 387 -
ABot-Earth 0.5: Generative 3D Earth Model
Paper • 2606.09967 • Published • 488
-
TradingAgents: Multi-Agents LLM Financial Trading Framework
Paper • 2412.20138 • Published • 119 -
PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training
Paper • 2606.03264 • Published • 27 -
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
Paper • 2605.23904 • Published • 264 -
VibeVoice Technical Report
Paper • 2508.19205 • Published • 177
-
Code as Agent Harness
Paper • 2605.18747 • Published • 225 -
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper • 2605.12500 • Published • 195 -
From Context to Skills: Can Language Models Learn from Context Skillfully?
Paper • 2604.27660 • Published • 171 -
PhysBrain 1.0 Technical Report
Paper • 2605.15298 • Published • 145
-
Towards Scalable Pre-training of Visual Tokenizers for Generation
Paper • 2512.13687 • Published • 108 -
MMGR: Multi-Modal Generative Reasoning
Paper • 2512.14691 • Published • 121 -
Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
Paper • 2512.23447 • Published • 100 -
LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation
Paper • 2512.23576 • Published • 66
-
Cosmos World Foundation Model Platform for Physical AI
Paper • 2501.03575 • Published • 84 -
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
Paper • 2502.11831 • Published • 20 -
PhysicsGen: Can Generative Models Learn from Images to Predict Complex Physical Relations?
Paper • 2503.05333 • Published • 8 -
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
Paper • 2503.15558 • Published • 52
-
LLM Pruning and Distillation in Practice: The Minitron Approach
Paper • 2408.11796 • Published • 61 -
TableBench: A Comprehensive and Complex Benchmark for Table Question Answering
Paper • 2408.09174 • Published • 53 -
To Code, or Not To Code? Exploring Impact of Code in Pre-training
Paper • 2408.10914 • Published • 45 -
Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications
Paper • 2408.11878 • Published • 64
-
Foundation Models in Robotics: Applications, Challenges, and the Future
Paper • 2312.07843 • Published • 16 -
Neural Fields in Robotics: A Survey
Paper • 2410.20220 • Published • 5 -
Robots Pre-train Robots: Manipulation-Centric Robotic Representation from Large-Scale Robot Dataset
Paper • 2410.22325 • Published • 10 -
Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning
Paper • 2410.21845 • Published • 16