N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation Paper • 2607.23783 • Published 10 days ago • 31
EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents Paper • 2607.28229 • Published 5 days ago • 7
RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models Paper • 2607.26991 • Published 6 days ago • 8
N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens Paper • 2607.23782 • Published 10 days ago • 56
Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory Paper • 2607.24368 • Published 9 days ago • 33
OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs Paper • 2607.25669 • Published 8 days ago • 9
MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Paper • 2607.25948 • Published 8 days ago • 18
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published 9 days ago • 33
WorldDiT: A Unified Diffusion Architecture for World and Action Modeling Paper • 2607.23909 • Published 9 days ago • 7
A Vocabulary for Multi-Agent Automated Research Systems Paper • 2607.22682 • Published 23 days ago • 4
UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Paper • 2607.23373 • Published 11 days ago • 6
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding Paper • 2607.24743 • Published 9 days ago • 11
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Paper • 2607.23855 • Published 10 days ago • 26
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification Paper • 2607.24027 • Published 9 days ago • 35
StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents Paper • 2607.22798 • Published 12 days ago • 61
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey Paper • 2607.21655 • Published 14 days ago • 192
From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search Paper • 2607.24280 • Published 9 days ago • 82
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents Paper • 2607.23588 • Published 10 days ago • 124