Stage-adaptive Token Selection for Efficient Omni-modal LLMs Paper • 2605.20035 • Published May 19 • 5
SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval Paper • 2603.08224 • Published Mar 9 • 2
Realtime-Venus: A full-duplex interaction system with asynchronous delegation Paper • 2609.13814 • Published 10 days ago • 10
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding Paper • 2605.18577 • Published May 18 • 6