HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
Abstract
Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale; robot-free UMI capture scales readily, and current practice uses the resulting data mainly for pre-training, adding a small real-robot "anchor" at post-training. We ask whether raising the fidelity of robot-free UMI data, rather than shrinking the real-robot fraction, can remove that anchor. We present HiFi-UMI, a portable UMI data-production system co-designed for trajectory accuracy, inter-gripper relative pose, synchronization, and field of view: head-mounted offline stereo-inertial SLAM, native rather than reconstructed relative pose, a shared microsecond GPIO trigger, and two wide-angle cameras per hand covering ~200 degrees. It reaches 3 mm workspace-local end-effector accuracy without external tracking infrastructure. Using this corpus, we demonstrate zero-robot post-training: a policy post-trained solely on HiFi-UMI demonstrations deploys directly on a real robot and matches in-domain teleoperation across three backbones spanning the vision-language-action and world-action-model families, with success-rate differences of -2.5, +3.1, and -0.6 percentage points on StarVLA-QwenPI, OpenPI-pi_0.5, and LingBot-VA; the strongest policy reaches 85% on a precision insertion task, even though the teleoperation baseline is collected in the evaluation scene and no HiFi-UMI trajectory is. Pre-training on 4,000 hours from the same corpus lowers action error on ten unseen tasks by 41% and, on StarVLA-QwenPI, raises real-robot success by a further 18.1 percentage points. We open-source HiFi-UMI-2K, 2,000 hours of microsecond-synchronized, ultra-wide-FoV demonstrations, each automatically reconstructed and validated through simulation replay, as a large-scale, high-fidelity resource for the robot-learning community.
Community
Hi everyone! We’re excited to share HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone.
The central question is simple: can we eliminate target-task robot teleoperation from post-training, rather than merely reduce it?
HiFi-UMI is a portable, robot-free data-production system co-designed for action fidelity. It achieves 3 mm workspace-local end-effector accuracy, <40 μs cross-sensor synchronization, and ultra-wide six-view sensing, together with automated trajectory reconstruction, simulation replay, and quality validation.
Our main findings:
Across three VLA and WAM backbones—StarVLA-QwenPI, OpenPI-π0.5, and LingBot-VA—post-training using only HiFi-UMI demonstrations matches in-domain robot teleoperation, with success-rate differences of −2.5, +3.1, and −0.6 percentage points.
The strongest policy reaches 85% success on precision insertion, despite no HiFi-UMI demonstration being collected in the evaluation scene.
Pre-training on 4,000 hours reduces action error on ten unseen tasks by 41% and improves real-robot success by 18.1 percentage points.
We release HiFi-UMI-2K: 2,000 hours and 482K+ replayable demonstrations across 110+ scenes under CC BY 4.0.
Our key takeaway: robot-free data can support deployment—not only pre-training—when it is sufficiently high-fidelity and action-aligned.
📄 Paper:https://arxiv.org/abs/2607.25895
🌐 主页:https://cloud.simpleai.tech/simple-world-lab/hifi-umi/
🤗 数据集:https://huggingface.co/datasets/simple-world-lab/HiFi-UMI-2K
🔥 Daily Paper:https://huggingface.co/papers/2607.25895
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- RealDexUMI: A Wearable Universal Manipulation Interface for Dexterous Robot Learning (2026)
- AXIS: A Growable Community-Driven Data Engine for Scalable Robot Manipulation (2026)
- Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models (2026)
- EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations (2026)
- ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining (2026)
- Native Video-Action Pretraining for Generalizable Robot Control (2026)
- Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2607.25895 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper