Abstract
Training vision-language-action (VLA) models for high-precision manipulation typically requires task-specific, high-quality data, which is slow and expensive to collect. To reduce this burden without compromising manipulation precision, we propose ε4P (Imperfection for Precision), a simple yet effective method that upcycles two otherwise discarded data sources: low-precision data from the target task and high-precision data from mismatched tasks.
Rather than naively mixing these imperfect data sources throughout co-training, ε4P controls where each source contributes along the flow-matching trajectory. Specifically, low-precision target-task data is used at high noise to preserve high-level task context, and high-precision task-mismatched data is used at low noise to transfer low-level action precision.
Through real-robot experiments on both sub-millimeter high-precision tasks and coarse-grained tasks, ε4P improves policy performance by up to 31.7 percentage points and can replace an equal amount of task-specific high-quality data with an average performance drop of only 4.2 percentage points. Overall, ε4P points toward a scalable paradigm in which heterogeneous imperfect data can be systematically repurposed to reduce reliance on costly task-specific high-quality data.