TL;DR
Vision-only samples use masked force inputs, while online correction samples retain measured force; the force pathway is trained only when real force data are available.
Source: [10]
On towel folding, π0.5 with online DAgger rose from 0 at 1.3K samples to 0.725 at 3.1K; LIFT reached 0.65 at 2.3K and peaked at 0.825 at 2.8K; LIFT without reactive force injection reached 0.925 at 2.4K and peaked at 0.95 at 2.8K.
Source: [3]
The authors identify human intervention as a limit on data-collection efficiency and an operator burden; intervention timing is difficult to judge, requires skilled operators, and can affect correction-data quality.
Source: [33]
Why This Matters
Source-paper contributions
Comparison baseline
The principal vision-only comparison is π0.5 with online DAgger, which uses the same online DAgger loop without force input.
Source: [5]
Comparison baselines
What the paper contributes
Evaluation datasets
Evaluation environment
Failure modes
For towel folding, a reported failure is that visual depth ambiguity can prevent the policy from determining whether the gripper has contacted or grasped the thin cloth.
Source: [3]
For book insertion, reported failures include pushing after the book has bottomed out or releasing before it is seated, which can cause the book to flip or fail to settle.
Source: [18]
Key Findings
Paper reports
On book insertion, LIFT reached a peak score of 0.6 at 4.6K samples, compared with 0.4 at 5.6K samples for π0.5 with online DAgger.
Source: [18]
On Hanoi ring placement, LIFT reached and maintained a score of 0.6 at the final checkpoint (2.0K samples); π0.5 with online DAgger peaked at 0.3 and ended at 0.2.
Source: [27]
The paper reports that LIFT showed no substantial degradation in most tested shifted settings involving objects, tablecloths, and lighting; this reported association does not establish that the method preserves generalization in untested conditions.
Source: [23]
On book insertion, LIFT reached a peak score of 0.6 at 4.6K samples, while the single-frame-force variant peaked at 0.4; on towel folding, the paper reports comparable performance between the single-frame-force and reactive variants.
The paper reports that LIFT with offline DAgger underperformed on all three tasks and fell to zero on book insertion, while online DAgger adds corrections on learner failure states.
The tested lightweight residual policy achieved substantially lower scores than LIFT; the authors attribute this in part to large required corrections from the weak base and sparse residual targets that make intervention timing difficult to learn.
Limitations
The experiments were conducted only on a single-arm platform; the authors identify bimanual tasks on dual-arm platforms and other force sensors as future directions.
Source: [33]
The authors note that post-training the full VLA carries substantial computational cost.
Source: [2]
How the method works
The reactive action stream decodes actions causally within a chunk, and the force pathway uses recent 6D end-effector force encoded as causal force memory with a latency-aligned causal mask.
LIFT caches the vision-language prefix and reuses it while reevaluating the action experts with updated force history within an action chunk.
Source: [25]
At initialization, copied action-expert weights and zero-initialized cross-attention output make the reactive force residual ineffective, so the augmented model starts with the same action output as the original model.
Evaluation metrics
Towel folding and book insertion use stage-based scores, while Hanoi ring placement uses binary success.
Source: [17]
Measurement conditions
Each checkpoint was evaluated with three groups of ten autonomous rollouts, and 95% confidence intervals were computed; the reported online sample counts represent human-intervention frames, the only online data used for training.
Source: [17]
Research question and scope
The paper asks how to inject force so a VLA can react quickly to contact while retaining force memory, how to preserve the pretrained VLA prior during adaptation, and how to post-train effectively when policy-dependent force distributions shift beyond offline coverage.
Source: [30]
Tested scope and boundaries
The evaluation covered three real-robot manipulation tasks: towel folding, book insertion, and Hanoi ring placement.
Source: [17]
The reported shifted-condition evaluation varied objects, tablecloths, and lighting; the paper reports results for most tested settings but does not claim substantial degradation was absent in every setting.
Source: [23]
Training setup
Paper Details
Machine Learning · Empirical
Original research: Never Too Late for Force: Accelerating VLA Post-Training with Reactive Force Injection · 2607.14236v2
Paper authors: Yi Wang, Wendi Chen, Zimo Wen, Han Xue, Xueqi Li, Wenye Yu, Zhijie Chen, Hao Yang, Jun Lv, Chuan Wen, Cewu Lu
Source license: CC BY-SA 4.0. This article summarizes and interprets the source using AI. Attribution does not imply endorsement by the source authors.
This adapted analysis is shared under the same CC BY-SA 4.0 license. This brief was checked offline against its admitted source evidence. A model semantic verdict was not obtained because of the retained capacity boundary. It is not presented as model-certified or independently replicated.
- Canonical source identity
- arXiv 2607.14236
- Analyzed source version
- v2
- Source retrieved
- BaitaPhish analysis published
- BaitaPhish analysis reviewed
Evidence & Provenance
Show evidence locators
Evidence labels locate support in the original paper; they do not establish independent replication.
- [1] · page 8 — Source passage: Admitted source passage
- [2] · page 9 — Source passage: Admitted source passage
- [3] · page 7 — Source passage: Admitted source passage
- [4] · page 7 — Source passage: Admitted source passage
- [5] · page 7 — Source passage: Admitted source passage
- [6] · page 8 — Source passage: Admitted source passage
- [7] · page 3 — Source passage: Admitted source passage
- [8] · page 9 — Source passage: Admitted source passage
- [9] · page 4 — Source passage: Admitted source passage
- [10] · page 6 — Source passage: Admitted source passage
- [11] · page 6 — Source passage: Admitted source passage
- [12] · page 2 — Source passage: Admitted source passage
- [13] · page 5 — Source passage: Admitted source passage
- [14] · page 6 — Source passage: Admitted source passage
- [15] · page 8 — Source passage: Admitted source passage
- [16] · page 7 — Source passage: Admitted source passage
- [17] · page 6 — Source passage: Admitted source passage
- [18] · page 7 — Source passage: Admitted source passage
- [19] · page 8 — Source passage: Admitted source passage
- [20] · page 9 — Source passage: Admitted source passage
- [21] · page 7 — Source passage: Admitted source passage
- [22] · page 5 — Source passage: Admitted source passage
- [23] · page 8 — Source passage: Admitted source passage
- [24] · page 5 — Source passage: Admitted source passage
- [25] · page 4 — Source passage: Admitted source passage
- [26] · page 7 — Source passage: Admitted source passage
- [27] · page 7 — Source passage: Admitted source passage
- [28] · page 2 — Source passage: Admitted source passage
- [29] · page 7 — Source passage: Admitted source passage
- [30] · page 2 — Source passage: Admitted source passage
- [31] · page 4 — Source passage: Admitted source passage
- [32] · page 4 — Source passage: Admitted source passage
- [33] · page 9 — Source passage: Admitted source passage