Back to Research papers
Research paper index

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction

Hongjin Ji, Guoyang Xia, Luoyang Sun, Fangxiang Feng, Lei Ren

arXiv:2608.09448Published August 10, 2026Updated August 12, 20260 citations
  • cs.RO
  • cs.CV
  • action
  • manipulation
  • vision-language
  • robot
  • policy

Abstract

Test-time training (TTT) offers a lightweight way to adapt vision--language--action (VLA) policies from unlabeled deployment streams, but it remains difficult to use reliably in closed-loop manipulation. A shared adaptation space can mix incompatible task corrections, while an online update can alter subsequent actions before its consequences are known. We introduce a reliable TTT framework for VLA policies (VANE). VANE conditions prompt adaptation on the current vision--language context and learns from the future visual consequences of executed actions. Candidate updates are isolated from the live policy, evaluated on subsequent observations, and committed only when supported by future evidence, making adaptation selective and reversible. On SimplerEnv WidowX, VANE improves average success by $3.2$ percentage points over the corresponding TTT baseline. Results on Google Robot further show that deployment-time gains remain task- and embodiment-dependent. Together, these results demonstrate a constrained, evidence-based approach to adapting VLA policies during interaction.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.