Back to Research papers
Research paper index

Early Warning Signals for OpenVLA Failure under Visual Distribution Shift

Dipesh Tharu Mahato, Rachel Ren

arXiv:2606.29699Published June 29, 2026Updated August 13, 20260 citations
  • cs.CV
  • cs.AI
  • cs.RO
  • action
  • policy
  • vision-language

Abstract

Visual shifts can cause a vision-language-action policy to fail after initially plausible behavior. We ask whether OpenVLA's internal activations contain signals associated with the steps before failure. We freeze the policy, record one MLP activation per LIBERO-10 step, and fit two linear monitors. Occlusion reduces task success from $57\%$ to $17\%$. Within failed matched-reset trajectories, a layer-16 logistic probe attains AUROC $0.972$ and AUPRC $0.352$, whereas action disagreement attains AUROC $0.496$. Without refitting, the occlusion-trained probe reaches AUROC $0.689$ on failed camera-jitter episodes. In a calibration check, however, the same layer-16 monitor averages 3.32 warning onsets per clean episode. This contrast shows that strong retrospective discrimination does not imply operationally quiet warning behavior. Because fitting and evaluation share tasks, resets, and seed, these results establish retrospective separability rather than prediction on independent episodes.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.