Back to Research papers
Research paper index

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models

Xingyu Ding, Yuzhong Zhao, Yang Wu, Chunhai Zhao, Chaoyang Zhao, Yifan Zhang, Jian Cheng

arXiv:2608.04633Published August 5, 2026Updated August 27, 20260 citations
  • cs.RO
  • action
  • manipulation
  • vision-language
  • robot

Abstract

Recent Vision-Language-Action (VLA) methods improve generalization by aligning their representations with 3D scene geometry. However, these methods are fundamentally instruction-agnostic: the representations align the entire scene uniformly, neglecting the 3D geometry of the specific target object designated by the language instruction. This causes failures on fine-grained manipulation and target occlusion tasks, where success depends on accurate 3D understanding of the target object rather than the entire scene. To address this, we present Mind-VLA, an instruction-aware spatial representation alignment method for VLA models. Specifically, Mind-VLA first obtains the target object specified by the language instruction, then prepares its canonical target views and extracts the corresponding VAE and VGGT features. Finally, the latent representation of the VLA model is aligned with these features to enable instruction-aware 3D understanding. Mind-VLA reaches 94.4% on LIBERO and 4.47 on CALVIN with a compact 345M-parameter backbone. On real-robot tasks with target occlusion, Mind-VLA reaches 54% average success, outperforming the matched scene-VGGT control by 26 percentage points.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.