Back to Research papers
Research paper index

ProgVLA: Progress-Aware Robot Manipulation Skill Learning

Seungsu Kim, Jinyoung Choi, Seungmin Baek, Jean-Michel Renders

arXiv:2605.28231Published May 27, 20260 citations
  • cs.RO
  • cs.LG
  • policy
  • reinforcement learning
  • robot
  • imitation learning
  • action
  • manipulation
  • vision-language

Abstract

We present ProgVLA, a compact vision-language-action (VLA) model designed for reliable robot manipulation under tight compute and memory budgets. The model specifically focuses on efficiently processing long multi-modal sequences by maintaining an explicit representation of task progress over extended horizons. To this end, ProgVLA integrates two key components. First, a multi-modal encoder with a two-stage Perceiver resampling scheme compresses variable-length visual, language, and proprioceptive streams into a fixed set of control-ready context tokens, substantially reducing sequence length while preserving cross-modal grounding. Second, an auxiliary set of progress heads is trained with offline reinforcement learning (RL) objectives to jointly learn critics over normalized remaining-horizon targets. This provides the policy with an internal estimate of task progress and enables advantage- and success-weighted flow-matching imitation learning. On two well-established multi-task robot manipulation benchmarks, a 0.1B-parameter ProgVLA model reaches success rates that are competitive with, and on long-horizon and harder task tiers exceed, substantially larger pretrained baselines. Ablations indicate that the learned context resampler and task-adaptive visual fine-tuning are the largest single contributors, while progress-aware training provides a consistent additional gain that is concentrated on long-horizon and multi-object tasks. We further validate the approach in real-world toy-kitchen environments.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.