Back to Research papers
Research paper index

Think, Look, and Revise: Inconsistency-Aware Visual Self-Correction in MLLMs

Yu Cheng, Arushi Goel, Hakan Bilen

arXiv:2608.29374Published August 29, 20260 citations
  • cs.CV
  • reinforcement learning

Abstract

Tool-augmented multimodal reasoning integrates external tools (e.g., object detection, depth estimation) into multimodal large language models (MLLMs) to address perceptual bottlenecks in complex visual tasks. However, existing approaches rarely verify tool outputs, limiting their ability to detect and recover from tool failures. We propose ReVISE, a framework that equips MLLMs with verification and dynamic error recovery for tool-augmented reasoning. ReVISE introduces (1) a curated training dataset that supervises reflective behaviors, enabling models to validate tool-derived evidence, reformulate queries when visual mismatches arise, and fall back to intrinsic grounding when external tools are unreliable; and (2) a reinforcement learning based targeted rewards that encourage internal reflection and penalize spatial misalignment. Experiments on several benchmarks demonstrate consistent improvements over existing methods, highlighting the importance of error detection and correction in tool-augmented multimodal reasoning.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.