Back to Research papers
Research paper index

Does Playing it Safe Count as Faithfulness? Reassessing LVLM Hallucination Mitigation Methods

Mehrdad Fazli, Sina Mansouri, Mohit Marvania, Ziwei Zhu

arXiv:2609.01888Published September 1, 20260 citations
  • cs.CV
  • vision-language

Abstract

Recent inference-time hallucination mitigation methods for large vision-language models (LVLMs) report strong gains on hallucination benchmarks. However, it remains unclear whether lower hallucination scores reflect improved multimodal grounding or more conservative generation. We evaluate six mitigation methods across three LVLMs and four benchmarks, including hallucination-focused evaluation and the diverse capability benchmark MMStar. Our analysis reveals two consistent patterns. First, hallucination reduction is often coupled with reduced informativeness: methods that lower hallucination rates also reduce object recall, visual coverage, or response detailedness. Second, improvements on hallucination benchmarks do not reliably transfer to broader multimodal capabilities, with methods showing inconsistent or degraded performance on fine-grained perception and reasoning tasks. Our findings suggest that current evaluation protocols may overestimate progress by rewarding conservative generation. We argue that hallucination mitigation should be evaluated as a faithfulness--informativeness--capability trade-off rather than through hallucination scores alone.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.