Back to Research papers
Research paper index

Slow to See, Slow to Suppress: Understanding the Effects of Modality in Context-Memory Conflicts

Athulith Paraselli, Etha Tianze Hua, Ellie Pavlick

arXiv:2609.00293Published August 31, 20260 citations
  • cs.CL
  • vision-language

Abstract

We investigate how vision-language models (VLMs) handle context-memory conflicts; that is, situations in which the model is given information in context that differs from what was stored parametrically during training. We document asymmetric biases: models tend to prefer in-context information about entities which appear in text, but prefer parametric information about entities which appear in images. We relate this asymmetry to the late representational alignment across modalities, showing that the longer processing time associated with resolving visual entities prevents the suppression of the model's usual factual recall mechanism, thus resulting in more parametric answers. Chain-of-thought reasoning does not appear to resolve the gap, but increasing the amount of visual information in the context does show an effect. These results illustrate the complexity of ensuring consistent behavior as models become increasingly multimodal and retrieval-augmented.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.