Back to Research papers
Research paper index

Context-Aware Multimodal Claim Verification in Spoken Dialogues

Chaewan Chun, Delvin Ce Zhang, Dongwon Lee

arXiv:2606.11420Published June 9, 20260 citations
  • cs.CL
  • cs.SI

Abstract

Every day, millions absorb claims from podcasts and streams that no fact-checker ever sees. Spoken misinformation is built through conversation, where credibility comes not from facts alone but from how claims are framed, reinforced, or left unchallenged across turns. Yet fact-checking has focused on isolated text, leaving dialogue audio under-studied. We introduce MAD2, a new Multi-turn Audio Dialogues benchmark for spoken claim verification, containing 1,000 two-speaker dialogues with 3,368 check-worthy claims and approximately 10 hours of audio, and propose calibrated multimodal fusion of a context-aware audio encoder and a dialogue-aware text model. Across settings, adding dialogue context improves verification, but the gains depend on scenario type. Using only preceding context often matches offline performance, supporting live-moderation settings, and audio contributes most when transcript-based models are destabilized by additional context. Overall, conversational structure matters more for verification than misinformation framing.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.