Back to Research papers
Research paper index

DoubleHelix: Structured Cross-Modal Fusion for Audio-Visual Speech Recognition with LLMs

Ziwei Cheng, Zhenhua Tan, Zhuomin Zhu

arXiv:2607.29112Published July 31, 20260 citations
  • cs.SD
  • cs.AI
  • action

Abstract

Audio-visual speech recognition (AVSR) relies on effective fusion of audio and visual modalities, yet existing approaches treat cross-modal interaction as a single-step operation without structured iterative refinement. We present DoubleHelix, a multimodal fusion framework that reformulates fusion as an iterative cross-modal interaction process with adaptive degradation-aware enhancement. The framework comprises three components including ReverseParallelHelix for multi-turn structured interaction with learned alignment constraints, QualitySensor for learning degradation-aware gating signals, and HelixReplication for consistency-guided conditional feature enhancement. Experiments on LRS3 demonstrate that DoubleHelix achieves 0.68% WER on clean audio, outperforming previous best results by 5.6% relative improvement under matched backbone settings. Comprehensive ablation studies validate each component contribution, including targeted analysis of design choices such as asymmetric pathway weighting. The framework shows improved robustness under evaluated babble-noise conditions, achieving 11.6% WER at SNR -5dB.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.