Back to Research papers
Research paper index

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement

Colombe Mboungou, Mostafa Sadeghi, Jean-Eudes Ayilo, Romain Serizel

arXiv:2606.23712Published June 16, 20260 citations
  • eess.SP
  • cs.AI

Abstract

Audio-visual speech enhancement (AVSE) exploits visual cues such as lip movements to recover speech in noisy environments. Recent work introduced diffusion-based unsupervised AVSE, where a speech diffusion model conditioned on visual features via cross-attention is trained and used as a data-driven prior for posterior sampling-based speech enhancement. Despite promising performance over its audio-only counterpart, the impact of explicitly enforcing cross-modal alignment in the fusion remains unclear. In this work, we propose to augment the diffusion training objective with a contrastive audio-visual loss to encourage stronger use of visual information while keeping the posterior sampling framework unchanged. Experiments across matched and mismatched test data show consistent improvements in interference suppression, signal reconstruction, and perceptual quality, with the largest gains at low SNRs. Code is available at https://github.com/ cexauce/AV-CA-DiffUSE

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.