Back to Research papers
Research paper index

Inside the Latent Flow: Causal Deciphering of Attention Dynamics in Audio Separation Foundation Models

Yuxuan Chen, Haoyuan Yu, Peize He

arXiv:2606.10046Published June 8, 2026Updated June 10, 20260 citations
  • cs.SD
  • cs.AI
  • foundation model

Abstract

Flow-matching transformers achieve strong audio separation, yet their attention dynamics are opaque. We adapt established causal-intervention principles into a deterministic, inference-time probing protocol for SAM Audio. Orthogonal probing uncovers a dual-pathway text-conditioning mechanism: additive injections control semantic identity, while cross-attention refines acoustic structure. We observe an asynchronous layerwise convergence: stable layers build temporal scaffolds early, whereas fast layers continue resolving artifacts during sampling. The model also attenuates temporal segmentation cues to maintain continuous-flow stability. Using these insights, we propose Layer-Selective Attention Caching (LSAC), a training-free acceleration method that caches attention in stable layers. Across acoustic complexities, LSAC cuts self-attention computation by about ~25% with negligible quality loss and yields up to 6.7x higher quality retention than naive step reduction.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.