Back to Research papers
Research paper index

Learning from Annotation Uncertainty: Entropy-Aware Curriculum for Speech Emotion Recognition

Zahra Omidi, John H. L. Hansen

arXiv:2606.27536Published June 25, 20260 citations
  • cs.SD
  • cs.LG

Abstract

Speech emotion recognition (SER) often relies on hard consensus labels that collapse annotator disagreement. We study distribution-based supervision for 9-class SER on MSP-Podcast 2.0 using a WavLM-Base multitask model for categorical emotion and dimensional VAD. Hard-label training is compared with targets from primary and merged primary--secondary annotator vote distributions. Distributional objectives improve alignment with human vote distributions, reducing JSD/KLD relative to hard-label training. Analysis shows that hard supervision partly benefits from assigning ambiguous utterances to the residual Other class, whereas distributional supervision redistributes uncertainty across emotion categories. Entropy-stratified evaluation shows that high-ambiguity utterances remain challenging, but distribution-based supervision better captures perceptual uncertainty. These findings support moving beyond hard labels toward targets that reflect listener disagreement.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.