Back to Research papers
Research paper index

LuxEmo: Expressive Text-to-Speech Corpus for Luxembourgish

Nina Hosseini-Kivanani, Sandipana Dowerah

arXiv:2606.31947Published June 30, 2026Updated July 2, 20260 citations
  • cs.CL

Abstract

State-of-the-art speech datasets predominantly focus on widely spoken languages, often overlooking low-resource languages such as Luxembourgish, which remain underrepresented in speech technology research. In this work, we introduce LuxEmo, a 21-hour conversational expressive speech corpus for Luxembourgish with 4 emotion categories. LuxEmo is derived from Radio Télévision Luxembourg (RTL) youth broadcasts, using automated detection followed by human validation. We propose a semi-automatic curation workflow combining voice activity detection, denoising, language identification, LuxASR-based segmentation, automatic emotion prediction, lexical cues, and targeted human review. Additionally, we benchmark five expressive TTS systems covering German-based cross-lingual transfer, multilingual Luxembourgish support, Luxembourgish adaptation, and non-parametric prosody transfer. Performance is evaluated using both objective metrics and human evaluation.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.