Back to Research papers
Research paper index

Evaluation of forced alignment of code-mixed speech: the case of Hindi-English

Ayushi Pandey, Pamir Gogoi, Kevin Tang

arXiv:2607.25581Published July 28, 20260 citations
  • cs.CL

Abstract

Code-mixed speech poses unique challenges to forced alignment: expanded inventories, orthographic errors, and speaker variation. We evaluate forced alignment of Hindi-English code-mixed speech using the Montreal Forced Aligner. We address 2 problems: (1) free variation involving native vs non-native pairs and (2) phonemic boundary detection for mid-utterance English words. Bootstrapping strategies substantially outperform unmodified lexicons. Acoustic models trained on sentence-level code-mixed data achieve a mean error of 4.15ms, ie. ten times lower than monolingual Hindi (38.18ms) or isolated English (37.58ms) alternatives. Principled lexicon design and code-mixed training data are both essential for reliable alignment of bilingual speech.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.