Back to Research papers
Research paper index

Massive Open-Vocabulary Keyword Spotting

Leonor Barreiros, Raul Monteiro, Afonso Mendes, Gonçalo M. Correia

arXiv:2606.11279Published June 9, 20260 citations
  • eess.AS
  • cs.CL
  • cs.LG
  • cs.SD

Abstract

Automatic speech recognition systems have been shown to under-perform when it comes to transcribing words rarely seen in the training data, namely specialized terminology. Open-vocabulary keyword spotting, combined with contextual biasing, has been shown to mitigate this issue. However, existing systems can only handle glossaries of a few hundred terms without becoming an infeasible bottleneck. We propose a system that stores features with a memory footprint up to 128 times smaller than a comparable baseline and allows users to process massive databases while remaining open-vocabulary. Without fine-tuning the speech recognition model, our system achieves a comparable entity recall as uncompressed solutions, even in languages not seen during training.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.