Back to Research papers
Research paper index

Phoneme- vs. Character-Level Targets and Selective State-Space Models for Intracortical Brain-to-Text

Lucas Zamora Vera, Jose A. Gonzalez-Lopez

arXiv:2607.26751Published July 29, 20260 citations
  • cs.CL
  • cs.AI
  • eess.SP

Abstract

State-of-the-art intracortical brain-to-text systems pair a neural-sequence phone decoder with an external language model. Two design axes remain underexplored: whether selective state-space models (Mamba) improve on recurrent decoders, and how the output target (phonetic vs.\ character) interacts with that choice. On the public Brain-to-Text '25 benchmark, we study a controlled 2x2 grid (GRU vs.\ hybrid Mamba decoder; phonetic vs.\ character targets) trained with a CTC objective under one reproducible protocol. The recurrent baseline remains strongest: the best phonetic GRU reaches 12.62\% PER and 21.19\% WER, while the best textual GRU after LM rescoring reaches 13.39\% CER and 26.28\% WER. The Mamba hybrid is competitive but does not surpass it. Ablations isolate architectural contributions, and error analysis shows representation-dependent failures: articulatory-like phoneme confusions vs.\ lexical and word-boundary errors.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.