Back to Research papers
Research paper index

Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs

Deniz Bayazit, Badr AlKhamissi, Antoine Bosselut

arXiv:2609.00155Published August 31, 20260 citations
  • cs.CL
  • cs.AI
  • cs.LG

Abstract

Latent language identification is often used to argue that multilingual language models route computation through language-specific states, such as English pivots. However, existing probes infer latent language from different signals, such as the geometry of hidden states or what can be decoded from intermediate representations. Since such claims shape conclusions about how models share and route information across languages, we ask whether these probes measure the same phenomenon or expose distinct aspects of multilingual computation. We study this question across model families, training regimes, domains, tasks, checkpoints, and up to 27 languages. We find that identification probes systematically disagree: the GMM-based representation probe, which draws evidence from hidden state geometry, shows earlier cross-lingual mixing, whereas decoding-based probes, which rely on output-space decodability, retain sharper language-specific and more English-biased signals. These differences track model multilinguality and training progression, but are comparatively stable across domains. Our results suggest a more cautious interpretation of latent language identification, where current probes expose different aspects of multilingual processing, rather than directly revealing a single internal lingua franca.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.