Back to Research papers
Research paper index

CyrillicQA: The Influence of Phonetically Encoded Secret Language on LLM Performance

Erik Thureck

arXiv:2608.21462Published August 20, 20260 citations
  • cs.CL
  • cs.AI
  • cs.LG
  • action

Abstract

Due to the selection of their training data, large language models (LLMs) perform best on standard-language inputs from languages using the Latin alphabet with large speaker populations, while disadvantaging other language varieties. Nevertheless, they can also be a versatile tool for preserving precisely such endangered languages. But do they also possess the necessary creativity and capacity for abstraction to decode phonetically encoded language the same way humans do?

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.