Back to Research papers
Research paper index

Sleight of Word Benchmark: Can Language Models Notice If Their Own Output Was Tampered With?

Alberto Cetoli

arXiv:2608.29921Published August 30, 20260 citations
  • cs.CL
  • cs.AI
  • action

Abstract

The output of a Language Model can be tampered with \emph{while} the model is writing it. A simple test can thus be constructed by evaluating the model's perception of this external perturbation. In this spirit, a simple benchmark is built in which a single word is consistently substituted with another in the generation process. We call this method \emph{Sleight of Word}. Two distinct axes are measured: metrics that relate to the model's surprise, as well as an evaluation of the textual reaction for 19 different open-weight language models.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.