Back to Research papers
Research paper index

Evaluating and Improving LLM Self-Modeling

Siqi Zeng, Andre N. Assis, Rowan Wang

arXiv:2608.30980Published August 31, 20260 citations
  • cs.CL
  • cs.AI

Abstract

We study self-modeling: an LLM's ability to answer questions about its own behavior. We focus on verifiable behavioral questions, such as whether a prompt edit would change the model's final answer. To measure this capability, we introduce a benchmark that tests diverse types of self-modeling questions. Current models show non-trivial but limited self-modeling skill, and make systematic mistakes on simple counterfactual questions about their own behavior. To improve self-modeling skill, we develop a scalable synthetic-data pipeline that produces self-modeling training data, and show that reinforcement-learning can improve aggregate self-modeling skill across three open-source model families with some transfer to held-out tasks. These gains, however, do not seem to constitute introspection consistently: improved self-modeling may not arise from privileged access to the model's internal decision process.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.