Back to Research papers
Research paper index

LLMs Can See the Smoke but not the Fire: Evaluating Abductive Reasoning with Elenchos

Julius Steiglechner, Lucas Mahler, Gabriele Lohmann

arXiv:2607.12733Published July 14, 20260 citations
  • cs.AI
  • cs.LG

Abstract

Large language models (LLMs) excel at pattern recognition and text generation, but their capacity for abductive inference - inferring latent hypotheses that explain observed behavior - remains poorly understood. Here, we introduce Elenchos (named after the Socratic method of cross-examination), a generative evaluation framework that measures abductive reasoning as a structural inverse problem. Given a reference formal system, such as the lambda-calculus, and a potentially mutated counterpart, agents must determine whether a mutation has occurred and infer the rule modifications responsible for the resulting behavioral differences. Evaluating frontier and mid-tier LLMs reveals a consistent detection-attribution dissociation: models often recognize that a system has been altered but struggle to identify the latent mutations causing the observed discrepancies. Performance degrades substantially under interacting mutations, where models frequently recover only a subset of the underlying mutations. Preliminary evidence also suggests diminishing returns from increased inference-time reasoning, with only modest improvements under larger reasoning budgets, though this finding requires further validation.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.