Back to Research papers
Research paper index

Separating perception from reasoning in vision-language models: a model-free render ceiling for crystal structures

Can Polat, Mustafa Kurban, Erchin Serpedin, Hasan Kurban

arXiv:2609.00663Published September 1, 20260 citations
  • cs.CV
  • cond-mat.mtrl-sci
  • physics.chem-ph
  • quant-ph
  • action
  • vision-language

Abstract

Multimodal evaluations cannot say whether a vision-language model misread an image or misreasoned about it, because every existing method for separating the two places a second model in the loop. We introduce the render ceiling, a model-free reference for benchmarks built by rendering known objects: inverting the frozen cameras and re-solving cross-view correspondence recovers exactly the answer the images support. We prove the ceiling fails only through an enumerable set of projection coincidences and certify that set empty on 2,160 rendered crystal structures, so every point of a model's deficit belongs to the model. Across fourteen vision-language models, supplying exact geometry as text lifts every model yet closes under half the gap for thirteen, while a supervised vision model with no language component reads the same images at 0.8952, above every vision-language model. The instrument exposes extraction-stage fabrication that downstream accuracy would misattribute to reasoning, yields camera-placement rules for benchmark builders, and transfers to any benchmark with an invertible forward rendering.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.