Back to Research papers
Research paper index

Learning to Refine Hidden States for Reliable LLM Reasoning

Chia-Hsuan Hsu, Jui-Ming Yao

arXiv:2606.17524Published June 16, 2026Updated June 25, 20260 citations
  • cs.LG
  • action
  • policy

Abstract

Large language models show strong reasoning ability, but their internal reasoning process can remain unstable in complex multi-step settings, where early hidden-state errors may propagate to incorrect predictions. We propose ReLAR, a reinforcement-guided latent refinement framework that iteratively updates hidden representations before decoding. ReLAR maintains a compact latent reasoning state and uses learned depth and action controllers to adaptively determine both the number and direction of refinement steps. The controllers are trained with a policy gradient objective based on step-wise likelihood improvement, enabling efficient input-dependent reasoning without explicit chain-of-thought generation. Experiments on medical, mathematical, multi-hop reasoning, and open-ended generation benchmarks show that ReLAR improves accuracy, generation quality, and reasoning stability with substantially lower inference overhead than explicit reasoning baselines.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.