Back to Research papers
Research paper index

HindSearch: Trajectory-Level Hindsight Critique for Search-Augmented Reinforcement Learning

Haowei Liu, Jiamian Wang, Hsin-Tai Wu, Zhiqiang Tao, Yi Fang

arXiv:2608.01597Published August 3, 20260 citations
  • cs.LG
  • cs.AI
  • cs.IR
  • reinforcement learning
  • trajectory
  • policy
  • action

Abstract

Search-augmented LM agents are typically trained with a binary exact-match reward, which throws away most of what a failed trajectory tells us about why it failed. We introduce HindSearch, a hindsight self-distillation procedure for GRPO: after each rollout, a frozen judge writes a short critique of every failed trajectory using the gold answer, and the critique supplies an auxiliary on-policy distillation signal on the student's search actions. On the standard seven-benchmark suite with Qwen2.5-3B-Instruct, HindSearch reaches 39.4% average EM, outperforming prior search-RL baselines. Removing the judge's access to the gold answer erases most of the gain, isolating hindsight as the source of the improvement.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.