Back to Research papers
Research paper index

PureTD: Reinforcement Learning for Backgammon Money Games with No Evaluation-time Search

Alexander L. Strehl

arXiv:2608.15146Published August 15, 20260 citations
  • cs.LG
  • reinforcement learning
  • action

Abstract

We revisit Tesauro's TD-Gammon for backgammon money games in the setting of no evaluation-time search. Both checker play and cube action (use of the doubling cube) are learned from scratch via self-play reinforcement learning (RL), with minimal hand-coded logic and no expert features. In this setting, we demonstrate that pure self-play RL suffices to train models that reach near-state-of-the-art playing strength. Specifically, for cubeful money games, our search-free model evaluates faster and is substantially stronger than the open-source engines GNU Backgammon and Open Sage running a one-move (1-ply) look-ahead search.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.