Back to Research papers
Research paper index

A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

Zijie Cheng, Xiang Li, Yang Peng, Zhihua Zhang

arXiv:2608.27313Published August 27, 20260 citations
  • stat.ML
  • cs.LG
  • reinforcement learning
  • action

Abstract

We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. The proof separates two stability mechanisms. A global comparison argument, based on the order monotonicity of reward cumulative distribution functions and the $W_\infty$ contraction of the distributional Bellman operator, brings an arbitrarily initialized iterate into a local neighborhood. Inside that neighborhood, we linearize the QTD mean field. Its Jacobian is a nonsingular $M$-matrix, and the associated positive semigroup permits a variance-sensitive martingale analysis. For stepsizes $α_t=c(t+1)^{-a}$ with $a\in(1/2,1)$, the leading last-iterate fluctuation is of order $\widetilde O\bigl(T^{-a/2}/\sqrt{1-γ}\bigr)$ and has no polynomial dependence on the number of quantiles. The deterministic transient and the required burn-in can still depend on the smallest Bellman-target density, which is of order $m^{-1}$ in the worst case. The result therefore distinguishes sharply between the local stochastic fluctuation and the global sample complexity.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.