Back to Research papers
Research paper index

Optimal Reward Shaping: Autonomous Car Parking Case Study

Emre Özkaya, Nicolas R. Gauger

arXiv:2607.23617Published July 26, 20260 citations
  • cs.LG
  • math.OC
  • policy
  • trajectory
  • reinforcement learning

Abstract

Designing effective reward functions for model-free reinforcement learning under non-holonomic constraints remains a persistent challenge, often resulting in severe local minima such as policy paralysis or over-conservative hazard avoidance. In this work, we present a parameterized reward shaping framework featuring coverage-gated alignment feedback, drive-direction switch regularization, and an aligned episode termination mechanism evaluated on an autonomous parallel parking task. Crucially, we show that environmental reward parameters and algorithmic hyperparameters are deeply co-dependent, requiring joint meta-optimization to achieve stable convergence. By employing surrogate-based Bayesian optimization, our co-optimized Deep Q-Network (DQN) agent resolves characteristic control failure modes, significantly outperforming uncalibrated baselines across both success rate and trajectory smoothness.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.