Interpretable Heuristic Priors Accelerate Deep Reinforcement Learning with Theory Guarantees
Main Article Content
Abstract
Deep reinforcement learning (RL) can excel on benchmarks but often wastes samples and trains unstably. We introduce a simple, theory-compatible way to add domain-agnostic prior structure to deep RL: (i) a Heuristic Policy prior (HP) that softly biases action selection using measurable strategic sub-scores, and (ii) a Nyāya-consistent potential (HR) for potential-based reward shaping that preserves optimal policies. Concretely, actions are sampled from while rewards are shaped via By evaluating on 10 tasks spanning Atari-RAM, grid/planning, Procgen, and Box2D, with n = 10 seeds per method, using baselines including DQN, A2C, PPO, SAC, PBRS-PPO, and RND-PPO. Aggregated across tasks, our HP/HR agents achieve a 29.4% median reduction in steps to reach 80% of each baseline’s asymptotic performance (the point where the moving-average curve over the final 10% of evaluations reaches 80% of its plateau), improve median final performance by 6.7%, and reduce run-to-run variance by 42.6%, all at ~5% training-time overhead. After Holm–Bonferroni correction, 9 out of 10 tasks show statistically significant sample-efficiency gains. Because HP and HR decompose into named components and predicates, decisions and shaped returns are auditable and abatable by design. These results demonstrate that lightweight, interpretable heuristic structure can reliably accelerate learning and stabilize training without sacrificing optimal control.