Interpretable Heuristic Priors Accelerate Deep Reinforcement Learning with Theory Guarantees

Main Article Content

Kajol Arora, Dr. Vijay Kumar

Abstract

Deep reinforcement learning (RL) can excel on benchmarks but often wastes samples and trains unstably. We introduce a simple, theory-compatible way to add domain-agnostic prior structure to deep RL: (i) a Heuristic Policy prior (HP) that softly biases action selection using measurable strategic sub-scores, and (ii) a Nyāya-consistent potential (HR) for potential-based reward shaping that preserves optimal policies. Concretely, actions are sampled from  while rewards are shaped via  By evaluating on 10 tasks spanning Atari-RAM, grid/planning, Procgen, and Box2D, with n = 10 seeds per method, using baselines including DQN, A2C, PPO, SAC, PBRS-PPO, and RND-PPO. Aggregated across tasks, our HP/HR agents achieve a 29.4% median reduction in steps to reach 80% of each baseline’s asymptotic performance (the point where the moving-average curve over the final 10% of evaluations reaches 80% of its plateau), improve median final performance by 6.7%, and reduce run-to-run variance by 42.6%, all at ~5% training-time overhead. After Holm–Bonferroni correction, 9 out of 10 tasks show statistically significant sample-efficiency gains. Because HP and HR decompose into named components and predicates, decisions and shaped returns are auditable and abatable by design. These results demonstrate that lightweight, interpretable heuristic structure can reliably accelerate learning and stabilize training without sacrificing optimal control.

Article Details

How to Cite
Kajol Arora, Dr. Vijay Kumar. (2026). Interpretable Heuristic Priors Accelerate Deep Reinforcement Learning with Theory Guarantees. Journal of Daoist Studies, 19(S4), 117–125. Retrieved from https://journalofdaoiststudies.org/index.php/journal/article/view/622
Section
Articles