COMPLIANCE-PRESERVING REWARD SHAPING FOR REINFORCEMENT LEARNING AGENTS

Authors

  • Rajesh Thonduru Author

Keywords:

Reinforcement Learning, Artificial Intelligence,

Abstract

Reinforcement learning often provides a natural fit within the landscape of the enterprise's decision systems, as it can be used to better optimize a sequence of actions within a constantly changing environment. Traditionally, the focus for a typical RL-based agent is on an optimization plan that attempts to maximize a certain "reward," generally not including considerations for regulations or governance within an organization.

The authors propose the use of the proposed framework, known as the 'compliance preserver reward shaping framework.' This framework incorporates the constraints associated with 'policy compliance' right into the reinforcement learning signal. This method avoids the use of 'compliance filters' and is also efficient for validation. Instead, the framework looks to directly incorporate hints for policy compliance, guiding the policy toward compliant behavior, not only in the learning scenario but also in the real-world application scenario.

Our simulation is informed by scenarios that could arise in an enterprise domain, e.g., access control, cloud resource management, automated approval, etc. Our results demonstrate that compliance-aware RS significantly are improving the policy adherence while keeping task performance top-notch. Simple visualizations facilitate easy interpretation of convergence trends, trade-offs between compliance and performance, and how compliance improves over time.

On the whole, the results show that reward shaping can be a potent method for ensuring alignment between reinforcement learning agents and the need for enterprise compliance.

 

Downloads

Published

2025-05-20

Issue

Section

Articles