HIERARCHICAL DRL FOR CONTINUOUS CYBER RISK OPTIMIZATION IN ENTERPRISE NETWORKS

Authors

  • Rajasekhar Reddy Arikatla Author

Keywords:

Cyber Risks, Hierarchical DRL, Enterprise Network, Network Security.

Abstract

Large-scale enterprise networks are facing cyber risks that are driven by continuously evolving sophisticated attack strategies, increasing system complexity, and changing operational conditions. Traditional cyber risk management is predominantly based on static policies, periodical assessments, and rule-based controls, which cannot deal with real-time threats within large-scale enterprise settings. Deep reinforcement learning has been  emerging as a promising approach in adaptive cyber defense; however, flat DRL architectures face severe challenges when it comes to scalability, reward attribution delay, and interpretability for such complex, multi-layered network infrastructures.

This paper presents a Hierarchical Deep Reinforcement Learning framework for continuous cyber risk optimization in enterprise networks. In the proposed approach, cyber defense decision-making is decomposed into two abstraction levels: that is, a high-level policy making strategic risk management decisions and choosing long-term security postures, and low-level policies that implement tactical defense actions such as adjusting access control, containing intrusions, and deploying patches. This hierarchical decomposition enables the study of efficient learning across multiple time scales while being able to reduce state-action complexity.

A risk-aware reward function is proposed, specifically intended to optimize cyber risk reduction, as well as satisfy cost and service availability constraints in the operation of the system. Evaluation is conducted in various simulation platforms, mimicking the dynamics of an enterprise network environment and its assets. The evaluation demonstrates the superiority of the proposed HDRL framework as it is able to attain quicker convergence, reduce cumulative risk, and improve policy stability in comparison to flat DRL and rule-based approaches. Visualization of the learning curves also demonstrates the capability of hierarchical control in maintaining continuous risk optimization. It is therefore evident that the HDRL is an efficient and proper platform for developing intelligent cyber defense mechanisms in real-world networks.

Downloads

Published

2024-09-19