Self-Healing Networks Using Reinforcement Learning: An Adaptive Framework for Autonomous Fault Detection and Recovery in SD-WAN Environments

Authors

  • Chetan Padshala

Keywords:

Self-Healing Networks, Reinforcement Learning, Deep Q-Network, SD-WAN, Autonomous Networking, Fault Recovery, Network Resilience.

Abstract

Software-defined wide area networks have simplified enterprise connectivity through centralized control and programmable policy, yet their growing scale and dependence on controllers and overlay tunnels have made fault management a first-order concern. Recovery mechanisms that rely on static thresholds and human intervention respond slowly and adapt poorly to conditions that were not anticipated when the rules were written. This paper proposes a self-healing framework that applies reinforcement learning to autonomous fault detection, diagnosis, and recovery in software-defined wide area networks. The framework couples continuous telemetry-driven monitoring with a Deep Q-Network agent that learns recovery policies by interacting with the network through the controller, and it organizes the process along the monitor, analyze, plan, execute, and knowledge stages of the autonomic loop. A reward that balances restored availability against recovery delay and packet loss steers the agent toward remediations that shorten disruption. Rather than assert measured gains from a single deployment, the framework is evaluated analytically: a parametric model links the recovery time an autonomous agent can achieve to network availability and annual downtime, showing how sub-minute remediation moves availability into a regime that manual recovery cannot reach. The analysis positions reinforcement learning as a credible foundation for autonomous, self-healing networks, subject to the safety and explainability constraints that the paper discusses.

Downloads

Download data is not yet available.

References

S. A. Fayaz, S. J. Sidiq, M. Zaman, and M. A. Butt, "Machine learning: An introduction to reinforcement learning," in Machine Learning and Data Science: Fundamentals and Applications, 2022, pp. 1–22, doi: 10.1002/9781119776499.ch1.

V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, et al., "Human-level control through deep reinforcement learning," Nature, vol. 518, no. 7540, pp. 529–533, 2015, doi: 10.1038/nature14236.

J. P. G. Sterbenz, D. Hutchison, E. K. Çetinkaya, A. Jabbar, J. P. Rohrer, M. Schöller, et al., "Resilience and survivability in communication networks: Strategies, principles, and survey of disciplines," Computer Networks, vol. 54, no. 8, pp. 1245–1265, 2010, doi: 10.1016/j.comnet.2010.03.005.

P. C. Fonseca and E. S. Mota, "A survey on fault management in software-defined networks," IEEE Commun. Surveys Tuts., vol. 19, no. 4, pp. 2284–2321, 2017, doi: 10.1109/COMST.2017.2719862.

L. Ochoa-Aday, C. Cervelló-Pastor, and A. Fernández-Fernández, "Self-healing and SDN: Bridging the gap," Digital Commun. Netw., vol. 6, no. 3, pp. 354–368, 2020, doi: 10.1016/j.dcan.2019.08.008.

N. C. Luong, D. T. Hoang, S. Gong, D. Niyato, P. Wang, Y.-C. Liang, et al., "Applications of deep reinforcement learning in communications and networking: A survey," IEEE Commun. Surveys Tuts., vol. 21, no. 4, pp. 3133–3174, 2019, doi: 10.1109/COMST.2019.2916583.

Y. Xiao, J. Liu, J. Wu, and N. Ansari, "Leveraging deep reinforcement learning for traffic engineering: A survey," IEEE Commun. Surveys Tuts., vol. 23, no. 4, pp. 2064–2097, 2021, doi: 10.1109/COMST.2021.3102580.

S. Wang, J. F. Balarezo, S. Kandeepan, A. Al-Hourani, K. Gomez Chavez, and B. Rubinstein, "Machine learning in network anomaly detection: A survey," IEEE Access, vol. 9, pp. 152379–152396, 2021, doi: 10.1109/ACCESS.2021.3126834.

A. Leivadeas and M. Falkner, "A survey on intent-based networking," IEEE Commun. Surveys Tuts., vol. 25, no. 1, pp. 625–655, 2022, doi: 10.1109/COMST.2022.3215919.

E. Coronado, R. Behravesh, T. Subramanya, A. Fernandez-Fernandez, M. S. Siddiqui, X. Costa-Pérez, et al., "Zero touch management: A survey of network automation solutions for 5G and 6G networks," IEEE Commun. Surveys Tuts., vol. 24, no. 4, pp. 2535–2578, 2022, doi: 10.1109/COMST.2022.3212586.

S. Troia, F. Sapienza, L. Varé, and G. Maier, "On deep reinforcement learning for traffic engineering in SD-WAN," IEEE J. Sel. Areas Commun., vol. 39, no. 7, pp. 2198–2212, 2020, doi: 10.1109/JSAC.2020.3041385.

J. Zhang, M. Ye, Z. Guo, C.-Y. Yen, and H. J. Chao, "CFR-RL: Traffic engineering with reinforcement learning in SDN," IEEE J. Sel. Areas Commun., vol. 38, no. 10, pp. 2249–2259, 2020, doi: 10.1109/JSAC.2020.3000371.

J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, "Proximal policy optimization algorithms," arXiv preprint arXiv:1707.06347, 2017, doi: 10.48550/arXiv.1707.06347.

A. Halbouni, T. S. Gunawan, M. H. Habaebi, M. Halbouni, M. Kartiwi, and R. Ahmad, "CNN-LSTM: Hybrid deep neural network for network intrusion detection system," IEEE Access, vol. 10, pp. 99837–99849, 2022, doi: 10.1109/ACCESS.2022.3206425.

A. Menaceur, H. Drid, and M. Rahouti, "Fault tolerance and failure recovery techniques in software-defined networking: A comprehensive approach," J. Netw. Syst. Manage., vol. 31, no. 4, p. 83, 2023, doi: 10.1007/s10922-023-09772-x.

Z. Zhu, H. Yu, Q. Liu, D. Liu, and B. Mei, "FFRLI: Fast fault recovery scheme based on link importance for data plane in SDN," Comput. Netw., vol. 237, p. 110062, 2023, doi: 10.1016/j.comnet.2023.110062.

T. Hu, P. Yi, J. Lan, Y. Hu, and P. Sun, "FTLink: Efficient and flexible link fault tolerance scheme for data plane in software-defined networking," Future Gener. Comput. Syst., vol. 111, pp. 381–400, 2020, doi: 10.1016/j.future.2019.11.015.

Y. Ma, Y. Guo, R. Yang, and H. Luo, "FRRL: A reinforcement learning approach for link failure recovery in a hybrid SDN," J. Netw. Comput. Appl., vol. 234, p. 104054, 2025, doi: 10.1016/j.jnca.2024.104054.

J. Li, X. Qi, W. Ma, and L. Liu, "Path selection for link failure protection in hybrid SDNs," Future Gener. Comput. Syst., vol. 137, pp. 201–215, 2022, doi: 10.1016/j.future.2022.07.016.

V. Balasubramanian, M. Aloqaily, and M. Reisslein, "Fed-TSN: Joint failure probability-based federated learning for fault-tolerant time-sensitive networks," IEEE Trans. Netw. Service Manage., vol. 20, no. 2, pp. 1470–1486, 2023, doi: 10.1109/TNSM.2023.3273396.

J. A. Hurtado Sánchez, K. Casilimas, and O. M. Caicedo Rendon, "Deep reinforcement learning for resource management on network slicing: A survey," Sensors, vol. 22, no. 8, p. 3031, 2022, doi: 10.3390/s22083031.

Downloads

Published

11.08.2026

How to Cite

Chetan Padshala. (2026). Self-Healing Networks Using Reinforcement Learning: An Adaptive Framework for Autonomous Fault Detection and Recovery in SD-WAN Environments. International Journal of Intelligent Systems and Applications in Engineering, 14(1s), 2221–2228. Retrieved from https://www.ijisae.org/index.php/IJISAE/article/view/8508

Issue

Section

Research Article