Optimized Test Scenario Identification Method based on Refined Warshall's Dynamic Programming for Validating Reinforcement Learning Model

Citations

SCOPUS

0

초록

In recent decades, Reinforcement learning (RL) has achieved remarkable performance in sequential decision-making tasks; however, validating RL-based software remains a significant challenge due to the exponential growth in the number of state–action combinations. We propose an optimized test-scenario identification method for a formal validation mechanism that integrates a refined Warshall's dynamic programming algorithm and path optimization to ensure efficient test coverage for RL systems. The method constructs a directed abstract graph from the RL model, applies transitive closure analysis to check reachability between states, and identifies missing states/transitions before testing. Our goal is to use test-scenario optimization to generate a minimal yet sufficient set of test cases, thereby achieving maximum coverage with minimal redundancy. This approach reduces verification complexity while maintaining mathematical rigor, making it well-suited for safety-critical applications such as autonomous driving. The proposed mechanism provides a scalable, interpretable validation process, offering a foundation for the reliable deployment of RL–based software in real-world systems. © 2026 KSII.

키워드

AI Software ValidationReinforcement LearningTest Scenario Optimization
제목
Optimized Test Scenario Identification Method based on Refined Warshall's Dynamic Programming for Validating Reinforcement Learning Model
저자
Kim, JanghwanKim, R. Young Chul
DOI
10.3837/tiis.2026.04.030
발행일
2026-04-30
유형
Article
저널명
KSII Transactions on Internet and Information Systems
20
4
페이지
2224 ~ 2241