We introduce State Vector Space Partitioning (SVSP), a novel method to mimic a black-box reinforcement learning policy using a set of human-interpretable sub-policies. By partitioning a distillation dataset of state–action pairs with linear support vector machine splits, SVSP constructs a compact and structured represen- tation of the original policy where linear models can be interpreted as a measure of feature importance. Our method improves mean return by +7.4\% over previous critic-driven state partitioning attempts such as Voronoi State Partitioning (VSP) and +2.8\% over the original TD3 policy, while reducing the number of required sub-policies against VSP by 82.1\%. Our results pave the path towards a more flexi- ble form of distillation where both the decision boundary and surrogate models can be chosen within a margin of the original black box behavior.
Deproost, S, Nowe, A & Asadi, M 2026, Hierarchical Support Vector State Partitioning for Distilling Black Box Reinforcement Learning Policies. in M Alimardani, T Lenaerts, A Meyer-Vitali, A Nowe, J Vennekens & S Wang (eds), HHAI 2026 Proceedings of the 5th International Conference on Hybrid Human-Artificial Intelligence. vol. 423, Frontiers in Artificial Intelligence and Applications, vol. 423, IOS Press, pp. 520-523, The 5th International Conference on Hybrid Human-Artificial Intelligence, Brussel, Belgium, 8/07/26. https://doi.org/10.3233/FAIA260549
Deproost, S., Nowe, A., & Asadi, M. (2026). Hierarchical Support Vector State Partitioning for Distilling Black Box Reinforcement Learning Policies. In M. Alimardani, T. Lenaerts, A. Meyer-Vitali, A. Nowe, J. Vennekens, & S. Wang (Eds.), HHAI 2026 Proceedings of the 5th International Conference on Hybrid Human-Artificial Intelligence (Vol. 423, pp. 520-523). (Frontiers in Artificial Intelligence and Applications; Vol. 423). IOS Press. https://doi.org/10.3233/FAIA260549
@inproceedings{bd09a5f502704084818067f04565c98f,
title = "Hierarchical Support Vector State Partitioning for Distilling Black Box Reinforcement Learning Policies",
abstract = "We introduce State Vector Space Partitioning (SVSP), a novel method to mimic a black-box reinforcement learning policy using a set of human-interpretable sub-policies. By partitioning a distillation dataset of state–action pairs with linear support vector machine splits, SVSP constructs a compact and structured represen- tation of the original policy where linear models can be interpreted as a measure of feature importance. Our method improves mean return by +7.4\% over previous critic-driven state partitioning attempts such as Voronoi State Partitioning (VSP) and +2.8\% over the original TD3 policy, while reducing the number of required sub-policies against VSP by 82.1\%. Our results pave the path towards a more flexi- ble form of distillation where both the decision boundary and surrogate models can be chosen within a margin of the original black box behavior.",
keywords = "Reinforcement Learning, Explainable AI",
author = "Senne Deproost and Ann Nowe and Mehrdad Asadi",
note = "Publisher Copyright: {\textcopyright} 2026 The Authors.; The 5th International Conference on Hybrid Human-Artificial Intelligence, HHAI'26 ; Conference date: 08-07-2026 Through 10-07-2026",
year = "2026",
month = jul,
day = "6",
doi = "10.3233/FAIA260549",
language = "English",
volume = "423",
series = "Frontiers in Artificial Intelligence and Applications",
publisher = "IOS Press",
pages = "520--523",
editor = "Maryam Alimardani and Tom Lenaerts and Andre Meyer-Vitali and Ann Nowe and Joost Vennekens and Shenghui Wang",
booktitle = "HHAI 2026 Proceedings of the 5th International Conference on Hybrid Human-Artificial Intelligence",
address = "Netherlands",
url = "https://hhai-conference.org/2026/",
}