Xingrui Yu, Bo Han, and Ivor W. Tsang. 2024. USN: A Robust Imitation Learning Method against Diverse Action Noise. J. Artif. Int. Res. 79 (Apr 2024). https://doi.org/10.1613/jair.1.15819
Abstract:
Learning from imperfect demonstrations is a crucial challenge in imitation learning (IL).
Unlike existing works that still rely on the enormous effort of expert demonstrators, we
consider a more cost-effective option for obtaining a large number of demonstrations. That
is, hire annotators to label actions for existing image records in realistic scenarios. However,
action noise can occur when annotators are not domain experts or encounter confusing states.
In this work, we introduce two particular forms of action noise, i.e., state-independent and
state-dependent action noise. Previous IL methods fail to achieve expert-level performance
when the demonstrations contain action noise, especially the state-dependent action noise.
To mitigate the harmful effects of action noises, we propose a robust learning paradigm
called USN (U ncertainty-aware S ample-selection with N egative learning). The model first
estimates the predictive uncertainty for all demonstration data and then selects samples
with high loss based on the uncertainty measures. Finally, it updates the model parameters
with additional negative learning on the selected samples. Empirical results in Box2D
tasks and Atari games show that USN consistently improves the final rewards of behavioral
cloning, online imitation learning, and offline imitation learning methods under various
action noises. The ratio of significant improvements is up to 94.44%. Moreover, our method
scales to conditional imitation learning with real-world noisy commands in urban driving.
License type:
Attribution 4.0 International (CC BY 4.0)
Funding Info:
This research / project is supported by the National Research Foundation, Singapore, and Maritime and Port Authority of Singapore / Singapore Maritime Institute - Maritime Transformation Programme (Maritime Artificial Intelligence (AI) Research Programme)
Grant Reference no. : SMI-2022-MTP-06
Description:
For the published article, refer here: https://dl.acm.org/doi/10.1613/jair.1.15819