HVIS: A Human-like Vision and Inference System for Human Motion Prediction

Page view(s)
17
Checked on Sep 09, 2025
HVIS: A Human-like Vision and Inference System for Human Motion Prediction
Title:
HVIS: A Human-like Vision and Inference System for Human Motion Prediction
Journal Title:
Proceedings of the AAAI Conference on Artificial Intelligence
Keywords:
Publication Date:
11 April 2025
Citation:
Lyu, K., Chen, H., Liu, Z., Yin, Y., Lin, Y., & Jiao, Y. (2025). HVIS: A Human-like Vision and Inference System for Human Motion Prediction. Proceedings of the AAAI Conference on Artificial Intelligence, 39(6), 5928–5936. https://doi.org/10.1609/aaai.v39i6.32633
Abstract:
Grasping the intricacies of human motion, which involve perceiving spatio-temporal dependence and multi-scale effects, is essential for predicting human motion. While humans inherently possess the requisite skills to navigate this issue, it proves to be markedly more challenging for machines to emulate. To bridge the gap, we propose the Human-like Vision and Inference System (HVIS) for human motion prediction, which is designed to emulate human observation and forecast future movements. HVIS comprises two components: the human-like vision encode (HVE) module and the human-like motion inference (HMI) module. The HVE module mimics and refines the human visual process, incorporating a retina-analog component that captures spatiotemporal information separately to avoid unnecessary crosstalk. Additionally, a visual cortex-analogy component is designed to hierarchically extract and treat complex motion features, focusing on both global and local features of human poses. The HMI is employed to simulate the multi-stage learning model of the human brain. The spontaneous learning network simulates the neuronal fracture generation process for the adversarial generation of future motions. Subsequently, the deliberate learning network is optimized for hard-to-train joints to prevent misleading learning. Experimental results demonstrate that our method achieves new state-of-the-art performance, significantly outperforming existing methods by 19.8 % on Human3.6M, 15.7 % on CMU Mocap, and 11.1 % on G3D.
License type:
Publisher Copyright
Funding Info:
National Natural Science Foundation of China

the Key R&D Program of Zhejiang Province
Description:
This material may not be retransmitted or redistributed without permission in writing from The Association for the Advancement of Artificial Intelligence. Permission to use document is granted, provided that (1) the copyright notice appears in all copies and that both the copyright notice and this permission notice appear, (2) use of such documents is for personal use only, and will not be copied or posted on any network computer or broadcast in any media, and (3) no modifications of any documents are made.
ISSN:
2374-3468
Files uploaded:

File Size Format Action
hvis.pdf 2.24 MB PDF Open