ADVersa: Abductive Driving Accident Video Understanding

Page view(s)
0
Checked on
ADVersa: Abductive Driving Accident Video Understanding
Title:
ADVersa: Abductive Driving Accident Video Understanding
Journal Title:
IEEE Transactions on Pattern Analysis and Machine Intelligence
Keywords:
Publication Date:
11 February 2026
Citation:
Li, L.-L., Fang, J., Xiao, J., Yu, H., Lv, C., Xue, J., Li, Z., & Chua, T.-S. (2026). ADVersa: Abductive Driving Accident Video Understanding. IEEE Transactions on Pattern Analysis and Machine Intelligence, 48(6), 6980–6998. https://doi.org/10.1109/tpami.2026.3663545
Abstract:
Understanding traffic accident scenes is a longstanding research for vision-based safe driving. It seeks to answer why accidents occur, how near-crash scenes develop, and what the key elements of an accident are. This research is challenging due to the scarcity and fragmentation of accident data, as well as the complex accident environments. To study this, we present a framework of Abductive Driving accident Video understanding (ADVersa), which infers a plausible visual and textual explanation for the absent near-crash scenes. ADVersa underscores three groups of tasks: 1) visual past recovery of near-crash scenes, 2) visual prediction of near-crash scenes, and 3) accident cause involved video synthesis. To support the study, we first contribute MM-AU, a novel dataset for Multi-Modal Accident video Understanding. MM-AU contains 11,727 in-the-wild driving accident videos with temporally aligned text descriptions, 2.23 million well-annotated object boxes, and 58,650 pairs of video based accident cause texts. We then propose an Abductive CLIP model and a Contrastive Graph Video Pre-training (CGVP) model which exploit relation-aware cross-modal semantic learning to drive spatially abductive and temporally abductive accident video diffusion. Extensive experiments verify the superiority of ADVersa to the state-of-the-art approaches on different tasks, i.e., historical near-crash video frame recovering, crashing video frame prediction, textual accident cause and category reasoning, normal-to-accident video synthesis and accident video editing. With these efforts, we hope this research can advance the progress on multimodal accident video understanding. The code, datasets, and more results are released at www.lotvsmmau.net.
License type:
Publisher Copyright
Funding Info:
There was no specific funding for the research done
Description:
© 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
ISSN:
0162-8828
2160-9292
1939-3539
Files uploaded: