Dynamic Interaction Dilation for Interactive Human Parsing

Page view(s)
63
Checked on Aug 05, 2025
Dynamic Interaction Dilation for Interactive Human Parsing
Title:
Dynamic Interaction Dilation for Interactive Human Parsing
Journal Title:
IEEE Transactions on Multimedia
Publication Date:
29 March 2023
Citation:
Gao, Y., Lang, C., Liu, F., Cao, Y., Sun, L., & Wei, Y. (2023). Dynamic Interaction Dilation for Interactive Human Parsing. IEEE Transactions on Multimedia, 1–12. https://doi.org/10.1109/tmm.2023.3262973
Abstract:
Interactive segmentation pursues generating high-quality pixel-level predictions with a few user-provided clicks, which is gaining attention for its convenience in segmentation data annotation. Users are allowed to iteratively refine the prediction by adding clicks until the result is satisfactory. Existing interactive methods usually transform the clicks into a set of localization maps by Euclidian distance computation or RGB texture extraction to guide the segmentation, which makes the click transformation a core module in interactive segmentation networks. However, when adopted in human images where large poses, occlusions, and bad illuminations are prevailing, prior transformation methods tend to cause uncorrectable overlapping across localization maps, i.e. , one click corresponds to multiple transformed values at the same position in different localization map channels, which are difficult to form a good match among human parts and limit the interaction efficiency. Furthermore, the inappropriately transformed information is hard to be refined with the static transformation manner, i.e. , based on the fixed formulas / RGB textures, which is out of tune with the dynamically refined interaction process. Hence, we design a dynamic transformation scheme for interactive human parsing (IHP) named Dynamic Interaction Dilation Net ( DID-Net ), which serves as an initial attempt to break the limitations of static transformation while capturing long-range dependencies of clicks within each human part. Specifically, we construct a Dynamic Dilation Module ( DD-Module ) to dilate clicks radially in several directions assisted by human body edge detection. The continually refined edges guide to improve the dilation quality in each interaction iteration, thereby better fitting user intention. Furthermore, we propose an Adaptive Interaction Excitation Block ( AIE-Block ) to exploit potential semantic clues buried in the dilated clicks and emphasize semantic expression for each human part by feature recalibration. Our DID-Net achieves state-of-the-art performance on 3 public human parsing benchmarks.
License type:
Publisher Copyright
Funding Info:
This research / project is supported by the A*STAR - AME Programmatic Funds
Grant Reference no. : A20H6b0151

This work is supported by the National Natural Science Foundation of China under Grant (No.62072027), the National Key Research and Development Program of China (No.2022ZD0118502) and the Major Projects of National Natural Science Foundation of China (No.72293583).
Description:
© 2023 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
ISSN:
1941-0077
1520-9210
Files uploaded:

File Size Format Action
final-version-tmm-removehead.pdf 11.65 MB PDF Open