EPSegFZ: Efficient Point Cloud Semantic Segmentation for Few- and Zero-Shot Scenarios with Language Guidance

Page view(s)
0
Checked on
EPSegFZ: Efficient Point Cloud Semantic Segmentation for Few- and Zero-Shot Scenarios with Language Guidance
Title:
EPSegFZ: Efficient Point Cloud Semantic Segmentation for Few- and Zero-Shot Scenarios with Language Guidance
Journal Title:
Proceedings of the AAAI Conference on Artificial Intelligence
Keywords:
Publication Date:
17 March 2026
Citation:
Wang, J., Zhu, H., Guo, H., Mamun, A. A., Xiang, C., & Lee, T. H. (2026). EPSegFZ: Efficient Point Cloud Semantic Segmentation for Few- and Zero-Shot Scenarios with Language Guidance. Proceedings of the AAAI Conference on Artificial Intelligence, 40(12), 9885–9893. https://doi.org/10.1609/aaai.v40i12.37953
Abstract:
Recent approaches for few-shot 3D point cloud semantic segmentation typically require a two-stage learning process, i.e., a pre-training stage followed by a few-shot training stage. While effective, these methods face overreliance on pre-training, which hinders model flexibility and adaptability. Some models tried to avoid pre-training yet failed to capture ample information. In addition, current approaches focus on visual information in the support set and neglect or do not fully exploit other useful data, such as textual annotations. This inadequate utilization of support information impairs the performance of the model and restricts its zero-shot ability. To address these limitations, we present a novel pre-training-free network, named Efficient Point Cloud Semantic Segmentation for Few- and Zero-shot scenarios. Our EPSegFZ incorporates three key components. A Prototype-Enhanced Registers Attention (ProERA) module and a Dual Relative Positional Encoding (DRPE)-based cross-attention mechanism for improved feature extraction and accurate query-prototype correspondence construction without pre-training. A Language-Guided Prototype Embedding (LGPE) module that effectively leverages textual information from the support set to improve few-shot performance and enable zero-shot inference.Extensive experiments show that our method outperforms the state-of-the-art method by 5.68% and 3.82% on the S3DIS and ScanNet benchmarks, respectively.
License type:
Publisher Copyright
Funding Info:
This research / project is supported by the National Research Foundation - “Centre for Advanced Robotics Technology Inno vation (CARTIN)”
Grant Reference no. :

This research / project is supported by the National Research Foundation - National Robotics Programme (NRP) 2.0
Grant Reference no. :
Description:
This material may not be retransmitted or redistributed without permission in writing from The Association for the Advancement of Artificial Intelligence. Permission to use document is granted, provided that (1) the copyright notice appears in all copies and that both the copyright notice and this permission notice appear, (2) use of such documents is for personal use only, and will not be copied or posted on any network computer or broadcast in any media, and (3) no modifications of any documents are made.
ISSN:
2374-3468
2159-5399
Files uploaded:

File Size Format Action
epsegfz-aaai-2026-copy.pdf 1.83 MB PDF Request a copy