Dimensionality's Blessing: Clustering Images by Underlying Distributions

Page view(s)

Checked on Aug 10, 2025

Please use this identifier to cite or link to this item: https://oar.a-star.edu.sg/communities-collections/articles/14499

Title:

Dimensionality's Blessing: Clustering Images by Underlying Distributions

Journal Title:

Computer Vision and Pattern Recognition (CVPR) 2018

DOI:

Publication URL:

https://arxiv.org/abs/1804.02624

Authors:

Wen-Yan Lin, Siying Liu, Yasuyuki Matsushita

Keywords:

high-dimensional space, clustering

Publication Date:

01 June 2018

Citation:

Abstract:

Many high dimensional vector distances tend to a constant. This is typically considered a negative "contrast-loss" phenomenon that hinders clustering and other machine learning techniques. We reinterpret "contrast-loss" as a blessing. Re-deriving "contrast-loss" using the law of large numbers, we show it results in a distribution's instances concentrating on a thin "hyper-shell". The hollow center means apparently chaotically overlapping distributions are actually intrinsically separable. We use this to develop distribution-clustering, an elegant algorithm for grouping of data points by their (unknown) underlying distribution. Distribution-clustering, creates notably clean clusters from raw unlabeled data, estimates the number of clusters for itself and is inherently robust to "outliers" which form their own clusters. This enables trawling for patterns in unorganized data and may be the key to enabling machine intelligence.

License type:

PublisherCopyrights

Funding Info:

Description:

URI:

https://oar.a-star.edu.sg/communities-collections/articles/14499

ISBN:

Collections:

Institute for Infocomm Research

Files uploaded:

Manuscripts in This Item:

File	Size	Format	Action
distribution-clustering.pdf	1.22 MB	PDF	Open