Data Science Wire

Statistical Properties of $k$-means Clustering for Data Missing Completely at Random

arXiv stat.ML4w4 min read

arXiv:2607.01945v1 Announce Type: new Abstract: The classical $k$-means clustering cannot be directly used to incomplete data, and existing $k$-means-based clustering for missing data primarily focus on improving the practical accuracy of clustering, whereas most of them lack theoretical guarantees in the asymptotic sense. In this paper, we investigate the statistical properties of $k$-means clustering in the presence of missing data. We first establish the $\sqrt{n}$-excess risk bound and prove the consistency of the estimated cluster centers under general missing mechanisms. For the Missing

Read the full story at arXiv stat.ML

More in Data Science