On the Asymptotics of Self-Supervised Pre-training: Two-Stage M-Estimation and Representation Symmetry
arXiv stat.ML4w4 min read
arXiv:2603.27631v2 Announce Type: replace-cross Abstract: Self-supervised pre-training, where large corpora of unlabeled data are used to learn representations for downstream fine-tuning, has become a cornerstone of modern machine learning. While a growing body of theoretical work has begun to analyze this paradigm, existing bounds leave open the question of how sharp the current rates are, and whether they accurately capture the complex interaction between pre-training and fine-tuning. In this paper, we address this gap by developing an asymptotic theory of pre-training via two-stage M-estima
