Stability and Generalization of Straight-Through Estimators for Training Two-Layer Quantized Neural Networks
arXiv stat.ML6d4 min read
arXiv:2609.06430v1 Announce Type: cross Abstract: We study the identity straight-through estimator (STE) for training a two-layer binary-activation network with hinge loss from the perspective of Statistical Learning Theory (SLT). Our central question is whether algorithmic stability can explain the statistical generalization of the estimator produced by the discontinuous STE training rule. In the saturated-output regime, the zero-initialized samplewise STE recursion is exactly the stochastic subgradient descent on the convex latent loss $(-yu^\top x)_+$. This representation makes a stability
