Data Science Wire

Contextual Slate GLM Bandits with Limited Adaptivity

arXiv stat.ML1mo4 min read

arXiv:2606.31449v1 Announce Type: cross Abstract: We investigate the contextual slate bandit problem with generalized linear rewards under limited adaptivity. At each round, the learner is presented with $N$ sets of items, where each item is represented by a $d$-dimensional feature vector. The learner then constructs a slate by selecting one item per set; the resulting slate yields a scalar reward sampled from a Generalized Linear Model (GLM). We propose algorithms under two limited-adaptivity settings: (a) Batched and (b) Rarely-Switching. For the batched setting, we introduce B-SlateGLinCB,

Read the full story at arXiv stat.ML

More in Data Science