Data Science Wire

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits

arXiv cs.LG5d4 min read

arXiv:2607.22012v1 Announce Type: new Abstract: Off-Policy Evaluation and Learning (OPE/L) in contextual bandits is rapidly gaining popularity in real systems because new policies can be evaluated and learned securely using only historical logged data. However, existing methods in OPE/L cannot handle many challenging but prevalent scenarios such as few-shot data, deterministic logging policies, and new actions. In many applications, such as personalized medicine, content recommendations, education, and advertising, we need to evaluate and learn new policies in the presence of these challenges.

Read the full story at arXiv cs.LG

More in Machine Learning