Back to papers
March 19, 2026cs.LGAdvanced
Best-of-Both-Worlds Multi-Dueling Bandits: Unified Algorithms for Stochastic and Adversarial Preferences under Condorcet and Borda Objectives
AI-Generated Summary
This paper solves a key challenge in recommendation systems by developing algorithms that work optimally in both unpredictable and random environments without knowing which one they face. The authors propose two main solutions: MetaDueling for ranking by comparing pairs of items, and AlgBorda for aggregating preferences across multiple items, with theoretical guarantees showing these algorithms perform as well as specialized algorithms designed for each environment separately.
Difficulty
Advanced
Categories
cs.LG
AI Tags
multi-armed banditsonline learningranking and recommendationadversarial robustnessstochastic optimizationpreference learning