Thompson Sampling for Stochastic Bandits with Noisy Contexts: An Information-Theoretic Regret Analysis
We study stochastic linear contextual bandits (CB) where the agent observes a <i>noisy</i> version of the true context through a noise channel with unknown channel parameters. Our objective is to design an action policy that can “approximate” that of a Bayesian oracle that has access to the reward m...
সংরক্ষণ করুন:
| প্রধান লেখক: | , |
|---|---|
| বিন্যাস: | Artigo |
| ভাষা: | Inglês |
| প্রকাশিত: |
MDPI AG
2024-07-01
|
| মালা: | Entropy |
| বিষয়গুলি: | |
| অনলাইন ব্যবহার করুন: | https://www.mdpi.com/1099-4300/26/7/606 |
| ট্যাগগুলো: |
কোনো ট্যাগ নেই, প্রথমজন হিসাবে ট্যাগ করুন!
|
