Thompson Sampling for Stochastic Bandits with Noisy Contexts: An Information-Theoretic Regret Analysis
We study stochastic linear contextual bandits (CB) where the agent observes a <i>noisy</i> version of the true context through a noise channel with unknown channel parameters. Our objective is to design an action policy that can “approximate” that of a Bayesian oracle that has access to the reward m...
Պահպանված է:
| Հիմնական հեղինակներ: | , |
|---|---|
| Ձևաչափ: | Artigo |
| Լեզու: | Inglês |
| Հրապարակվել է: |
MDPI AG
2024-07-01
|
| Շարք: | Entropy |
| Խորագրեր: | |
| Առցանց հասանելիություն: | https://www.mdpi.com/1099-4300/26/7/606 |
| Ցուցիչներ: |
Չկան պիտակներ, Եղեք առաջինը, ով նշում է այս գրառումը!
|
