QR կոդ

Thompson Sampling for Stochastic Bandits with Noisy Contexts: An Information-Theoretic Regret Analysis

We study stochastic linear contextual bandits (CB) where the agent observes a <i>noisy</i> version of the true context through a noise channel with unknown channel parameters. Our objective is to design an action policy that can “approximate” that of a Bayesian oracle that has access to the reward m...

Ամբողջական նկարագրություն

Պահպանված է:
Մատենագիտական մանրամասներ
Հիմնական հեղինակներ: Sharu Theresa Jose, Shana Moothedath
Ձևաչափ: Artigo
Լեզու:Inglês
Հրապարակվել է: MDPI AG 2024-07-01
Շարք:Entropy
Խորագրեր:
Առցանց հասանելիություն:https://www.mdpi.com/1099-4300/26/7/606
Ցուցիչներ: Ավելացրեք ցուցիչ
Չկան պիտակներ, Եղեք առաջինը, ով նշում է այս գրառումը!