QR-kod

An <i>ϵ</i>-Greedy Multiarmed Bandit Approach to Markov Decision Processes

We present REGA, a new adaptive-sampling-based algorithm for the control of finite-horizon Markov decision processes (MDPs) with very large state spaces and small action spaces. We apply a variant of the <inline-formula><math xmlns="http://www.w3.org/1998/Math/MathML" display="inline"><semantics><mi...

Full beskrivning

Sparad:
Bibliografiska uppgifter
Huvudupphov: Isa Muqattash, Jiaqiao Hu
Materialtyp: Artigo
Språk:Inglês
Utgiven: MDPI AG 2023-01-01
Serie:Stats
Ämnen:
Länkar:https://www.mdpi.com/2571-905X/6/1/6
Taggar: Lägg till en tagg
Inga taggar, Lägg till första taggen!