An <i>ϵ</i>-Greedy Multiarmed Bandit Approach to Markov Decision Processes
We present REGA, a new adaptive-sampling-based algorithm for the control of finite-horizon Markov decision processes (MDPs) with very large state spaces and small action spaces. We apply a variant of the <inline-formula><math xmlns="http://www.w3.org/1998/Math/MathML" display="inline"><semantics><mi...
Sparad:
| Huvudupphov: | , |
|---|---|
| Materialtyp: | Artigo |
| Språk: | Inglês |
| Utgiven: |
MDPI AG
2023-01-01
|
| Serie: | Stats |
| Ämnen: | |
| Länkar: | https://www.mdpi.com/2571-905X/6/1/6 |
| Taggar: |
Inga taggar, Lägg till första taggen!
|
