An <i>ϵ</i>-Greedy Multiarmed Bandit Approach to Markov Decision Processes
We present REGA, a new adaptive-sampling-based algorithm for the control of finite-horizon Markov decision processes (MDPs) with very large state spaces and small action spaces. We apply a variant of the <inline-formula><math xmlns="http://www.w3.org/1998/Math/MathML" display="inline"><semantics><mi...
Na minha lista:
| Principais autores: | , |
|---|---|
| 格式: | Artigo |
| 語言: | Inglês |
| 出版: |
MDPI AG
2023-01-01
|
| 叢編: | Stats |
| 主題: | |
| 在線閱讀: | https://www.mdpi.com/2571-905X/6/1/6 |
| 標簽: |
沒有標簽, 成為第一個標記此記錄!
|
