Segmenting Action-Value Functions over Time Scales in SARSA via TD(Δ)
In numerous episodic reinforcement learning (RL) environments, SARSA-based methodologies are employed to enhance policies aimed at maximizing returns over long horizons. Traditional SARSA algorithms face challenges in achieving an optimal balance between bias and variation, primarily due to their de...
Shranjeno v:
| Principais autores: | , , , , , , , , , , |
|---|---|
| Format: | Artigo |
| Jezik: | Inglês |
| Izdano: |
MDPI AG
2025-11-01
|
| Serija: | Algorithms |
| Teme: | |
| Online dostop: | https://www.mdpi.com/1999-4893/18/11/729 |
| Oznake: |
Brez oznak, prvi označite!
|
