Código QR (código de barras bidimensional)

Policy Return: A New Method for Reducing the Number of Experimental Trials in Deep Reinforcement Learning

Using the same algorithm and hyperparameter configurations, deep reinforcement learning (DRL) will derive drastically different results from multiple experimental trials, and most of these results are unsatisfactory. Because of the instability of the results, researchers have to perform many trials...

Fuld beskrivelse

Na minha lista:
Bibliografiske detaljer
Principais autores: Feng Liu, Shuling Dai, Yongjia Zhao
Format: Artigo
Sprog:Inglês
Udgivet: IEEE 2020-01-01
Serier:IEEE Access
Fag:
Online adgang:https://ieeexplore.ieee.org/document/9298771/
Tags: Tilføj Tag
Ingen Tags, Vær først til at tagge denne postø!