QR Kod

Policy Return: A New Method for Reducing the Number of Experimental Trials in Deep Reinforcement Learning

Using the same algorithm and hyperparameter configurations, deep reinforcement learning (DRL) will derive drastically different results from multiple experimental trials, and most of these results are unsatisfactory. Because of the instability of the results, researchers have to perform many trials...

Ful tanımlama

Kaydedildi:
Detaylı Bibliyografya
Asıl Yazarlar: Feng Liu, Shuling Dai, Yongjia Zhao
Materyal Türü: Artigo
Dil:Inglês
Baskı/Yayın Bilgisi: IEEE 2020-01-01
Seri Bilgileri:IEEE Access
Konular:
Online Erişim:https://ieeexplore.ieee.org/document/9298771/
Etiketler: Etiketle
Etiket eklenmemiş, İlk siz ekleyin!