QR Code

Policy Return: A New Method for Reducing the Number of Experimental Trials in Deep Reinforcement Learning

Using the same algorithm and hyperparameter configurations, deep reinforcement learning (DRL) will derive drastically different results from multiple experimental trials, and most of these results are unsatisfactory. Because of the instability of the results, researchers have to perform many trials...

Whakaahuatanga katoa

I tiakina i:
Ngā taipitopito rārangi puna kōrero
Ngā kaituhi matua: Feng Liu, Shuling Dai, Yongjia Zhao
Hōputu: Artigo
Reo:Inglês
I whakaputaina: IEEE 2020-01-01
Rangatū:IEEE Access
Ngā marau:
Urunga tuihono:https://ieeexplore.ieee.org/document/9298771/
Ngā Tūtohu: Tāpirihia he Tūtohu
Kāore He Tūtohu, Me noho koe te mea tuatahi ki te tūtohu i tēnei pūkete!