Human skill knowledge guided global trajectory policy reinforcement learning method
Traditional trajectory learning methods based on Imitation Learning (IL) only learn the existing trajectory knowledge from human demonstration. In this way, it can not adapt the trajectory knowledge to the task environment by interacting with the environment and fine-tuning the policy. To address th...
保存先:
| 主要な著者: | , , , , , |
|---|---|
| フォーマット: | Artigo |
| 言語: | Inglês |
| 出版事項: |
Frontiers Media S.A.
2024-03-01
|
| シリーズ: | Frontiers in Neurorobotics |
| 主題: | |
| オンライン・アクセス: | https://www.frontiersin.org/articles/10.3389/fnbot.2024.1368243/full |
| タグ: |
タグなし, このレコードへの初めてのタグを付けませんか!
|
