HPRS: hierarchical potential-based reward shaping from task specifications
The automatic synthesis of policies for robotics systems through reinforcement learning relies upon, and is intimately guided by, a reward signal. Consequently, this signal should faithfully reflect the designer’s intentions, which are often expressed as a collection of high-level requirements. Seve...
Đã lưu trong:
| Những tác giả chính: | , , , |
|---|---|
| Định dạng: | Artigo |
| Ngôn ngữ: | Inglês |
| Được phát hành: |
Frontiers Media S.A.
2025-02-01
|
| Loạt: | Frontiers in Robotics and AI |
| Những chủ đề: | |
| Truy cập trực tuyến: | https://www.frontiersin.org/articles/10.3389/frobt.2024.1444188/full |
| Các nhãn: |
Không có thẻ, Là người đầu tiên thẻ bản ghi này!
|
