TSS GAZ PTP: Towards Improving Gumbel AlphaZero with Two-Stage Self-Play for Multi-Constrained Electric Vehicle Routing Problems
Deep reinforcement learning (DRL) with self-play has emerged as a promising paradigm for solving combinatorial optimization (CO) problems. The recently proposed Gumbel AlphaZero Plan-to-Play (GAZ PTP) framework adopts a competitive training setup between a learning agent and an opponent to tackle cl...
محفوظ في:
| المؤلفون الرئيسيون: | , , |
|---|---|
| التنسيق: | Artigo |
| اللغة: | Inglês |
| منشور في: |
MDPI AG
2026-01-01
|
| سلاسل: | Smart Cities |
| الموضوعات: | |
| الوصول للمادة أونلاين: | https://www.mdpi.com/2624-6511/9/2/21 |
| الوسوم: |
لا توجد وسوم, كن أول من يضع وسما على هذه التسجيلة!
|
