Journal of Automotive Safety and Energy ›› 2025, Vol. 16 ›› Issue (2): 326-333.DOI: 10.3969/j.issn.1674-8484.2025.02.016
• Intelligent Driving and Intelligent Transportation • Previous Articles Next Articles
HU Zhilong1(
), PEI Xiaofei1,2,*(
), ZHOU Honglong1, WEI Weiran2
Received:2024-07-08
Revised:2024-09-24
Online:2025-04-30
Published:2025-04-22
CLC Number:
HU Zhilong, PEI Xiaofei, ZHOU Honglong, WEI Weiran. Risk-sensitive hierarchical reinforcement learning decision-making for autonomous vehicles[J]. Journal of Automotive Safety and Energy, 2025, 16(2): 326-333.
Add to citation manager EndNote|Ris|BibTeX
URL: https://www.journalase.com/EN/10.3969/j.issn.1674-8484.2025.02.016
| 参数名称 | 描述 | 参数值 |
|---|---|---|
| 隐藏层(Linear)参数 | 各层神经元数 | (256, 128) |
| 优化器 | 用于梯度下降 | Adam |
| 折扣系数 | 计算折扣奖励 | 0.995 |
| 网络学习率 | 策略梯度更新步长 | 0.000 1 |
| 激活函数 | 增加神经网络非线性 | Relu |
| 批量大小 | 批量梯度下降中样本数量 | 32 |
| 分布最大最小值 | 将Q值分布为一个区间 | [-100, 100] |
| 分布个数 | 将Q值所属区间等分 | 51 |
| 分位数数量 | 估计Q分布的分位数数量 | 200 |
| 置信水平 | 条件风险价值参数 | 0.7 |
| 参数名称 | 描述 | 参数值 |
|---|---|---|
| 隐藏层(Linear)参数 | 各层神经元数 | (256, 128) |
| 优化器 | 用于梯度下降 | Adam |
| 折扣系数 | 计算折扣奖励 | 0.995 |
| 网络学习率 | 策略梯度更新步长 | 0.000 1 |
| 激活函数 | 增加神经网络非线性 | Relu |
| 批量大小 | 批量梯度下降中样本数量 | 32 |
| 分布最大最小值 | 将Q值分布为一个区间 | [-100, 100] |
| 分布个数 | 将Q值所属区间等分 | 51 |
| 分位数数量 | 估计Q分布的分位数数量 | 200 |
| 置信水平 | 条件风险价值参数 | 0.7 |
| 场景 | 算法模型 | 安全性 得分 | 效率 得分 | 舒适性 得分 | 总分 |
|---|---|---|---|---|---|
| 汇入汇出 场景 | RainbowDQN | 16.41 | 8.35 | 16.64 | 41.40 |
| RainbowDQN-QR | 26.88 | 10.35 | 17.43 | 54.66 | |
| DSAC-T | 36.97 | 7.75 | 17.83 | 62.55 | |
| RainbowDQN-CVaR | 34.58 | 13.00 | 16.76 | 64.34 | |
| 交叉口 场景 | DSAC-T | 7.87 | 17.89 | 18.9 | 44.66 |
| RainbowDQN | 13.46 | 16.69 | 19.28 | 49.43 | |
| RainbowDQN-QR | 18.16 | 16.56 | 19.40 | 54.12 | |
| RainbowDQN-CVaR | 34.39 | 19.41 | 18.90 | 72.70 |
| 场景 | 算法模型 | 安全性 得分 | 效率 得分 | 舒适性 得分 | 总分 |
|---|---|---|---|---|---|
| 汇入汇出 场景 | RainbowDQN | 16.41 | 8.35 | 16.64 | 41.40 |
| RainbowDQN-QR | 26.88 | 10.35 | 17.43 | 54.66 | |
| DSAC-T | 36.97 | 7.75 | 17.83 | 62.55 | |
| RainbowDQN-CVaR | 34.58 | 13.00 | 16.76 | 64.34 | |
| 交叉口 场景 | DSAC-T | 7.87 | 17.89 | 18.9 | 44.66 |
| RainbowDQN | 13.46 | 16.69 | 19.28 | 49.43 | |
| RainbowDQN-QR | 18.16 | 16.56 | 19.40 | 54.12 | |
| RainbowDQN-CVaR | 34.39 | 19.41 | 18.90 | 72.70 |
| 场景 | 置信 水平,α | 安全性 得分 | 效率 得分 | 舒适性 得分 | 总分 |
|---|---|---|---|---|---|
| 汇入汇出 场景 | 0.6 | 22.32 | 10.15 | 17.37 | 49.84 |
| 0.7 | 34.58 | 13.00 | 16.76 | 64.34 | |
| 0.8 | 28.4 | 10.96 | 17.61 | 56.97 | |
| 0.9 | 24.41 | 9.99 | 16.74 | 51.14 | |
| 交叉口 场景 | 0.6 | 19.08 | 16.68 | 19.09 | 54.85 |
| 0.7 | 34.39 | 19.41 | 18.9 | 72.7 | |
| 0.8 | 22.16 | 17.14 | 18.95 | 58.25 | |
| 0.9 | 19.85 | 16.95 | 19.06 | 55.86 |
| 场景 | 置信 水平,α | 安全性 得分 | 效率 得分 | 舒适性 得分 | 总分 |
|---|---|---|---|---|---|
| 汇入汇出 场景 | 0.6 | 22.32 | 10.15 | 17.37 | 49.84 |
| 0.7 | 34.58 | 13.00 | 16.76 | 64.34 | |
| 0.8 | 28.4 | 10.96 | 17.61 | 56.97 | |
| 0.9 | 24.41 | 9.99 | 16.74 | 51.14 | |
| 交叉口 场景 | 0.6 | 19.08 | 16.68 | 19.09 | 54.85 |
| 0.7 | 34.39 | 19.41 | 18.9 | 72.7 | |
| 0.8 | 22.16 | 17.14 | 18.95 | 58.25 | |
| 0.9 | 19.85 | 16.95 | 19.06 | 55.86 |
| [1] | Bojarski M, Testa D, Dworakowski D, et al. End to end learning for self-driving car[J/OL]. (2016-04-25) https://doi.org/10.48550/arXiv.1604.07316. |
| [2] | 杨顺. 从虚拟到现实的智能车辆深度强化学习控制研究[D]. 长春: 吉林大学, 2019. |
| YANG Shun. Research on deep reinforcement learning control of intelligent vehicles from virtual to real[D]. Changchun: Jilin University, 2019. (in Chinese) | |
| [3] | 李伟东, 马草原, 史浩, 等. 基于分层强化学习的自动驾驶决策控制算法[J/OL]. 吉林大学学报(工学版), 2023: 1-8. (2023-12-19) https://doi.org/10.13229/j.cnki.jdxbgxb.20230891. |
| LI Weidong, MA Caoyuan, SHI Hao, et al. Autonomous driving decision-making control algorithm based on hierarchical reinforcement learning[J/OL]. J Jilin University (Engi Tech Edit), 2023: 1-8. (2023-12-19) https://doi.org/10.13229/j.cnki.jdxbgxb.20230891. (in Chinese) | |
| [4] | YANG Lan, HU Zhiqiang, WANG Liang, et al. Entire route eco-driving method for electric bus based on rule-based reinforcement learning[J]. Transport Res, Part E: Logist Transport Rev, 2024, 189: 1-24. |
| [5] | Min K, Kim H, Huh K. Deep distributional reinforcement learning based high-Level driving policy determination[J]. IEEE Transport Intel Vehi, 2019, 4(3): 416-424. |
| [6] | ZHOU Lun, WANG Ke, YU Huang, et al. Path planning of improved DQN based on quantile regression [C]// 2022 Int’l Conf Artif Intel Comput Info Tech (AICIT). IEEE, 2022: 1-4. |
| [7] | MA Xiaoteng, XIA Li, ZHOU Zhengyuan, et al. DSAC: Distributional soft actor critic for risk-sensitive reinforcement learning[J/OL]. (2020-04-30) https://doi.org/10.48550/arXiv.2004.14547. |
| [8] | Ugur Y, Tufan K, Kemal U. A new approach for tactical decision making in lane changing:Sample efficient deep Q learning with a safety feedback reward [C]// 2020 IEEE Intel Vehi Symp (IV). Las Vegas, NV, USA, 2020: 1156-1161. |
| [9] | Bouton M, Nakhaei A, Fujimura K, et al. Safe reinforcement learning with scene decomposition for navigating complex urban environments [C]// 2019 IEEE Intel Vehi Symp (IV). Paris, France, 2019: 1469-1476. |
| [10] | CHEN Dong, JIANG Longsheng, WANG Yue, et al. Autonomous driving using safe reinforcement learning by incorporating a regret-based human lane-changing decision model [C]// 2020 Ame Contr Conf (ACC). Denver, CO, USA, 2020: 4355-4361. |
| [11] | Baheri A, Nageshrao S, Tseng H, et al. Deep reinforcement learning with enhanced safety for autonomous highway driving [C]// 2020 IEEE Intel Vehi Symp (IV). Las Vegas, NV, USA, 2020: 1550-1555. |
| [12] | Bellemare M G, Dabney W, Munos R. A distributional perspective on reinforcement learning[C]// Int’l Conf Mach Learn. PMLR, 2017: 449-458. |
| [13] | Danial K, Carlos L, Martin L, et al. Risk-aware high-level decisions for automated driving at occluded intersections with reinforcement learning [C]// 2020 IEEE Intel Vehi Symp (IV). Las Vegas, NV, USA, 2020: 1205-1212. |
| [14] | WEN Lu, DUAN Jingliang, LI Shengbo, et al. Safe reinforcement learning for autonomous vehicles through parallel constrained policy optimization [C]// 2020 IEEE 23rd Int’l Conf Intel Transport Syst (ITSC). Rhodes, Greece, 2020: 1-7. |
| [15] | 詹吟霄, 刘潇, 梁军. 基于深度强化学习与风险矫正的智能车辆决策研究[J]. 汽车工程学报, 2023, 13(5): 656-667. |
| ZHAN Yinxiao, LIU Xiao, LIANG Jun. Research on intelligent vehicle decision-making based on deep reinforcement learning and risk correction[J]. China J Autom Engi, 2023, 13(5): 656-667. (in Chinese) | |
| [16] | Dabney W, Rowland M, Bellemare M, et al. Distributional reinforcement learning with quantile regression[C]// Proc AAAI Conf Artif Intel. New Orleans, USA, 2018: 2892-2901. |
| [17] | DUAN Jingliang, WANG Wenxuan, XIAO Liming, et al. DSAC-T: Distributional soft actor-critic with three refinements[J/OL]. (2023-10-09) https://doi.org/10.48550/arXiv.2310.05858. |
| [1] | WANG Yue, DUAN Hongwei, ZHONG Wei, YANG Lu, HE Lei, CHAI Fulai, SHI Xiaoyang. Path planning method for leader-follower multi-vehicle formation with integrating GoT-SAC [J]. Journal of Automotive Safety and Energy, 2026, 17(1): 122-129. |
| [2] | WU Hangzhe, JIAO Yizhou, LIU Yang, ZHONG Wei, WANG Shuihe, GUO Jinghua, ZHAO Jian. Predictive trajectory tracking control by a linear time-varying model for emergency collision avoidance of autonomous vehicles [J]. Journal of Automotive Safety and Energy, 2025, 16(6): 934-944. |
| [3] | ZHENG Xunjia, CAO Zeyi, CHEN Xing, LIU Hui, GAO Jianjie. Trajectory tracking control based on adaptive prediction time-domain MPC [J]. Journal of Automotive Safety and Energy, 2025, 16(5): 773-783. |
| [4] | PAN Yuheng, REN Chen, LU Weijia, LI Yang. DV-PointPillars 3D object detection model based on dual pooling attention mechanism and vertical feature fusion [J]. Journal of Automotive Safety and Energy, 2025, 16(5): 793-801. |
| [5] | HAN Yu, CHEN Zhixuan, WANG Yixuan, LI Chunjie, LEI Wei, JIAO Yanli, LIU Pan. Deep reinforcement learning-based strategy for freeway ramp metering [J]. Journal of Automotive Safety and Energy, 2025, 16(4): 587-597. |
| [6] | OUYANG Delin, QIU Yifan, WANG Yingchen, YANG Liang, MIN Haigen, WANG Wenjun, LI Guofa. End-to-end decision-making model for multi-task autonomous driving [J]. Journal of Automotive Safety and Energy, 2025, 16(4): 610-619. |
| [7] | LI Ziyuan, LIU Qiang, LI Dingli, LI Zilong. Blind spot traffic strategy for intelligent connected vehicles based on deep reinforcement learning [J]. Journal of Automotive Safety and Energy, 2025, 16(3): 470-477. |
| [8] | LI Guofa, OUYANG Delin, CHEN Chen, NIE Binging, ZHANG Wei, YU Huili, Liu Bin, ZHANG Qiang, WANG Wenjun, CHENG Bo, LI Shengbo. Review on driving risk monitoring and intervention technologies [J]. Journal of Automotive Safety and Energy, 2025, 16(2): 181-196. |
| [9] | YANG Junru, ZHENG Sifa, XU Shucai, TIAN Ye, SUN Jian, SUN Chuan, LI Haoran. Design and research of an automated parking evaluation tool based on the OnSite platform [J]. Journal of Automotive Safety and Energy, 2025, 16(2): 334-343. |
| [10] | ZHANG Fuchun, YIN Yanli, MA Yongjuan, XIAO Hangyang, CHEN Haixin, YU Kai. Ecological driving and hierarchical control of energy management for networked hybrid electric vehicle queues [J]. Journal of Automotive Safety and Energy, 2025, 16(1): 159-169. |
| [11] | ZHANG Jinxiu, YAN Caihong, REN Guizhou. Estimation on state of health of lithium battery based on Gaussian process quantile regression model [J]. Journal of Automotive Safety and Energy, 2024, 15(6): 886-894. |
| [12] | CAI Tianmao, KONG Weiwei, LUO Yugong, SHI Jia, JI Pengxiao, LI Congmin. Multi-vehicle cooperative control in ramp merging area based on MADDPG algorithm [J]. Journal of Automotive Safety and Energy, 2024, 15(6): 923-933. |
| [13] | CAO Liling, LIU Junli, JIN Shengye, CAO Shouqi, ZHOU Guofeng. Design of a remote multidimensional information real time interaction system for autonomous driving [J]. Journal of Automotive Safety and Energy, 2024, 15(6): 934-942. |
| [14] | LIU Yang, ZHAN Jiahao, LI Shen, LI Xiaopeng, CHEN Jun. Future of autonomous driving: Single autonomous driving and intelligent vehicle-infrastructure collaboration systems [J]. Journal of Automotive Safety and Energy, 2024, 15(5): 611-633. |
| [15] | QU Guangyue, YANG Lan, YUAN Meng, FANG Shan, LIU Songyan. A multimodal trajectory prediction method of pedestrians at signalized intersections for autonomous vehicles [J]. Journal of Automotive Safety and Energy, 2024, 15(5): 689-701. |
| Viewed | ||||||
|
Full text |
|
|||||
|
Abstract |
|
|||||