期刊文献+
共找到4篇文章
< 1 >
每页显示 20 50 100
基于随机集成网络-TD3的四足机器人步态学习方法 认领 引用
1
作者 朱晓庆 朱晓宇 +2 位作者 阮晓钢 南博睿 毕兰越 《北京工业大学学报》 CAS CSCD 北大核心 2026年第4期371-379,共9页
为解决四足机器人技能学习领域中双延迟深度确定性策略梯度(twin delayed deep deterministic policy gradient,TD3)算法中存在Q值低估导致价值估计不准确,从而出现学习效果恶化的问题,提出一种随机集成网络-TD3(randomized ensembled n... 为解决四足机器人技能学习领域中双延迟深度确定性策略梯度(twin delayed deep deterministic policy gradient,TD3)算法中存在Q值低估导致价值估计不准确,从而出现学习效果恶化的问题,提出一种随机集成网络-TD3(randomized ensembled network-TD3,RE-TD3)算法。首先,该算法集成多个Q值网络,并随机选取Q值网络进行评估,缓解价值估计不准确的问题,有效提高策略性能;其次,设计合适的奖励函数以正确引导四足机器人的步态学习任务;最后,设置仿真实验进行验证。实验结果表明,该算法能够使四足机器人学习到良好的运动步态,与TD3算法相比,奖励值提高了32%,机体稳定性提高了约67%,期望方向偏离量提高了60%。 展开更多
关键词 强化学习 四足机器人 双延迟深度确定性策略梯度(twin delayed deep deterministic policy gradient,TD3) 奖励函数 步态学习 集成网络
暂未订购 下载PDF
Three-degree-of-freedom motion posture stabilization control of platform based on DTW-LSTM-MATD3 under high and low frequency disturbances of ships 认领 引用
2
作者 Qin ZHANG Jingyi ZHOU +1 位作者 Bangping GU Xiong HU 《Journal of Zhejiang University-SCIENCE A》 SCIE EI CAS CSCD 2026年第3期246-261,共16页
In the complex and variable deep-sea environment,the compensation control of ship motion ensures the safety and efficiency of equipment installation and transportation in offshore wind farms.However,the ship motion po... In the complex and variable deep-sea environment,the compensation control of ship motion ensures the safety and efficiency of equipment installation and transportation in offshore wind farms.However,the ship motion posture compensation control system is severely affected by uncertainties,which significantly impact the accuracy of compensation control.In this paper,we propose a ship three-degree-of-freedom(3-DoF)motion posture stabilization control method based on the DTW-LSTM-MATD3 algorithm.We use the multi-agent twin delayed deep deterministic policy gradient(MATD3)to control a platform with six electric cylinders to achieve stable control.However,owing to random noise affecting the ship’s motion posture,we use a dynamic time warping(DTW)algorithm to distinguish between high-frequency noise and low-frequency tracking signals.Further,we embed a long short-term memory(LSTM)network into the MATD3 network to better align the Critic network’s training with the true Q-value.We use a combined reward function to enhance the agent’s exploration capability in complex dynamic environments.Finally,verification was conducted under sixth-level,abrupt sea conditions with high-frequency noise,as well as under real abrupt sea conditions,and a generalization test was also carried out.Simulation results show that the proposed DTW-LSTM-MATD3 method has great compensation control ability. 展开更多
关键词 Compensation control Multi-agent twin delayed deep deterministic policy gradient(MATD3)algorithm Dynamic time warping(DTW)algorithm Long short-term memory(LSTM)network
暂未订购 下载PDF
Efficient sensorimotor cues for training a glider to soar autonomously 认领 引用
3
作者 Siyuan ZHENG Jiachi ZHAO +2 位作者 Lifang ZENG Zhouhong WANG Jun LI 《Journal of Zhejiang University-SCIENCE A》 SCIE EI CAS CSCD 2026年第2期128-141,共14页
Migratory birds depend on the perception of atmospheric updraft for long-distance flight.To realize more efficient autonomous soaring in an unpowered glider,different strategies for using potential sensorimotor cues t... Migratory birds depend on the perception of atmospheric updraft for long-distance flight.To realize more efficient autonomous soaring in an unpowered glider,different strategies for using potential sensorimotor cues to achieve autonomous soaring efficiency were compared and optimized.A simulation framework of autonomous soaring for an unpowered glider was developed based on a reinforcement learning algorithm.The framework was composed of three models:an updraft environment model,the glider's dynamics and control model,and a reinforcement learning agent,which learns to harvest more energy in flight.Based on the simulation,effects of different combinations of 12 potential sensorimotor cues on soaring efficiency were studied.Firstly,the absence of one particular sensorimotor cue and the use of only a single valid cue in autonomous soaring were analyzed.The results showed that the vertical airflow velocity gradient(aw)and the wing-tip updraft velocity difference(τ)have advantages over the other cues.Secondly,strategies combining aw orτwith other cues were analyzed to achieve more effective autonomous soaring,and seven potentially effective combinations of sensorimotor cues were identified.The final results showed that,among the tested combinations,the combination of vertical airflow velocity(Vw)andτ,enables the most efficient autonomous soaring.This study identified a highly effective sensorimotor cue strategy to guide an intelligent glider to achieve long-distance autonomous soaring flight. 展开更多
关键词 Autonomous soaring Glider Reinforcement learning Twin delayed deep deterministic policy gradient(TD3) Sensorimotor cues
暂未订购 下载PDF
增强型深度强化学习方法应用于化工过程控制 认领 引用 被引量:3
4
作者 张佳鑫 董立春 《化工进展》 EI CAS CSCD 北大核心 2025年第10期5563-5569,共7页
深度强化学习(DRL)算法因其无须依赖历史数据和先验知识,仅通过环境与智能体的互动即可实现策略优化和自主学习,在工业过程控制领域表现出良好的应用前景。其中,基于双延迟深度确定性策略梯度(TD3)算法的控制策略可有效克服深度确定性... 深度强化学习(DRL)算法因其无须依赖历史数据和先验知识,仅通过环境与智能体的互动即可实现策略优化和自主学习,在工业过程控制领域表现出良好的应用前景。其中,基于双延迟深度确定性策略梯度(TD3)算法的控制策略可有效克服深度确定性策略梯度(DDPG)模型中Q值易被高估,导致次优策略和鲁棒性不佳的缺陷,成为目前最领先的基于深度强化学习的控制模型。然而,原始TD3方法在应用于具有较显著策略波动的工业过程控制时仍显示出局限性,特别是其Q值低估问题会导致模型控制性能不佳。为了解决这些限制,本文提出了一种适用于工业过程控制的增强型TD3控制模型(ETD3),该模型首先建立评估指标来判断行动者(Actor)网络参数的高估或低估情况,并根据评估结果调整输入到批评家(Critic)网络的损失函数。然后,通过替换原始TD3中的固定学习率为三角衰减周期学习率,以提升模型的训练收敛性和控制性能。本文最后通过将增强型TD3算法应用于工业天然气脱水过程的控制过程验证了其有效性。 展开更多
关键词 过程控制 深度强化学习 双延时深度确定性策略梯度 三角衰减周期
暂未订购 下载PDF
上一页 1 下一页 到第
在线咨询 使用帮助 返回顶部 意见反馈