期刊文献+
共找到471篇文章
< 1 2 24 >
每页显示 20 50 100
Multi-agent deep deterministic policy gradient algorithm for predictive UAV deployment in CF-mMIMO for identifying coverage holes 认领 引用
1
作者 Kenneth Okello Elijah Mwangi Dominic Bernard Onyango Konditi 《Journal of Electronic Science and Technology》 EI CAS CSCD 2026年第2期91-108,共18页
Unmanned aerial vehicles(UAVs),due to their adaptable mobility and various applications,including supporting communication infrastructure,monitoring,and rescue,are becoming increasingly valuable,making them a valuable... Unmanned aerial vehicles(UAVs),due to their adaptable mobility and various applications,including supporting communication infrastructure,monitoring,and rescue,are becoming increasingly valuable,making them a valuable addition to emergency communication networks.Even though cell-free massive multiple-input multiple-output(CF-mMIMO)networks provide high communication data rates,their immobility makes it difficult to maintain quality network continuity in emergency,unpredictable,and congested areas where users’equipment is located.To mitigate this challenge,the integration of aerial access points(AAPs)into CF-mMIMO networks is proposed by using the multi-agent deep deterministic policy gradient(MADDPG)framework,which teaches several UAVs to jointly learn the best deployment plans by estimating user distributions and traffic demand trends on invitations to provide tremendous dynamic coverage,increased spectral efficiency(SE),and throughput maximization.The predictive component framework utilizes a long short-term memory(LSTM)network model incorporating concepts of learning,association,movement,and service provision for temporal traffic forecasting,thereby ensuring proactive UAV positioning before coverage holes emerge.Our extensive simulation results demonstrate that the MADDPG-based throughput deployment strategy achieves approximately 45 Gbps for 50 UAVs,the SE for downlink and uplink of 10.2 bps/Hz and 15.2 bps/Hz,respectively,and the minimal transmit power of 3.5 kJ as compared with the multi-agent soft actorcritic(MASAC)method,traditional heuristic-LSTM,and single-agent reinforcement learning approaches. 展开更多
关键词 Aerial access points Cell-free massive multiple-input multiple-output Long short-term memory Multi-agent deep deterministic policy gradient Unmanned aerial vehicles
暂未订购 下载PDF
A Dynamic Deceptive Defense Framework for Zero-Day Attacks in IIoT:Integrating Stackelberg Game and Multi-Agent Distributed Deep Deterministic Policy Gradient 认领 引用
2
作者 Shigen Shen Xiaojun Ji Yimeng Liu 《Computers, Materials & Continua》 SCIE EI 2025年第11期3997-4021,共25页
The Industrial Internet of Things(IIoT)is increasingly vulnerable to sophisticated cyber threats,particularly zero-day attacks that exploit unknown vulnerabilities and evade traditional security measures.To address th... The Industrial Internet of Things(IIoT)is increasingly vulnerable to sophisticated cyber threats,particularly zero-day attacks that exploit unknown vulnerabilities and evade traditional security measures.To address this critical challenge,this paper proposes a dynamic defense framework named Zero-day-aware Stackelberg Game-based Multi-Agent Distributed Deep Deterministic Policy Gradient(ZSG-MAD3PG).The framework integrates Stackelberg game modeling with the Multi-Agent Distributed Deep Deterministic Policy Gradient(MAD3PG)algorithm and incorporates defensive deception(DD)strategies to achieve adaptive and efficient protection.While conventional methods typically incur considerable resource overhead and exhibit higher latency due to static or rigid defensive mechanisms,the proposed ZSG-MAD3PG framework mitigates these limitations through multi-stage game modeling and adaptive learning,enabling more efficient resource utilization and faster response times.The Stackelberg-based architecture allows defenders to dynamically optimize packet sampling strategies,while attackers adjust their tactics to reach rapid equilibrium.Furthermore,dynamic deception techniques reduce the time required for the concealment of attacks and the overall system burden.A lightweight behavioral fingerprinting detection mechanism further enhances real-time zero-day attack identification within industrial device clusters.ZSG-MAD3PG demonstrates higher true positive rates(TPR)and lower false alarm rates(FAR)compared to existing methods,while also achieving improved latency,resource efficiency,and stealth adaptability in IIoT zero-day defense scenarios. 展开更多
关键词 Industrial internet of things zero-day attacks Stackelberg game distributed deep deterministic policy gradient defensive spoofing dynamic defense
暂未订购 下载PDF
Perception Enhanced Deep Deterministic Policy Gradient for Autonomous Driving in Complex Scenarios 认领 引用
3
作者 Lyuchao Liao Hankun Xiao +3 位作者 Pengqi Xing Zhenhua Gan Youpeng He Jiajun Wang 《Computer Modeling in Engineering & Sciences》 SCIE EI 2024年第7期557-576,共20页
Autonomous driving has witnessed rapid advancement;however,ensuring safe and efficient driving in intricate scenarios remains a critical challenge.In particular,traffic roundabouts bring a set of challenges to autonom... Autonomous driving has witnessed rapid advancement;however,ensuring safe and efficient driving in intricate scenarios remains a critical challenge.In particular,traffic roundabouts bring a set of challenges to autonomous driving due to the unpredictable entry and exit of vehicles,susceptibility to traffic flow bottlenecks,and imperfect data in perceiving environmental information,rendering them a vital issue in the practical application of autonomous driving.To address the traffic challenges,this work focused on complex roundabouts with multi-lane and proposed a Perception EnhancedDeepDeterministic Policy Gradient(PE-DDPG)for AutonomousDriving in the Roundabouts.Specifically,themodel incorporates an enhanced variational autoencoder featuring an integrated spatial attention mechanism alongside the Deep Deterministic Policy Gradient framework,enhancing the vehicle’s capability to comprehend complex roundabout environments and make decisions.Furthermore,the PE-DDPG model combines a dynamic path optimization strategy for roundabout scenarios,effectively mitigating traffic bottlenecks and augmenting throughput efficiency.Extensive experiments were conducted with the collaborative simulation platform of CARLA and SUMO,and the experimental results show that the proposed PE-DDPG outperforms the baseline methods in terms of the convergence capacity of the training process,the smoothness of driving and the traffic efficiency with diverse traffic flow patterns and penetration rates of autonomous vehicles(AVs).Generally,the proposed PE-DDPGmodel could be employed for autonomous driving in complex scenarios with imperfect data. 展开更多
关键词 Autonomous driving traffic roundabouts deep deterministic policy gradient spatial attention mechanisms
暂未订购 下载PDF
Optimizing the Multi-Objective Discrete Particle Swarm Optimization Algorithm by Deep Deterministic Policy Gradient Algorithm 认领 引用
4
作者 Sun Yang-Yang Yao Jun-Ping +2 位作者 Li Xiao-Jun Fan Shou-Xiang Wang Zi-Wei 《Journal on Artificial Intelligence》 2022年第1期27-35,共9页
Deep deterministic policy gradient(DDPG)has been proved to be effective in optimizing particle swarm optimization(PSO),but whether DDPG can optimize multi-objective discrete particle swarm optimization(MODPSO)remains ... Deep deterministic policy gradient(DDPG)has been proved to be effective in optimizing particle swarm optimization(PSO),but whether DDPG can optimize multi-objective discrete particle swarm optimization(MODPSO)remains to be determined.The present work aims to probe into this topic.Experiments showed that the DDPG can not only quickly improve the convergence speed of MODPSO,but also overcome the problem of local optimal solution that MODPSO may suffer.The research findings are of great significance for the theoretical research and application of MODPSO. 展开更多
关键词 Deep deterministic policy gradient multi-objective discrete particle swarm optimization deep reinforcement learning machine learning
暂未订购 下载PDF
Noise-driven enhancement for exploration:Deep reinforcement learning for UAV autonomous navigation in complex environments 认领 引用
5
作者 Haotian ZHANG Yiyang LI +1 位作者 Lingquan CHENG Jianliang AI 《Chinese Journal of Aeronautics》 SCIE EI CAS CSCD 2026年第1期454-471,共18页
Unmanned Aerial Vehicle(UAV)plays a prominent role in various fields,and autonomous navigation is a crucial component of UAV intelligence.Deep Reinforcement Learning(DRL)has expanded the research avenues for addressin... Unmanned Aerial Vehicle(UAV)plays a prominent role in various fields,and autonomous navigation is a crucial component of UAV intelligence.Deep Reinforcement Learning(DRL)has expanded the research avenues for addressing challenges in autonomous navigation.Nonetheless,challenges persist,including getting stuck in local optima,consuming excessive computations during action space exploration,and neglecting deterministic experience.This paper proposes a noise-driven enhancement strategy.In accordance with the overall learning phases,a global noise control method is designed,while a differentiated local noise control method is developed by analyzing the exploration demands of four typical situations encountered by UAV during navigation.Both methods are integrated into a dual-model for noise control to regulate action space exploration.Furthermore,noise dual experience replay buffers are designed to optimize the rational utilization of both deterministic and noisy experience.In uncertain environments,based on the Twin Delay Deep Deterministic Policy Gradient(TD3)algorithm with Long Short-Term Memory(LSTM)network and Priority Experience Replay(PER),a Noise-Driven Enhancement Priority Memory TD3(NDE-PMTD3)is developed.We established a simulation environment to compare different algorithms,and the performance of the algorithms is analyzed in various scenarios.The training results indicate that the proposed algorithm accelerates the convergence speed and enhances the convergence stability.In test experiments,the proposed algorithm successfully and efficiently performs autonomous navigation tasks in diverse environments,demonstrating superior generalization results. 展开更多
关键词 Action space exploration Autonomous navigation Deep reinforcement learning Twin delay deep deterministic policy gradient Unmanned aerial vehicle
暂未订购 下载PDF
基于STL-MTL-DDPG自适应动态组合的综合能源系统多元负荷短期预测 认领 引用
6
作者 张玉敏 孙猛 +3 位作者 吉兴全 杨明 叶平峰 孟祥剑 《高电压技术》 EI CAS CSCD 北大核心 2026年第3期1178-1187,I0033-I0036,共10页
针对不同时刻多元负荷间耦合系数的峰谷变化以及气象因素对多元负荷变化的感应差异对多任务学习(multi-task learning,MTL)预测模型精度的影响,提出一种基于单任务学习(single-task learning,STL)-MTL-深度确定性策略梯度算法(deep dete... 针对不同时刻多元负荷间耦合系数的峰谷变化以及气象因素对多元负荷变化的感应差异对多任务学习(multi-task learning,MTL)预测模型精度的影响,提出一种基于单任务学习(single-task learning,STL)-MTL-深度确定性策略梯度算法(deep deterministic policy gradient,DDPG)的自适应动态组合综合能源系统多元负荷短期预测方法。首先,明晰电、冷、热多元负荷的峰谷特征指标,基于皮尔森相关系数逐时刻提取气象数据与电、冷、热负荷间的耦合关系,为后续预测模型构建提供有效的数据支撑;其次,为解决MTL模型存在耦合特征峰谷差的问题,构建STL模型相对独立的提取负荷数据自时序特征,同时解决环境因素对多元负荷变化的灵敏度干扰;然后,对MTL与STL预测数据进行重构,采用DDPG构建动态权重自适应分配模型,拟合子模型预测结果,感知外部环境变化,实现预测模型的逐时刻最优组合。最后,以亚利桑那州立大学Tempe校区综合能源系统为例进行验证,结果表明,模型电、冷、热负荷预测结果的平均绝对百分比误差分别为0.63%、0.98%和1.08%,预测精度相比其他预测模型具有一定提升。 展开更多
关键词 综合能源系统 多任务学习 单任务学习 深度确定性策略梯度算法 负荷预测
暂未订购 下载PDF
基于LLM-DDPG协同决控实现闭环自动驾驶 认领 引用
7
作者 郭彤颖 樊烁 郑岩 《汽车技术》 CSCD 北大核心 2026年第4期10-16,共7页
针对自动驾驶领域中规则驱动方法在长尾场景下泛化能力有限,数据驱动模型存在数据过拟合、决策过程缺乏可解释性等问题,提出一种基于大语言模型(LLM)和深度确定性策略梯度(DDPG)驱动的自动驾驶框架。通过整合常识与环境数据,生成可解释... 针对自动驾驶领域中规则驱动方法在长尾场景下泛化能力有限,数据驱动模型存在数据过拟合、决策过程缺乏可解释性等问题,提出一种基于大语言模型(LLM)和深度确定性策略梯度(DDPG)驱动的自动驾驶框架。通过整合常识与环境数据,生成可解释性决策;利用DDPG算法实现底层控制;在存储模块中,设计数值+语义特征相似性检索机制,为当前决策任务实时匹配相似历史案例辅助动态决策。试验结果表明:相较于Rule-Driving和DQN-Driving,提出的方法在异构测试场景A中的成功率分别提升约33%和14%;在异构测试场景B中的成功率分别提升约70%和42%,表现出更强的跨场景泛化能力和环境适应能力。 展开更多
关键词 自动驾驶 大语言模型 思维链 深度确定性策略梯度算法 案例推理
暂未订购 下载PDF
DDPG优化算法的改进型自抗扰风电机组桨距角控制 认领 引用
8
作者 徐晓宁 范召强 +3 位作者 周雪松 陶珑 问虎龙 杨风霞 《太阳能学报》 EI CAS CSCD 北大核心 2026年第1期575-584,共10页
为解决传统风电机组桨距角控制策略面对风速变化时存在动态响应差以及控制器参数适应性不足导致输出功率波动大的问题,提出一种基于深度确定性策略梯度(DDPG)算法的改进型线性自抗扰桨距角控制策略。该策略在线性扩张状态观测器(LESO)... 为解决传统风电机组桨距角控制策略面对风速变化时存在动态响应差以及控制器参数适应性不足导致输出功率波动大的问题,提出一种基于深度确定性策略梯度(DDPG)算法的改进型线性自抗扰桨距角控制策略。该策略在线性扩张状态观测器(LESO)基础上引入自由扩张维度的状态变量,并对增阶后的参数基于比例微分形式进行改进,以提高对扰动的顺馈矫正能力。随后根据发电机转速误差设计合适的奖励函数,利用DDPG算法使改进后的线性自抗扰控制(LADRC)参数能够自适应调整,实现最优的控制效果。仿真结果表明,所提策略能有效应对风速剧烈波动,使桨距角能快速适应风速变化,从而维持风电机组的稳定运行和电能的高效输出。 展开更多
关键词 风电机组 桨距角 线性自抗扰控制 深度确定性策略梯度 奖励函数 参数整定
暂未订购 下载PDF
基于DDPG的新型配电网调频功率分配仿真与控制研究 认领 引用
9
作者 李明 魏承志 +5 位作者 郭小易 赵瑞峰 卢建刚 陈益哲 高宜凡 甘锴 《电气传动》 2026年第8期39-48,共10页
新型电力系统中,新能源的发电方式与运行特性显著区别于传统同步机,频率安全稳定运行面临新的挑战和机遇。当电网出现功率扰动引发的频率波动时,为了调节系统频率,需要及时地调整机组出力,调频机组功率分配的合理性对能源利用效率和电... 新型电力系统中,新能源的发电方式与运行特性显著区别于传统同步机,频率安全稳定运行面临新的挑战和机遇。当电网出现功率扰动引发的频率波动时,为了调节系统频率,需要及时地调整机组出力,调频机组功率分配的合理性对能源利用效率和电网稳定性具有重要意义。因此,基于深度强化学习中的深度确定性策略梯度算法(DDPG),以区域系统的频率偏差和区域控制误差设计奖励函数,设置多个场景训练智能体,按照扰动大小赋予场景权重,输出通用策略集的加权功率分配因子,最终实现机组调频功率的合理分配。建立基于Python调用PSS/E的联合仿真平台,以北欧地区44节点IEEE标准电力系统模型为仿真算例,利用联合仿真平台验证了所提优化策略的可行性与优越性。 展开更多
关键词 新型电力系统 频率稳定 调频功率分配 深度确定性策略梯度算法
暂未订购 下载PDF
Transfer Learning for Deep Reinforcement Learning-Based Path Following of Autonomous Surface Vessels 认领 引用 被引量:1
10
作者 Aniket Malviya Suresh Rajendran Xueqian Zhou 《哈尔滨工程大学学报(英文版)》 CSCD 2026年第3期728-744,共17页
Deep Reinforcement Learning(DRL)offers a powerful,model-free,and data-driven approach for the navigation and control of Autonomous Surface Vessels(ASVs).The primary challenge,however,lies in the extensive training req... Deep Reinforcement Learning(DRL)offers a powerful,model-free,and data-driven approach for the navigation and control of Autonomous Surface Vessels(ASVs).The primary challenge,however,lies in the extensive training required for an agent to converge to an effective policy within a complex simulation,leading to significant computational overhead.This paper presents a multi-stage training framework that uses Transfer Learning to pass knowledge between different simulation models,resulting in a highly robust DRL controller for ASVs.The proposed framework utilizes the Deep Deterministic Policy Gradient(DDPG)algorithm to develop the data-driven controller.First,a foundational policy is efficiently learned using a simplified first-order Nomoto dynamics and second-order Nomoto dynamics,which captures the fundamental vessel dynamics.This pre-trained policy is then transferred to a complex,nonlinear Manoeuvring Modelling Group(MMG)model,significantly accelerating training convergence.Subsequently,the agent is fine-tuned within the MMG simulation with environmental disturbances.The models are evaluated on various trajectories during testing to ensure robust performance.The accuracy of the DRL controller is assessed by measuring heading error(eψ)and cross-track error(ye).A traditional Proportional-Integral-Derivative(PID)controller is implemented and compared to benchmark the DRL controller's effectiveness,to highlight the relative advantages and limitations of each approach. 展开更多
关键词 Deep reinforcement learning(DRL) Autonomous surface vessels(ASVs) Deep deterministic policy gradient(DDPG) Transfer learning Proportional-integral-derivative(PID)controller Line of sight(LOS)guidance algorithm
暂未订购 下载PDF
Multi-UAV Cooperative Path Planning Based on the Improved MADDPG 认领 引用
11
作者 Cailong Wu Caiyi Chen +2 位作者 Zhengyu Guo Jian Zhang Delin Luo 《Journal of Beijing Institute of Technology》 EI CAS 2026年第1期31-43,共13页
To address real-time path planning requirements for multi-unmanned aerial vehicle(multi-UAV)collaboration in environments,this study proposes an improved multi-agent deep deterministic policy gradient algorithm with p... To address real-time path planning requirements for multi-unmanned aerial vehicle(multi-UAV)collaboration in environments,this study proposes an improved multi-agent deep deterministic policy gradient algorithm with prioritized experience replay(PER-MADDPG).By designing a multi-dimensional state representation incorporating relative positions,velocity vectors,and obstacle distance fields,we construct a composite reward function integrating safe obstacle avoidance,formation maintenance,and energy efficiency for environment perception and multiobjective collaborative optimization.The prioritized experience replay mechanism dynamically adjusts sampling weights based on temporal difference(TD)errors,enhancing learning efficiency for high-value samples.Simulation experiments demonstrate that our method generates real-time collaborative paths in 3D complex obstacle environments,reducing training time by 25.3%and 16.8%compared to traditional MADDPG and multi-agent twin delayed deep deterministic policy gradient(MATD3)algorithms respectively,while achieving smaller path length variances among UAVs.Results validate the effectiveness of prioritized experience replay in multi-agent collaborative decision-making. 展开更多
关键词 multi-unmanned aerial vehicle(multi-UAV) path planning deep deterministic policy gradient prioritized experience replay
暂未订购 下载PDF
Multi-UAV Collaborative Energy Charging for Battery-Free SWIPT-Enabled Sensor Networks Based on MADDPG 认领 引用
12
作者 Xiangyi Le Deyu Lin +2 位作者 Yufei Zhao Wang Miao Yong Liang Guan 《Computers, Materials & Continua》 SCIE EI 2026年第9期2376-2395,共20页
The emergence of Unmanned Aerial Vehicle(UAV)-enabled Wireless Energy Transfer(WET)and Simultaneous Wireless Information and Power Transfer(SWIPT)technology provide a promising solution to overcome the energy sustaina... The emergence of Unmanned Aerial Vehicle(UAV)-enabled Wireless Energy Transfer(WET)and Simultaneous Wireless Information and Power Transfer(SWIPT)technology provide a promising solution to overcome the energy sustainability limitations of traditional harvesting-reliant sensor networks.However,in large-scale Battery-free SWIPT-enabled Sensor Networks(BSSN)characterized by sparse node distribution and heterogeneous energy consumption and harvesting rates,employing a single UAV for energy replenishment often suffers from insufficient operation continuity and low charging efficiency.To overcome these challenges,a Multi-UAV Collaborative Energy Charging for BSSN Based on Multi-Agent Deep Deterministic Policy Gradient(MCEC-MADDPG)is proposed in this paper.Specifically,we construct a collaborative one-to-one precision energy supply model where UAVs hover directly above specific nodes to achieve power transmission without complex beamforming requirements.To achieve collaborative scheduling among multiple UAVs in wide-area dynamic environments,the energy replenishment problem is first formulated as a Partially Observable Markov Decision Process(POMDP).Subsequently,the Centralized Training with Decentralized Execution(CTDE)architecture of the MADDPG algorithm is leveraged to solve this POMDP,which effectively tackles the non-stationarity challenge inherent in multi-agent environments.Simulation results demonstrate that MCEC-MADDPG exhibits superior performance in terms of convergence speed and stability.It enables the adaptive emergence of spatial-division collaborative strategies,significantly enhances the average residual energy of the network,and elevates the node survival rate to nearly 90%.Compared with Deep Deterministic Policy Gradient(DDPG),the traditional static Partition-Greedy method,the heuristic K-Means algorithm and the dynamic Two-Layer task allocation strategy,the proposed approach demonstrates substantial advantages. 展开更多
关键词 Battery-free SWIPT-enabled sensor networks multi-agent deep deterministic policy gradient multi-unmanned aerial vehicle collaborative energy charging partially observable Markov decision process centralized training with decentralized execution
暂未订购 下载PDF
基于改进DDPG算法的无人船避碰路径规划 认领 引用
13
作者 任重霖 许志远 《舰船科学技术》 北大核心 2026年第9期125-132,共8页
针对传统路径规划算法在无人船避碰过程中存在路径不符合船舶运动特性、对《国际海上避碰规则》(COLREGs)遵循性不足等问题,以及深度确定性策略梯度(Deep Deterministic Policy Gradient,DDPG)算法在训练过程中收敛效率低的缺陷,本文提... 针对传统路径规划算法在无人船避碰过程中存在路径不符合船舶运动特性、对《国际海上避碰规则》(COLREGs)遵循性不足等问题,以及深度确定性策略梯度(Deep Deterministic Policy Gradient,DDPG)算法在训练过程中收敛效率低的缺陷,本文提出一种基于改进DDPG算法的无人船自主避碰路径规划方法。该方法融合优先级经验回放(Prioritized Experience Replay,PER)机制与长短期记忆网络(Long Short-Term Memory,LSTM)时序建模模块,构建具备动态样本筛选能力和时空特征提取能力的Actor-Critic网络架构,从而提升算法在连续动作空间中的策略学习效率与泛化能力。同时,设计基于COLREGs的奖励函数,实现对路径安全性与航行合规性的协同优化。仿真实验结果表明,相较于传统路径规划算法,所提方法在路径长度和平滑度方面表现更优;在与动态船舶交互的复杂场景中,能够严格遵循COLREGs完成高效避碰决策。此外,与传统DDPG算法相比,改进后的算法在收敛速度和最终性能上均有显著提升。本研究为复杂海况下无人船的避碰路径规划提供了新的技术思路与方法支撑。 展开更多
关键词 无人船 路径规划 深度确定性策略梯度算法 《国际海上避碰规则》
暂未订购 下载PDF
基于CLP-DDPG算法的复杂环境下无人机路径规划(特邀) 认领 引用
14
作者 谢伟宁 陈龙胜 +3 位作者 何国毅 宋伟 王凯 陈敬玮 《南昌航空大学学报(自然科学版)》 CAS 2026年第2期1-10,共10页
针对复杂环境下无人机路径规划存在的探索效率低、收敛速度慢与路径平滑度不佳等问题,本文提出一种融合课程学习与嵌入优先回放机制的深度确定性策略梯度算法的路径规划改进方法。首先,通过设计一套从易到难的障碍物环境课程,引导无人... 针对复杂环境下无人机路径规划存在的探索效率低、收敛速度慢与路径平滑度不佳等问题,本文提出一种融合课程学习与嵌入优先回放机制的深度确定性策略梯度算法的路径规划改进方法。首先,通过设计一套从易到难的障碍物环境课程,引导无人机逐步学习从简单到复杂的路径规划任务,提高训练效率和稳定性。然后,将优先回放机制嵌入算法中,保证在冗杂的经验池中快速提取有效经验,进一步提升收敛速度和稳定性,确保路径效率更高和平滑度更优。仿真结果表明,融合课程学习与嵌入优先回放机制的强化学习方法能有效提升无人机在未知复杂环境下的自主避障与路径规划能力,相比于传统的DDPG算法路径规划效率提高了26.84%,平滑度提升了66.19%,训练的收敛速度更快且后期训练的稳定性更好。 展开更多
关键词 课程学习 深度确定性策略梯度算法 路径规划 无人机 强化学习
暂未订购 下载PDF
基于动态权重多指标经验回放的MADDPG算法研究 认领 引用
15
作者 胡金泽 唐宏伟 +2 位作者 程翰超 谢培淼 贺露谊 《农业装备与车辆工程》 2026年第1期73-80,共8页
针对多智能体深度强化学习中传统经验回放机制存在的评估指标单一与权重策略静态化问题,提出一种基于动态权重多指标经验回放的改进MADDPG算法。设计了多维度经验评估体系,将时序差分误差、经验年龄和合作贡献度3个指标系统融合,实现对... 针对多智能体深度强化学习中传统经验回放机制存在的评估指标单一与权重策略静态化问题,提出一种基于动态权重多指标经验回放的改进MADDPG算法。设计了多维度经验评估体系,将时序差分误差、经验年龄和合作贡献度3个指标系统融合,实现对经验样本价值的全面评估;提出了动态权重调整机制,通过训练进程自适应的权重系数调整,使算法在训练初期注重个体价值函数准确性,后期偏向团队协作优化;构建了协作感知的优先级框架,通过合作贡献度指标显式量化经验在多智能体协作中的价值,提升团队协作效率。在OpenAI多智能体粒子环境的3个典型场景中的实验结果表明:与对比算法相比,所提算法在平均回合奖励、目标达成率与冲突规避率等关键性能指标上均有提升,收敛速度更快,验证了其有效性与优越性。 展开更多
关键词 多智能体强化学习 多智能体深度确定性策略梯度算法 经验回放 动态权重 合作贡献度 协作探索
暂未订购 下载PDF
基于DDPG算法的无轴承永磁薄片电机磁悬浮控制 认领 引用
16
作者 林景灿 周扬忠 《微特电机》 2026年第5期41-46,共6页
由于无轴承永磁薄片电机(bearingless permanent magnet slice motor,BPMSM)动态偏心磁扰动、参数变化等特点使得传统的固定参数比例-积分-微分(proportional-integral-derivative,PID)磁悬浮控制性能不佳。本文提出一种结合深度确定性... 由于无轴承永磁薄片电机(bearingless permanent magnet slice motor,BPMSM)动态偏心磁扰动、参数变化等特点使得传统的固定参数比例-积分-微分(proportional-integral-derivative,PID)磁悬浮控制性能不佳。本文提出一种结合深度确定性策略梯度(deep deterministic policy gradient,DDPG)算法的无轴承永磁薄片电机磁悬浮PID控制方法,该方法利用DDPG智能体动态调整PID控制器参数,以提高磁悬浮控制精度和响应速度。本文分析了BPMSM磁悬浮数学模型,并在此基础上,构建磁悬浮PID控制策略;提出了基于DDPG算法的PID参数自适应策略,并设计了由误差惩罚、精度奖励和超调惩罚构成的奖励函数;通过MATLAB/Simulink仿真验证,结果表明,相比固定参数PID控制,深度确定性策略梯度-比例-积分-微分(deep deterministic policy gradient-proportional-integral-derivative,DDPG-PID)控制器在径向位移跟随和径向负载扰动实验中表现出更小的超调量、更短的调节时间以及更强的抗干扰和自适应能力。 展开更多
关键词 无轴承永磁薄片电机 深度确定性策略梯度 比例-积分-微分控制
暂未订购 下载PDF
基于DDPG控制算法的电厂热控系统优化研究 认领 引用
17
作者 郑磊 郝宏伟 +4 位作者 袁国平 刘龙 蒋波 焦明明 牛博涵 《科技资讯》 2026年第11期96-98,共3页
为解决电厂热控系统多变量、强耦合等复杂特性导致传统比例-积分-微分(Proportional-Integral-Derivative,PID)控制自适应与优化能力不足的问题,本研究将深度确定性策略梯度(Deep Deterministic Policy Gradient,DDPG)算法引入电厂热控... 为解决电厂热控系统多变量、强耦合等复杂特性导致传统比例-积分-微分(Proportional-Integral-Derivative,PID)控制自适应与优化能力不足的问题,本研究将深度确定性策略梯度(Deep Deterministic Policy Gradient,DDPG)算法引入电厂热控系统优化。通过构建马尔可夫决策过程模型,明确状态、动作空间和多目标加权奖励函数,在600 MW超临界机组仿真平台开展实验,对比DDPG算法与传统PID控制性能。结果显示,DDPG算法控制下,主蒸汽参数超调更小、稳定更快,机组煤耗降低,自助发电控制响应速率提升。DDPG算法能够有效提升电厂热控系统性能,为火电智能化转型提供支撑。 展开更多
关键词 深度确定性策略梯度算法 电厂热控系统 马尔可夫决策过程 比例-积分-微分控制 机组优化
暂未订购 下载PDF
基于注意力增强DDPG的短波协同测向与定位方法 认领 引用
18
作者 冯祺玥 唐涛 +2 位作者 张昀普 赵排航 陈晓云 《信息工程大学学报》 2026年第3期259-266,274,共8页
复杂场景下的短波定位在战场通信中发挥着至关重要的作用,而现有基于深度学习的短波智能测向方法主要依赖人工标注且优质短波信号样本较少。针对上述问题,在短波信号空时高分辨时频图基础上,提出了一种基于注意力增强深度确定性策略梯度... 复杂场景下的短波定位在战场通信中发挥着至关重要的作用,而现有基于深度学习的短波智能测向方法主要依赖人工标注且优质短波信号样本较少。针对上述问题,在短波信号空时高分辨时频图基础上,提出了一种基于注意力增强深度确定性策略梯度(DDPG)算法的短波协同测向与定位方法。该方法首先设计了基于空时高分辨时频图的强化学习环境和具有混合动作空间的智能体,再依据估计定位结果设计奖励函数,以此建立可观测马尔可夫决策过程。其次,利用注意力机制增强深度强化学习的多层次非精细化评估,实现自动测向定位,可以减少人工标注,实现边工作边学习的自主进化,逐步提升短波信号的智能测向定位能力。实验结果表明,所提算法在保证定位精度与参数估计算法性能相当的情况下,使得测向定位时间缩短了77.7%。 展开更多
关键词 测向定位 马尔可夫决策过程 空时高分辨 混合动作空间 深度确定性策略梯度算法
暂未订购 下载PDF
DDPG-Based Intelligent Computation Offloading and Resource Allocation for LEO Satellite Edge Computing Network 认领 引用 被引量:2
19
作者 Jia Min Wu Jian +2 位作者 Zhang Liang Wang Xinyu Guo Qing 《China Communications》 SCIE EI CSCD 2025年第3期1-15,共15页
Low earth orbit(LEO)satellites with wide coverage can carry the mobile edge computing(MEC)servers with powerful computing capabilities to form the LEO satellite edge computing system,providing computing services for t... Low earth orbit(LEO)satellites with wide coverage can carry the mobile edge computing(MEC)servers with powerful computing capabilities to form the LEO satellite edge computing system,providing computing services for the global ground users.In this paper,the computation offloading problem and resource allocation problem are formulated as a mixed integer nonlinear program(MINLP)problem.This paper proposes a computation offloading algorithm based on deep deterministic policy gradient(DDPG)to obtain the user offloading decisions and user uplink transmission power.This paper uses the convex optimization algorithm based on Lagrange multiplier method to obtain the optimal MEC server resource allocation scheme.In addition,the expression of suboptimal user local CPU cycles is derived by relaxation method.Simulation results show that the proposed algorithm can achieve excellent convergence effect,and the proposed algorithm significantly reduces the system utility values at considerable time cost compared with other algorithms. 展开更多
关键词 computation offloading deep deterministic policy gradient low earth orbit satellite mobile edge computing resource allocation
暂未订购 下载PDF
基于改进DDPG算法的无人船自主避碰决策方法 认领 引用 被引量:7
20
作者 关巍 郝淑慧 +1 位作者 崔哲闻 王淼淼 《中国舰船研究》 CSCD 北大核心 2025年第1期172-180,共9页
[目的]针对传统深度确定性策略梯度(DDPG)算法数据利用率低、收敛性差的特点,改进并提出一种新的无人船自主避碰决策方法。[方法]利用优先经验回放(PER)自适应调节经验优先级,降低样本的相关性,并利用长短期记忆(LSTM)网络提高算法的收... [目的]针对传统深度确定性策略梯度(DDPG)算法数据利用率低、收敛性差的特点,改进并提出一种新的无人船自主避碰决策方法。[方法]利用优先经验回放(PER)自适应调节经验优先级,降低样本的相关性,并利用长短期记忆(LSTM)网络提高算法的收敛性。基于船舶领域和《国际海上避碰规则》(COLREGs),设置会遇情况判定模型和一组新定义的奖励函数,并考虑了紧迫危险以应对他船不遵守规则的情况。为验证所提方法的有效性,在两船和多船会遇局面下进行仿真实验。[结果]结果表明,改进的DDPG算法相比于传统DDPG算法在收敛速度上提升约28.8%,[结论]训练好的自主避碰模型可以使无人船在遵守COLREGs的同时实现自主决策和导航,为实现更加安全、高效的海上交通智能化决策提供参考。 展开更多
关键词 无人船 深度确定性策略梯度算法 自主避碰决策 优先经验回放 国际海上避碰规则 避碰
暂未订购 下载PDF
上一页 1 2 24 下一页 到第
在线咨询 使用帮助 返回顶部 意见反馈