The deep deterministic policy gradient(DDPG)algo-rithm is an off-policy method that combines two mainstream reinforcement learning methods based on value iteration and policy iteration.Using the DDPG algorithm,agents ...The deep deterministic policy gradient(DDPG)algo-rithm is an off-policy method that combines two mainstream reinforcement learning methods based on value iteration and policy iteration.Using the DDPG algorithm,agents can explore and summarize the environment to achieve autonomous deci-sions in the continuous state space and action space.In this paper,a cooperative defense with DDPG via swarms of unmanned aerial vehicle(UAV)is developed and validated,which has shown promising practical value in the effect of defending.We solve the sparse rewards problem of reinforcement learning pair in a long-term task by building the reward function of UAV swarms and optimizing the learning process of artificial neural network based on the DDPG algorithm to reduce the vibration in the learning process.The experimental results show that the DDPG algorithm can guide the UAVs swarm to perform the defense task efficiently,meeting the requirements of a UAV swarm for non-centralization,autonomy,and promoting the intelligent development of UAVs swarm as well as the decision-making process.展开更多
针对可移动阵元同时透射和反射可重构智能表面(Movable Elements Based Simultaneous Transmitting and Reflecting Reconfigurable Intelligent Surface,ME-STAR-RIS)辅助抗干扰系统中信道估计开销巨大的问题,提出一种基于深度确定性...针对可移动阵元同时透射和反射可重构智能表面(Movable Elements Based Simultaneous Transmitting and Reflecting Reconfigurable Intelligent Surface,ME-STAR-RIS)辅助抗干扰系统中信道估计开销巨大的问题,提出一种基于深度确定性策略梯度(Deep Deterministic Policy Gradient,DDPG)算法的双尺度协同优化抗干扰传输方法。首先利用统计信道状态信息(Channel State Information,CSI)优化长时阵元位置,再基于优化后的阵元位置估计瞬时CSI,进而优化短时波束成形。为解决高维连续状态空间以及阵元位置、相移系数等连续动作空间带来的优化难题,引入DDPG算法实现动态策略学习。仿真结果表明,所以方法相较于瞬时CSI联合优化方案虽存在约1.5 b/s/Hz的性能损失,但显著降低了信道估计开销。展开更多
基金supported by the Key Research and Development Program of Shaanxi(2022GY-089)the Natural Science Basic Research Program of Shaanxi(2022JQ-593).
摘要The deep deterministic policy gradient(DDPG)algo-rithm is an off-policy method that combines two mainstream reinforcement learning methods based on value iteration and policy iteration.Using the DDPG algorithm,agents can explore and summarize the environment to achieve autonomous deci-sions in the continuous state space and action space.In this paper,a cooperative defense with DDPG via swarms of unmanned aerial vehicle(UAV)is developed and validated,which has shown promising practical value in the effect of defending.We solve the sparse rewards problem of reinforcement learning pair in a long-term task by building the reward function of UAV swarms and optimizing the learning process of artificial neural network based on the DDPG algorithm to reduce the vibration in the learning process.The experimental results show that the DDPG algorithm can guide the UAVs swarm to perform the defense task efficiently,meeting the requirements of a UAV swarm for non-centralization,autonomy,and promoting the intelligent development of UAVs swarm as well as the decision-making process.
摘要针对可移动阵元同时透射和反射可重构智能表面(Movable Elements Based Simultaneous Transmitting and Reflecting Reconfigurable Intelligent Surface,ME-STAR-RIS)辅助抗干扰系统中信道估计开销巨大的问题,提出一种基于深度确定性策略梯度(Deep Deterministic Policy Gradient,DDPG)算法的双尺度协同优化抗干扰传输方法。首先利用统计信道状态信息(Channel State Information,CSI)优化长时阵元位置,再基于优化后的阵元位置估计瞬时CSI,进而优化短时波束成形。为解决高维连续状态空间以及阵元位置、相移系数等连续动作空间带来的优化难题,引入DDPG算法实现动态策略学习。仿真结果表明,所以方法相较于瞬时CSI联合优化方案虽存在约1.5 b/s/Hz的性能损失,但显著降低了信道估计开销。