This study proposes an automatic control system for Autonomous Underwater Vehicle(AUV)docking,utilizing a digital twin(DT)environment based on the HoloOcean platform,which integrates six-degree-of-freedom(6-DOF)motion...This study proposes an automatic control system for Autonomous Underwater Vehicle(AUV)docking,utilizing a digital twin(DT)environment based on the HoloOcean platform,which integrates six-degree-of-freedom(6-DOF)motion equations and hydrodynamic coefficients to create a realistic simulation.Although conventional model-based and visual servoing approaches often struggle in dynamic underwater environments due to limited adaptability and extensive parameter tuning requirements,deep reinforcement learning(DRL)offers a promising alternative.In the positioning stage,the Twin Delayed Deep Deterministic Policy Gradient(TD3)algorithm is employed for synchronized depth and heading control,which offers stable training,reduced overestimation bias,and superior handling of continuous control compared to other DRL methods.During the searching stage,zig-zag heading motion combined with a state-of-the-art object detection algorithm facilitates docking station localization.For the docking stage,this study proposes an innovative Image-based DDPG(I-DDPG),enhanced and trained in a Unity-MATLAB simulation environment,to achieve visual target tracking.Furthermore,integrating a DT environment enables efficient and safe policy training,reduces dependence on costly real-world tests,and improves sim-to-real transfer performance.Both simulation and real-world experiments were conducted,demonstrating the effectiveness of the system in improving AUV control strategies and supporting the transition from simulation to real-world operations in underwater environments.The results highlight the scalability and robustness of the proposed system,as evidenced by the TD3 controller achieving 25%less oscillation than the adaptive fuzzy controller when reaching the target depth,thereby demonstrating superior stability,accuracy,and potential for broader and more complex autonomous underwater tasks.展开更多
Unmanned Aerial Vehicle(UAV)plays a prominent role in various fields,and autonomous navigation is a crucial component of UAV intelligence.Deep Reinforcement Learning(DRL)has expanded the research avenues for addressin...Unmanned Aerial Vehicle(UAV)plays a prominent role in various fields,and autonomous navigation is a crucial component of UAV intelligence.Deep Reinforcement Learning(DRL)has expanded the research avenues for addressing challenges in autonomous navigation.Nonetheless,challenges persist,including getting stuck in local optima,consuming excessive computations during action space exploration,and neglecting deterministic experience.This paper proposes a noise-driven enhancement strategy.In accordance with the overall learning phases,a global noise control method is designed,while a differentiated local noise control method is developed by analyzing the exploration demands of four typical situations encountered by UAV during navigation.Both methods are integrated into a dual-model for noise control to regulate action space exploration.Furthermore,noise dual experience replay buffers are designed to optimize the rational utilization of both deterministic and noisy experience.In uncertain environments,based on the Twin Delay Deep Deterministic Policy Gradient(TD3)algorithm with Long Short-Term Memory(LSTM)network and Priority Experience Replay(PER),a Noise-Driven Enhancement Priority Memory TD3(NDE-PMTD3)is developed.We established a simulation environment to compare different algorithms,and the performance of the algorithms is analyzed in various scenarios.The training results indicate that the proposed algorithm accelerates the convergence speed and enhances the convergence stability.In test experiments,the proposed algorithm successfully and efficiently performs autonomous navigation tasks in diverse environments,demonstrating superior generalization results.展开更多
Deep Reinforcement Learning(DRL)offers a powerful,model-free,and data-driven approach for the navigation and control of Autonomous Surface Vessels(ASVs).The primary challenge,however,lies in the extensive training req...Deep Reinforcement Learning(DRL)offers a powerful,model-free,and data-driven approach for the navigation and control of Autonomous Surface Vessels(ASVs).The primary challenge,however,lies in the extensive training required for an agent to converge to an effective policy within a complex simulation,leading to significant computational overhead.This paper presents a multi-stage training framework that uses Transfer Learning to pass knowledge between different simulation models,resulting in a highly robust DRL controller for ASVs.The proposed framework utilizes the Deep Deterministic Policy Gradient(DDPG)algorithm to develop the data-driven controller.First,a foundational policy is efficiently learned using a simplified first-order Nomoto dynamics and second-order Nomoto dynamics,which captures the fundamental vessel dynamics.This pre-trained policy is then transferred to a complex,nonlinear Manoeuvring Modelling Group(MMG)model,significantly accelerating training convergence.Subsequently,the agent is fine-tuned within the MMG simulation with environmental disturbances.The models are evaluated on various trajectories during testing to ensure robust performance.The accuracy of the DRL controller is assessed by measuring heading error(eψ)and cross-track error(ye).A traditional Proportional-Integral-Derivative(PID)controller is implemented and compared to benchmark the DRL controller's effectiveness,to highlight the relative advantages and limitations of each approach.展开更多
针对柔性直流输电系统(voltage source converter based high voltage direct current transmission,VSC-HVDC)控制参数设计过程中存在的鲁棒性差、依赖已知电路参数、工程设计经验化等问题,提出一种基于马尔科夫转换场(Markov transiti...针对柔性直流输电系统(voltage source converter based high voltage direct current transmission,VSC-HVDC)控制参数设计过程中存在的鲁棒性差、依赖已知电路参数、工程设计经验化等问题,提出一种基于马尔科夫转换场(Markov transition field,MTF)与深度确定性策略梯度算法(deep deterministic policy gradient,DDPG)结合的鲁棒性强、不依赖电路参数特性以及可视化的VSC-HVDC控制参数优化设计方法。首先,采用马尔科夫转换场将电路功率、电压等一维时序波形数据转换为二维马尔科夫转换场域图像并使用马尔科夫转换场损失函数(Markov transition field loss,MTFL)判断二维转换域图的数据波动性;其次,将MTFL损失函数与DDPG算法相结合,综合利用MTFL损失函数对系统输出时序数据动态特性评价能力更强的优点和DDPG算法泛化性能优秀的特点,实现VSC-HVDC系统控制参数优化;最后,通过MATLAB模拟和实验结果验证该方法的有效性。展开更多
In the complex and variable deep-sea environment,the compensation control of ship motion ensures the safety and efficiency of equipment installation and transportation in offshore wind farms.However,the ship motion po...In the complex and variable deep-sea environment,the compensation control of ship motion ensures the safety and efficiency of equipment installation and transportation in offshore wind farms.However,the ship motion posture compensation control system is severely affected by uncertainties,which significantly impact the accuracy of compensation control.In this paper,we propose a ship three-degree-of-freedom(3-DoF)motion posture stabilization control method based on the DTW-LSTM-MATD3 algorithm.We use the multi-agent twin delayed deep deterministic policy gradient(MATD3)to control a platform with six electric cylinders to achieve stable control.However,owing to random noise affecting the ship’s motion posture,we use a dynamic time warping(DTW)algorithm to distinguish between high-frequency noise and low-frequency tracking signals.Further,we embed a long short-term memory(LSTM)network into the MATD3 network to better align the Critic network’s training with the true Q-value.We use a combined reward function to enhance the agent’s exploration capability in complex dynamic environments.Finally,verification was conducted under sixth-level,abrupt sea conditions with high-frequency noise,as well as under real abrupt sea conditions,and a generalization test was also carried out.Simulation results show that the proposed DTW-LSTM-MATD3 method has great compensation control ability.展开更多
Migratory birds depend on the perception of atmospheric updraft for long-distance flight.To realize more efficient autonomous soaring in an unpowered glider,different strategies for using potential sensorimotor cues t...Migratory birds depend on the perception of atmospheric updraft for long-distance flight.To realize more efficient autonomous soaring in an unpowered glider,different strategies for using potential sensorimotor cues to achieve autonomous soaring efficiency were compared and optimized.A simulation framework of autonomous soaring for an unpowered glider was developed based on a reinforcement learning algorithm.The framework was composed of three models:an updraft environment model,the glider's dynamics and control model,and a reinforcement learning agent,which learns to harvest more energy in flight.Based on the simulation,effects of different combinations of 12 potential sensorimotor cues on soaring efficiency were studied.Firstly,the absence of one particular sensorimotor cue and the use of only a single valid cue in autonomous soaring were analyzed.The results showed that the vertical airflow velocity gradient(aw)and the wing-tip updraft velocity difference(τ)have advantages over the other cues.Secondly,strategies combining aw orτwith other cues were analyzed to achieve more effective autonomous soaring,and seven potentially effective combinations of sensorimotor cues were identified.The final results showed that,among the tested combinations,the combination of vertical airflow velocity(Vw)andτ,enables the most efficient autonomous soaring.This study identified a highly effective sensorimotor cue strategy to guide an intelligent glider to achieve long-distance autonomous soaring flight.展开更多
基金supported by the National Science and Technology Council,Taiwan[Grant NSTC 111-2628-E-006-005-MY3]supported by the Ocean Affairs Council,Taiwansponsored in part by Higher Education Sprout Project,Ministry of Education to the Headquarters of University Advancement at National Cheng Kung University(NCKU).
摘要This study proposes an automatic control system for Autonomous Underwater Vehicle(AUV)docking,utilizing a digital twin(DT)environment based on the HoloOcean platform,which integrates six-degree-of-freedom(6-DOF)motion equations and hydrodynamic coefficients to create a realistic simulation.Although conventional model-based and visual servoing approaches often struggle in dynamic underwater environments due to limited adaptability and extensive parameter tuning requirements,deep reinforcement learning(DRL)offers a promising alternative.In the positioning stage,the Twin Delayed Deep Deterministic Policy Gradient(TD3)algorithm is employed for synchronized depth and heading control,which offers stable training,reduced overestimation bias,and superior handling of continuous control compared to other DRL methods.During the searching stage,zig-zag heading motion combined with a state-of-the-art object detection algorithm facilitates docking station localization.For the docking stage,this study proposes an innovative Image-based DDPG(I-DDPG),enhanced and trained in a Unity-MATLAB simulation environment,to achieve visual target tracking.Furthermore,integrating a DT environment enables efficient and safe policy training,reduces dependence on costly real-world tests,and improves sim-to-real transfer performance.Both simulation and real-world experiments were conducted,demonstrating the effectiveness of the system in improving AUV control strategies and supporting the transition from simulation to real-world operations in underwater environments.The results highlight the scalability and robustness of the proposed system,as evidenced by the TD3 controller achieving 25%less oscillation than the adaptive fuzzy controller when reaching the target depth,thereby demonstrating superior stability,accuracy,and potential for broader and more complex autonomous underwater tasks.
基金the Collaborative Innovation Project of Shanghai,China for the financial support。
摘要Unmanned Aerial Vehicle(UAV)plays a prominent role in various fields,and autonomous navigation is a crucial component of UAV intelligence.Deep Reinforcement Learning(DRL)has expanded the research avenues for addressing challenges in autonomous navigation.Nonetheless,challenges persist,including getting stuck in local optima,consuming excessive computations during action space exploration,and neglecting deterministic experience.This paper proposes a noise-driven enhancement strategy.In accordance with the overall learning phases,a global noise control method is designed,while a differentiated local noise control method is developed by analyzing the exploration demands of four typical situations encountered by UAV during navigation.Both methods are integrated into a dual-model for noise control to regulate action space exploration.Furthermore,noise dual experience replay buffers are designed to optimize the rational utilization of both deterministic and noisy experience.In uncertain environments,based on the Twin Delay Deep Deterministic Policy Gradient(TD3)algorithm with Long Short-Term Memory(LSTM)network and Priority Experience Replay(PER),a Noise-Driven Enhancement Priority Memory TD3(NDE-PMTD3)is developed.We established a simulation environment to compare different algorithms,and the performance of the algorithms is analyzed in various scenarios.The training results indicate that the proposed algorithm accelerates the convergence speed and enhances the convergence stability.In test experiments,the proposed algorithm successfully and efficiently performs autonomous navigation tasks in diverse environments,demonstrating superior generalization results.
摘要Deep Reinforcement Learning(DRL)offers a powerful,model-free,and data-driven approach for the navigation and control of Autonomous Surface Vessels(ASVs).The primary challenge,however,lies in the extensive training required for an agent to converge to an effective policy within a complex simulation,leading to significant computational overhead.This paper presents a multi-stage training framework that uses Transfer Learning to pass knowledge between different simulation models,resulting in a highly robust DRL controller for ASVs.The proposed framework utilizes the Deep Deterministic Policy Gradient(DDPG)algorithm to develop the data-driven controller.First,a foundational policy is efficiently learned using a simplified first-order Nomoto dynamics and second-order Nomoto dynamics,which captures the fundamental vessel dynamics.This pre-trained policy is then transferred to a complex,nonlinear Manoeuvring Modelling Group(MMG)model,significantly accelerating training convergence.Subsequently,the agent is fine-tuned within the MMG simulation with environmental disturbances.The models are evaluated on various trajectories during testing to ensure robust performance.The accuracy of the DRL controller is assessed by measuring heading error(eψ)and cross-track error(ye).A traditional Proportional-Integral-Derivative(PID)controller is implemented and compared to benchmark the DRL controller's effectiveness,to highlight the relative advantages and limitations of each approach.
摘要针对柔性直流输电系统(voltage source converter based high voltage direct current transmission,VSC-HVDC)控制参数设计过程中存在的鲁棒性差、依赖已知电路参数、工程设计经验化等问题,提出一种基于马尔科夫转换场(Markov transition field,MTF)与深度确定性策略梯度算法(deep deterministic policy gradient,DDPG)结合的鲁棒性强、不依赖电路参数特性以及可视化的VSC-HVDC控制参数优化设计方法。首先,采用马尔科夫转换场将电路功率、电压等一维时序波形数据转换为二维马尔科夫转换场域图像并使用马尔科夫转换场损失函数(Markov transition field loss,MTFL)判断二维转换域图的数据波动性;其次,将MTFL损失函数与DDPG算法相结合,综合利用MTFL损失函数对系统输出时序数据动态特性评价能力更强的优点和DDPG算法泛化性能优秀的特点,实现VSC-HVDC系统控制参数优化;最后,通过MATLAB模拟和实验结果验证该方法的有效性。
基金supported by the National Natural Science Foundation of China(No.52105466).
摘要In the complex and variable deep-sea environment,the compensation control of ship motion ensures the safety and efficiency of equipment installation and transportation in offshore wind farms.However,the ship motion posture compensation control system is severely affected by uncertainties,which significantly impact the accuracy of compensation control.In this paper,we propose a ship three-degree-of-freedom(3-DoF)motion posture stabilization control method based on the DTW-LSTM-MATD3 algorithm.We use the multi-agent twin delayed deep deterministic policy gradient(MATD3)to control a platform with six electric cylinders to achieve stable control.However,owing to random noise affecting the ship’s motion posture,we use a dynamic time warping(DTW)algorithm to distinguish between high-frequency noise and low-frequency tracking signals.Further,we embed a long short-term memory(LSTM)network into the MATD3 network to better align the Critic network’s training with the true Q-value.We use a combined reward function to enhance the agent’s exploration capability in complex dynamic environments.Finally,verification was conducted under sixth-level,abrupt sea conditions with high-frequency noise,as well as under real abrupt sea conditions,and a generalization test was also carried out.Simulation results show that the proposed DTW-LSTM-MATD3 method has great compensation control ability.
基金supported by the National Natural Science Foundation of China(Nos.12202384 and U2241274)the Leading Talent Project for Scientific and Technological Innovation in Zhejiang Province(No.2023R5220)the Specialized Research Projects of Huanjiang Laboratory,China。
摘要Migratory birds depend on the perception of atmospheric updraft for long-distance flight.To realize more efficient autonomous soaring in an unpowered glider,different strategies for using potential sensorimotor cues to achieve autonomous soaring efficiency were compared and optimized.A simulation framework of autonomous soaring for an unpowered glider was developed based on a reinforcement learning algorithm.The framework was composed of three models:an updraft environment model,the glider's dynamics and control model,and a reinforcement learning agent,which learns to harvest more energy in flight.Based on the simulation,effects of different combinations of 12 potential sensorimotor cues on soaring efficiency were studied.Firstly,the absence of one particular sensorimotor cue and the use of only a single valid cue in autonomous soaring were analyzed.The results showed that the vertical airflow velocity gradient(aw)and the wing-tip updraft velocity difference(τ)have advantages over the other cues.Secondly,strategies combining aw orτwith other cues were analyzed to achieve more effective autonomous soaring,and seven potentially effective combinations of sensorimotor cues were identified.The final results showed that,among the tested combinations,the combination of vertical airflow velocity(Vw)andτ,enables the most efficient autonomous soaring.This study identified a highly effective sensorimotor cue strategy to guide an intelligent glider to achieve long-distance autonomous soaring flight.