Unmanned Aerial Vehicles(UAVs)have become integral components in smart city infrastructures,supporting applications such as emergency response,surveillance,and data collection.However,the high mobility and dynamic top...Unmanned Aerial Vehicles(UAVs)have become integral components in smart city infrastructures,supporting applications such as emergency response,surveillance,and data collection.However,the high mobility and dynamic topology of Flying Ad Hoc Networks(FANETs)present significant challenges for maintaining reliable,low-latency communication.Conventional geographic routing protocols often struggle in situations where link quality varies and mobility patterns are unpredictable.To overcome these limitations,this paper proposes an improved routing protocol based on reinforcement learning.This new approach integrates Q-learning with mechanisms that are both link-aware and mobility-aware.The proposed method optimizes the selection of relay nodes by using an adaptive reward function that takes into account energy consumption,delay,and link quality.Additionally,a Kalman filter is integrated to predict UAV mobility,improving the stability of communication links under dynamic network conditions.Simulation experiments were conducted using realistic scenarios,varying the number of UAVs to assess scalability.An analysis was conducted on key performance metrics,including the packet delivery ratio,end-to-end delay,and total energy consumption.The results demonstrate that the proposed approach significantly improves the packet delivery ratio by 12%–15%and reduces delay by up to 25.5%when compared to conventional GEO and QGEO protocols.However,this improvement comes at the cost of higher energy consumption due to additional computations and control overhead.Despite this trade-off,the proposed solution ensures reliable and efficient communication,making it well-suited for large-scale UAV networks operating in complex urban environments.展开更多
For unmanned surface vehicles(USVs),how to find an effective,feasible path that substantially improves mission success rates and time efficiency in dynamic marine environments is a critical issue.To address the path p...For unmanned surface vehicles(USVs),how to find an effective,feasible path that substantially improves mission success rates and time efficiency in dynamic marine environments is a critical issue.To address the path planning problem for USVs using deep reinforcement learning(DRL)in dynamic ocean environments,an improved algorithm based on Deep Q-Networks(DQN)is proposed,which is called Fast Guided Deep Q-Network Algorithm(FG-DQN).This algorithm combines DQN with the artificial potential field(APF)method and uses the A*algorithm to initialize a guiding path in a global static environment and to provide prior knowledge for the USVs.Additionally,the configuration of the reward function using APF and the guiding path effectively reduces the frequency of random movements during the early exploration phase of the DQN algorithm,which accelerates convergence,improves the computational efficiency of path planning,and increases path safety.Finally,the performance of the presented algorithm is validated through experiments in a 2D environment.Compared with traditional reinforcement learning methods such as Q-learning and Sarsa,as well as the original DQN algorithm,FG-DQN is more effective for USV path planning.展开更多
The wireless cloud robotic system(WCRS),which fully integrates sensing,communication,computing,and control capabilities as an intelligent agent,is a promising way to achieve intelligent manufacturing due to easy deplo...The wireless cloud robotic system(WCRS),which fully integrates sensing,communication,computing,and control capabilities as an intelligent agent,is a promising way to achieve intelligent manufacturing due to easy deployment and flexible expansion.However,the high-precision control of WCRS requires deterministic wireless communication,which is always challenging in the complex and dynamic radio space.This paper employs the reconfigurable intelligent surface(RIS)to establish a novel RIS-assisted WCRS architecture,where the radio channel is controlled to achieve ultra-reliable,low-delay,and low-jitter communication for high-precision closed-loop motion control.However,control and communication are strongly coupled and should be co-optimized.Fully considering the constraints of control input threshold,control delay deadline,beam phase,antenna power,and information distortion,we establish a stability maximization problem to jointly optimize control input compensation,RIS phase shift,and beamforming.Herein,a new jitter-oriented system stability objective with respect to control error and communication jitter is defined and the closed-form expression of control delay deadline is derived based on the Jensen Inequality and Lyapunov-Krasovskii functional.Due to the time-varying and partial observability of the channel and robot states,we model the problem as a partially observable Markov decision process(POMDP).To solve this complex problem,we propose a multi-agent transfer reinforcement learning algorithm named LSTM-PPO-MATRL,where the LSTM-enhanced proximal policy optimization(PPO)is designed to approximate an optimal solution and the option-guided policy transfer learning is proposed to facilitate the learning process.By centralized training and decentralized execution,LSTM-PPO-MATRL is validated by extensive experiments on MuJoCo tasks for both low-mobility and high-mobility robotic control scenarios.The results demonstrate that LSTM-PPO-MATRL not only realizes high learning efficiency,but also supports low-delay,low-jitter communication for low error control,where 71.9%control accuracy improvement and 68.7%delay jitter reduction are achieved compared to the PPO-MADRL baseline.展开更多
Cooperation,fairness,trust,and resource coordination are cornerstones of modern civilization,yet their emergence remains inadequately explained,largely due to persistent discrepancies between theoretical predictions a...Cooperation,fairness,trust,and resource coordination are cornerstones of modern civilization,yet their emergence remains inadequately explained,largely due to persistent discrepancies between theoretical predictions and behavioral experiments.Part of this gap may arise from the imitation learning paradigm commonly used in prior theoretical models,which assumes individuals merely copy successful neighbors according to predetermined,fixed rules.This review examines recent advances in evolutionary game dynamics that employ reinforcement learning(RL)as an alternative paradigm.In RL,individuals learn through trial and error and intro spec tively refine their strategies based on environmental feedback.We begin by introducing key concepts in evolutionary game theory and the two learning paradigms,then synthesize progress in applying RL to elucidate cooperation,trust,fairness,optimal resource coordination,and ecological dynamics.Collectively,these studies indicate that RL offers a promising unified framework for understanding the diverse social and ecological phenomena observed in human and natural systems.展开更多
Deep reinforcement learning(DRL)has demonstrated exceptional capabilities in combinatorial optimization,which automatically devises policies for solution construction and optimizer refinement.DRL is particularly adept...Deep reinforcement learning(DRL)has demonstrated exceptional capabilities in combinatorial optimization,which automatically devises policies for solution construction and optimizer refinement.DRL is particularly adept in generating training samples by itself,thereby providing the flexibility to solve a variety of combinatorial optimization problems without supervision.While DRL takes actions according to states extracted from problem-specific information,it cannot be directly applied to black-box continuous optimization lacking explicit information.To address this issue,this paper proposes a search space independent operator based DRL method for black-box continuous optimization.It conceptualizes the optimization process driven by search space independent operators as a Markov decision process,wherein actions are defined as operators and states are extracted from solutions generated by operators.In contrast to other DRLassisted metaheuristics,the proposed method does not rely on any existing metaheuristic.Instead,it innovates by creating totally new operators,able to surpass the performance boundaries of existing metaheuristics.Compared with state-of-the-art metaheuristics and DRL methods,the proposed method shows significantly faster convergence speed on challenging continuous optimization problems.展开更多
Intrinsic motivation serves as the predominant paradigm of exploration in reinforcement learning.In pursuit of an informative and robust state representation,the behavioural metric groups behaviourally equivalent stat...Intrinsic motivation serves as the predominant paradigm of exploration in reinforcement learning.In pursuit of an informative and robust state representation,the behavioural metric groups behaviourally equivalent states together,which share the same single-step reward and transition distribution.However,due to the presence of uninformative rewards and the dynamic nature of procedurally generated environments,these behavioural metric-based approaches could limit the effectiveness of the learnt state representations,potentially leading to a representation collapse and an ineffective exploration.Therefore,a more comprehensive and generalisable behavioural metric is needed to overcome the above issues.In this work,we approach the exploration problem from a novel perspective,extending beyond the conventional single-step assessments to encompass a longterm consideration of the whole trajectory.Specifically,we propose a novel trajectory-level behavioural metric(TBM)that exploits temporal dependencies of the trajectory and captures the underlying sequential information of behaviour patterns.To achieve an effective trajectory representation for exploration,we develop a pivotal state identifier(PSI)and a trajectory return estimator(TRE)to distinguish the diverse contributions of individual states in the trajectory.Moreover,an auxiliary representation regulariser is developed to promote the diversity and informativeness of the trajectory representation,mitigating the risk of representation mode collapse.Extensive experiments and empirical analysis conducted on procedurally generated environments showcase the superior performance of our proposed framework.展开更多
With the increasing maturity of multi-UAV technology and its broad applications in scenarios such as UAV roundup tasks,this paper proposes a novel approach to enhance interception efficiency and system robustness by a...With the increasing maturity of multi-UAV technology and its broad applications in scenarios such as UAV roundup tasks,this paper proposes a novel approach to enhance interception efficiency and system robustness by addressing insufficient historical data utilization and inadequate environmental explo-ration.The multi-UAV roundup problem is formulated as a Markov Decision Process(MDP),and an Improved Cross-Entropy Method with Intrinsic Curiosity-enhanced Multi-Agent Twin Delayed Deep Deterministic Policy Gradient(I2C-MATD3)is designed.Specifically,an Improved Cross-Entropy Method(ICEM)based on global elite samples rapidly optimizes training strategies while generating extensive experience for a Multi-Agent Twin Delayed Deep Deterministic Policy Gradient algorithm augmented with intrinsic curiosity rewards(IC-MATD3).In turn,IC-MATD3 guides the optimization direction of ICEM,enabling a synergistic interaction that facilitates effective historical data exploitation and pro-active environmental exploration for UAV agents to accomplish roundup tasks.Experiments in complex scenarios demonstrate that the proposed algorithm achieves superior training efficiency and conver-gence performance compared to state-of-the-art multi-agent reinforcement learning(MARL)methods.Robustness tests and ablation experiments further validate its enhanced generalizability and robustness.展开更多
As a special type of dynamic game,Pursuit-Evasion Games(PEGs)have expanded their application range from initial military confrontations to areas such as navigation control and aerospace,demonstrating broad applicabili...As a special type of dynamic game,Pursuit-Evasion Games(PEGs)have expanded their application range from initial military confrontations to areas such as navigation control and aerospace,demonstrating broad applicability and significant value in addressing a wide array of modern complex decision-making problems.Traditional optimal control methods based on differential game theory are classic approaches to solve PEG problems.However,these methods often struggle to perform well in complex environments,nonlinear systems,and situations involving highly uncertain participant behaviors.In recent years,rapidly developing Reinforcement Learning(RL)techniques has provided new avenues for PEG research.RL is capable of adapting to environmental changes through efficient online computation and feedback-driven learning,exhibiting strong generalization capabilities.Therefore,this survey presents a detailed and systematic review of PEG research based on RL methods.First,it classifies and discusses key RL algorithms and theoretical foundations in PEGs according to different forms of strategy learning.Then,it summarizes typical application scenarios,including tactical combat,unmanned systems control,and spacecraft interception,demonstrating the potential and effectiveness of RL in addressing real-world challenges.Finally,the survey explores current challenges and future opportunities in applying RL to PEGs,with the aim of promoting further research on more effective and practical solutions.展开更多
Carrier-borne aircraft are the primary formidable assets in aircraft carrier combat,and their sortie rate is a pivotal metric for evaluating the carrier's combat capability.Enhancing the efficiency of aircraft sup...Carrier-borne aircraft are the primary formidable assets in aircraft carrier combat,and their sortie rate is a pivotal metric for evaluating the carrier's combat capability.Enhancing the efficiency of aircraft support operations scheduling is a significant means to improve the sortie rate,where the central issue is to assign multi-wave aircraft to support stations according to the flight plan,and then obtain support resources for completing support operations.The existing studies primarily focus on considering partial operation processes(e.g.,ammunition transfer,disturbance handling,support personnel deployment,and deck arrangement),and lack modeling of the entire process of multi-wave aircraft support.In this paper,we investigate the multi-wave Aircraft Dispatch and Recovery Planning(ADRP)problem that aims to reasonably plan the operation processes,stations and resources of multi-wave aircraft with the goal of minimizing the total support operation time,and present a three-layer solution framework based on hierarchical reinforcement learning to address it.Specifically,we first abstract the multi-wave aircraft support operation planning process into the process layer,station layer and resource layer,and model it as a Decentralized Partially Observable Markov Decision Process(Dec-POMDP).Then,we propose a three-layer solution framework based on hierarchical reinforcement learning to solve the ADRP problem.To further improve planning results from long-term and global perspective,we design an inter-layer communication mechanism to allow efficient information exchange between stations.Extensive experimental results demonstrate that our proposed approach can achieve a high-quality operating schedule while meeting real-time demands.展开更多
Theintegration of human factors into artificial intelligence(AI)systems has emerged as a critical research frontier,particularly in reinforcement learning(RL),where human-AI interaction(HAII)presents both opportunitie...Theintegration of human factors into artificial intelligence(AI)systems has emerged as a critical research frontier,particularly in reinforcement learning(RL),where human-AI interaction(HAII)presents both opportunities and challenges.As RL continues to demonstrate remarkable success in model-free and partially observable environments,its real-world deployment increasingly requires effective collaboration with human operators and stakeholders.This article systematically examines HAII techniques in RL through both theoretical analysis and practical case studies.We establish a conceptual framework built upon three fundamental pillars of effective human-AI collaboration:computational trust modeling,system usability,and decision understandability.Our comprehensive review organizes HAII methods into five key categories:(1)learning from human feedback,including various shaping approaches;(2)learning from human demonstration through inverse RL and imitation learning;(3)shared autonomy architectures for dynamic control allocation;(4)human-in-the-loop querying strategies for active learning;and(5)explainable RL techniques for interpretable policy generation.Recent state-of-the-art works are critically reviewed,with particular emphasis on advances incorporating large language models in human-AI interaction research.To illustrate some concepts,we present three detailed case studies:an empirical trust model for farmers adopting AI-driven agricultural management systems,the implementation of ethical constraints in roboticmotion planning through human-guided RL,and an experimental investigation of human trust dynamics using a multi-armed bandit paradigm.These applications demonstrate how HAII principles can enhance RL systems’practical utility while bridging the gap between theoretical RL and real-world human-centered applications,ultimately contributing to more deployable and socially beneficial intelligent systems.展开更多
As a government-regulated public service,traffic signal control(TSC)requires reliable and transparent decision-making.However,existing deep reinforcement learning(DRL)methods,despite improvements in control accuracy,s...As a government-regulated public service,traffic signal control(TSC)requires reliable and transparent decision-making.However,existing deep reinforcement learning(DRL)methods,despite improvements in control accuracy,still lack explainability and generalisation,severely limiting their applicability in real-world environments.To address the challenges above,this paper proposes GenEx-TSC,a generalisable and explainable TSC method that integrates deep reinforcement learning with large language models(LLMs).First,starting from vehicle-level states,we train a DRL agent incorporating intersection physical heterogeneity and neighbourhood information,which lays the evaluation foundation for constructing a high-quality LLM dataset.Subsequently,the LLM agent is optimised through a two-stage training mechanism.In the distillation stage,a lightweight LLM agent is trained using the reasoning trajectories of a larger-scale LLM agent,inheriting its semantic understanding and decision-generation capabilities and in the alignment stage,the DRL evaluation network is employed to calibrate the outputs of the distilled LLM agent,ensuring that the generated cycle-level signal timing strategies are both efficient and interpretable.We synthesise 10 intersection networks with different physical attributes in SUMO and set traffic flows of varying scales.Experimental results across diverse traffic environments demonstrate that the proposed GenEx-TSC exhibits clear advantages over traditional methods,mainstream DRL methods and LLM baselines in terms of control accuracy,generalisation and explainability.展开更多
Intelligent routing plays a key role in modern communication infrastructure,including data centers,computing networks,and future 6G networks.Although reinforcement learning(RL)has shown great potential for intelligent...Intelligent routing plays a key role in modern communication infrastructure,including data centers,computing networks,and future 6G networks.Although reinforcement learning(RL)has shown great potential for intelligent routing,its practical deployment remains constrained by high energy consumption and decision latency.Here,we propose a photonic spiking RL architecture that implements a proximal policy optimization(PPO)–based intelligent routing algorithm.The performance of the proposed approach is systematically evaluated on a softwaredefined network(SDN)with a fat-tree topology.The results demonstrate that,under various baseline traffic rate conditions,the PPO-based routing strategy significantly outperforms the conventional Dijkstra algorithm in key performance metrics,including throughput,packet loss rate,average latency,and load balance.Furthermore,a hardware-software collaborative framework of the spiking Actor network is realized for three typical baseline traffic rates,utilizing a photonic synapse chip based on a Mach-Zehnder interferometer(MZI)array and a photonic spiking neuron chip based on distributed feedback lasers with a saturable absorber(DFB-SAs).Experimental validation on 640 state–action pairs shows that the inference accuracy of the hardware-software collaborative framework is consistent with that of the pure algorithmic implementation.The impacts of different hidden-layer scales in the spiking Actor network and varying network size of fat-tree topology are further analyzed.The integration of photonic spiking RL with SDN-based routing establishes a novel paradigm for intelligent routing optimization,featuring ultralow latency and high energy efficiency.This approach exhibits broad application prospects in real-time network optimization scenarios,including large-scale data centers,computing networks,satellite Internet systems,and future 6G networks.展开更多
With the advent of sixth-generation mobile communications(6G),space-air-ground integrated networks have become mainstream.This paper focuses on collaborative scheduling for mobile edge computing(MEC)under a three-tier...With the advent of sixth-generation mobile communications(6G),space-air-ground integrated networks have become mainstream.This paper focuses on collaborative scheduling for mobile edge computing(MEC)under a three-tier heterogeneous architecture composed of mobile devices,unmanned aerial vehicles(UAVs),and macro base stations(BSs).This scenario typically faces fast channel fading,dynamic computational loads,and energy constraints,whereas classical queuing-theoretic or convex-optimization approaches struggle to yield robust solutions in highly dynamic settings.To address this issue,we formulate a multi-agent Markov decision process(MDP)for an air-ground-fused MEC system,unify link selection,bandwidth/power allocation,and task offloading into a continuous action space and propose a joint scheduling strategy that is based on an improved MATD3 algorithm.The improvements include Alternating Layer Normalization(ALN)in the actor to suppress gradient variance,Residual Orthogonalization(RO)in the critic to reduce the correlation between the twin Q-value estimates,and a dynamic-temperature reward to enable adaptive trade-offs during training.On a multi-user,dual-link simulation platform,we conduct ablation and baseline comparisons.The results reveal that the proposed method has better convergence and stability.Compared with MADDPG,TD3,and DSAC,our algorithm achieves more robust performance across key metrics.展开更多
To meet the requirement of simultaneous arrival for multiple hypersonic glide vehicles(HGVs),we propose a time control entry guidance(TCEG)method leveraging deep reinforcement learning.First,the entry guidance problem...To meet the requirement of simultaneous arrival for multiple hypersonic glide vehicles(HGVs),we propose a time control entry guidance(TCEG)method leveraging deep reinforcement learning.First,the entry guidance problem is solved with a reinforcement learning framework based on a designed reference flight profile.By appropriately designing the observation space and training environment,the well-trained agent demonstrates robust guidance performance under varying widths of the heading error corridor.Then,a novel method for predicting the remaining flight time is established,which consists of two main components.The first component estimates the remaining flight time using an analytical formula,while the second component employs a deep neural network(DNN)to predict the residual error between the estimated and the true value.Subsequently,based on the predicted terminal time error,the threshold of the heading error and the observation vector are corrected in real time,thereby guiding the agent to dynamically adjust its output actions.This enables precise control of the terminal time.Since the generation of guidance commands only requires forward computations by the neural network,the proposed method exhibits excellent real-time performance.Finally,the effectiveness and robustness of the method are demonstrated through numerical simulations in various scenarios.展开更多
This paper presents HealthNet,a novel framework for the dynamic optimisation of healthcare transportation networks using multi-agent reinforcement learning.HealthNet leverages a spatiotemporal dependency module to cap...This paper presents HealthNet,a novel framework for the dynamic optimisation of healthcare transportation networks using multi-agent reinforcement learning.HealthNet leverages a spatiotemporal dependency module to capture complex spatiotemporal relationships in healthcare demand and resource allocation patterns,combined with centralised training and a decentralised execution approach.The system is modelled as a Markov game and solved using a deep reinforcement learning algorithm.Extensive simulations demonstrate that HealthNet outperforms eight state-of-the-art baseline methods across multiple network configurations and evaluation metrics.In a 4×4 grid network,HealthNet reduces average waiting times by 47.6%compared to model predictive control and 22.1%compared to the best-performing baseline.Traffic congestion rates are reduced to 16.7%compared to 42.3%for the worst baseline and 23.1%for the best baseline.Under irregular network topologies with stochastic disruptions,including demand surges and vehicle unavailability,HealthNet maintains superior performance with 42.1%lower average waiting time and 51.1%improvement in peak response times compared to competing approaches.These findings indicate that HealthNet can enhance both efficiency and resilience in healthcare transportation systems,potentially improving patient outcomes in complex urban environments.展开更多
While reinforcement learning-based underwater acoustic adaptive modulation shows promise for enabling environment-adaptive communication as supported by extensive simulation-based research,its practical performance re...While reinforcement learning-based underwater acoustic adaptive modulation shows promise for enabling environment-adaptive communication as supported by extensive simulation-based research,its practical performance remains underexplored in field investigations.To evaluate the practical applicability of this emerging technique in adverse shallow sea channels,a field experiment was conducted using three communication modes:orthogonal frequency division multiplexing(OFDM),M-ary frequency-shift keying(MFSK),and direct sequence spread spectrum(DSSS)for reinforcement learning-driven adaptive modulation.Specifically,a Q-learning method is used to select the optimal modulation mode according to the channel quality quantified by signal-to-noise ratio,multipath spread length,and Doppler frequency offset.Experimental results demonstrate that the reinforcement learning-based adaptive modulation scheme outperformed fixed threshold detection in terms of total throughput and average bit error rate,surpassing conventional adaptive modulation strategies.展开更多
Dear Editor,This letter deals with the security control for nonlinear cyber-physical systems(CPSs)under mixed deception attacks.Both sensors and actuators are assumed to be injected deception data during the data tran...Dear Editor,This letter deals with the security control for nonlinear cyber-physical systems(CPSs)under mixed deception attacks.Both sensors and actuators are assumed to be injected deception data during the data transmission via networks.In order to identify the unknown dynamics of the attacked system,a neural network(NN)is adopted,on basis of which an NN-based secure observer is designed to diminish the attack impact on state estimation.Then,by resorting to the reinforcement learning approach,the secure control strategy is presented via actor-critic and zero-sum games.At last,the designed control scheme is proved via a numerical simulation.展开更多
Vehicle Edge Computing(VEC)and Cloud Computing(CC)significantly enhance the processing efficiency of delay-sensitive and computation-intensive applications by offloading compute-intensive tasks from resource-constrain...Vehicle Edge Computing(VEC)and Cloud Computing(CC)significantly enhance the processing efficiency of delay-sensitive and computation-intensive applications by offloading compute-intensive tasks from resource-constrained onboard devices to nearby Roadside Unit(RSU),thereby achieving lower delay and energy consumption.However,due to the limited storage capacity and energy budget of RSUs,it is challenging to meet the demands of the highly dynamic Internet of Vehicles(IoV)environment.Therefore,determining reasonable service caching and computation offloading strategies is crucial.To address this,this paper proposes a joint service caching scheme for cloud-edge collaborative IoV computation offloading.By modeling the dynamic optimization problem using Markov Decision Processes(MDP),the scheme jointly optimizes task delay,energy consumption,load balancing,and privacy entropy to achieve better quality of service.Additionally,a dynamic adaptive multi-objective deep reinforcement learning algorithm is proposed.Each Double Deep Q-Network(DDQN)agent obtains rewards for different objectives based on distinct reward functions and dynamically updates the objective weights by learning the value changes between objectives using Radial Basis Function Networks(RBFN),thereby efficiently approximating the Pareto-optimal decisions for multiple objectives.Extensive experiments demonstrate that the proposed algorithm can better coordinate the three-tier computing resources of cloud,edge,and vehicles.Compared to existing algorithms,the proposed method reduces task delay and energy consumption by 10.64%and 5.1%,respectively.展开更多
Effective partitioning is crucial for enabling parallel restoration of power systems after blackouts.This paper proposes a novel partitioning method based on deep reinforcement learning.First,the partitioning decision...Effective partitioning is crucial for enabling parallel restoration of power systems after blackouts.This paper proposes a novel partitioning method based on deep reinforcement learning.First,the partitioning decision process is formulated as a Markov decision process(MDP)model to maximize the modularity.Corresponding key partitioning constraints on parallel restoration are considered.Second,based on the partitioning objective and constraints,the reward function of the partitioning MDP model is set by adopting a relative deviation normalization scheme to reduce mutual interference between the reward and penalty in the reward function.The soft bonus scaling mechanism is introduced to mitigate overestimation caused by abrupt jumps in the reward.Then,the deep Q network method is applied to solve the partitioning MDP model and generate partitioning schemes.Two experience replay buffers are employed to speed up the training process of the method.Finally,case studies on the IEEE 39-bus test system demonstrate that the proposed method can generate a high-modularity partitioning result that meets all key partitioning constraints,thereby improving the parallelism and reliability of the restoration process.Moreover,simulation results demonstrate that an appropriate discount factor is crucial for ensuring both the convergence speed and the stability of the partitioning training.展开更多
基金funded by Hung Yen University of Technology and Education under grand number UTEHY.L.2025.62.
摘要Unmanned Aerial Vehicles(UAVs)have become integral components in smart city infrastructures,supporting applications such as emergency response,surveillance,and data collection.However,the high mobility and dynamic topology of Flying Ad Hoc Networks(FANETs)present significant challenges for maintaining reliable,low-latency communication.Conventional geographic routing protocols often struggle in situations where link quality varies and mobility patterns are unpredictable.To overcome these limitations,this paper proposes an improved routing protocol based on reinforcement learning.This new approach integrates Q-learning with mechanisms that are both link-aware and mobility-aware.The proposed method optimizes the selection of relay nodes by using an adaptive reward function that takes into account energy consumption,delay,and link quality.Additionally,a Kalman filter is integrated to predict UAV mobility,improving the stability of communication links under dynamic network conditions.Simulation experiments were conducted using realistic scenarios,varying the number of UAVs to assess scalability.An analysis was conducted on key performance metrics,including the packet delivery ratio,end-to-end delay,and total energy consumption.The results demonstrate that the proposed approach significantly improves the packet delivery ratio by 12%–15%and reduces delay by up to 25.5%when compared to conventional GEO and QGEO protocols.However,this improvement comes at the cost of higher energy consumption due to additional computations and control overhead.Despite this trade-off,the proposed solution ensures reliable and efficient communication,making it well-suited for large-scale UAV networks operating in complex urban environments.
基金Supported by the Science Research Foundation for Introduced Talents,Fujian Province of China under Grant Nos.GY-Z21215,GY-Z21216.
摘要For unmanned surface vehicles(USVs),how to find an effective,feasible path that substantially improves mission success rates and time efficiency in dynamic marine environments is a critical issue.To address the path planning problem for USVs using deep reinforcement learning(DRL)in dynamic ocean environments,an improved algorithm based on Deep Q-Networks(DQN)is proposed,which is called Fast Guided Deep Q-Network Algorithm(FG-DQN).This algorithm combines DQN with the artificial potential field(APF)method and uses the A*algorithm to initialize a guiding path in a global static environment and to provide prior knowledge for the USVs.Additionally,the configuration of the reward function using APF and the guiding path effectively reduces the frequency of random movements during the early exploration phase of the DQN algorithm,which accelerates convergence,improves the computational efficiency of path planning,and increases path safety.Finally,the performance of the presented algorithm is validated through experiments in a 2D environment.Compared with traditional reinforcement learning methods such as Q-learning and Sarsa,as well as the original DQN algorithm,FG-DQN is more effective for USV path planning.
基金supported in part by the National Natural Science Foundation of China(62522320,92267108,62173322)Liaoning Revitalization Talents Program(XLYC2403062)the Science and Technology Program of Liaoning Province(2023JH3/10200004,2022JH25/10100005)。
摘要The wireless cloud robotic system(WCRS),which fully integrates sensing,communication,computing,and control capabilities as an intelligent agent,is a promising way to achieve intelligent manufacturing due to easy deployment and flexible expansion.However,the high-precision control of WCRS requires deterministic wireless communication,which is always challenging in the complex and dynamic radio space.This paper employs the reconfigurable intelligent surface(RIS)to establish a novel RIS-assisted WCRS architecture,where the radio channel is controlled to achieve ultra-reliable,low-delay,and low-jitter communication for high-precision closed-loop motion control.However,control and communication are strongly coupled and should be co-optimized.Fully considering the constraints of control input threshold,control delay deadline,beam phase,antenna power,and information distortion,we establish a stability maximization problem to jointly optimize control input compensation,RIS phase shift,and beamforming.Herein,a new jitter-oriented system stability objective with respect to control error and communication jitter is defined and the closed-form expression of control delay deadline is derived based on the Jensen Inequality and Lyapunov-Krasovskii functional.Due to the time-varying and partial observability of the channel and robot states,we model the problem as a partially observable Markov decision process(POMDP).To solve this complex problem,we propose a multi-agent transfer reinforcement learning algorithm named LSTM-PPO-MATRL,where the LSTM-enhanced proximal policy optimization(PPO)is designed to approximate an optimal solution and the option-guided policy transfer learning is proposed to facilitate the learning process.By centralized training and decentralized execution,LSTM-PPO-MATRL is validated by extensive experiments on MuJoCo tasks for both low-mobility and high-mobility robotic control scenarios.The results demonstrate that LSTM-PPO-MATRL not only realizes high learning efficiency,but also supports low-delay,low-jitter communication for low error control,where 71.9%control accuracy improvement and 68.7%delay jitter reduction are achieved compared to the PPO-MADRL baseline.
基金supported by the National Natural Science Foundation of China(Grants Nos.12075144,12165014)the Fundamental Research Funds for the Central Universities(Grant No.GK202401002)the Key Research and Development Program of Ningxia in China(Grant No.2021BEB04032)。
摘要Cooperation,fairness,trust,and resource coordination are cornerstones of modern civilization,yet their emergence remains inadequately explained,largely due to persistent discrepancies between theoretical predictions and behavioral experiments.Part of this gap may arise from the imitation learning paradigm commonly used in prior theoretical models,which assumes individuals merely copy successful neighbors according to predetermined,fixed rules.This review examines recent advances in evolutionary game dynamics that employ reinforcement learning(RL)as an alternative paradigm.In RL,individuals learn through trial and error and intro spec tively refine their strategies based on environmental feedback.We begin by introducing key concepts in evolutionary game theory and the two learning paradigms,then synthesize progress in applying RL to elucidate cooperation,trust,fairness,optimal resource coordination,and ecological dynamics.Collectively,these studies indicate that RL offers a promising unified framework for understanding the diverse social and ecological phenomena observed in human and natural systems.
基金supported in part by the National Natural Science Foundation of China(62136008,62276001,U21A20512,W2441019)the Anhui Provincial Natural Science Foundation(2308085J03)the Excellent Youth Foundation of Anhui Provincial Colleges(2022AH030013)。
摘要Deep reinforcement learning(DRL)has demonstrated exceptional capabilities in combinatorial optimization,which automatically devises policies for solution construction and optimizer refinement.DRL is particularly adept in generating training samples by itself,thereby providing the flexibility to solve a variety of combinatorial optimization problems without supervision.While DRL takes actions according to states extracted from problem-specific information,it cannot be directly applied to black-box continuous optimization lacking explicit information.To address this issue,this paper proposes a search space independent operator based DRL method for black-box continuous optimization.It conceptualizes the optimization process driven by search space independent operators as a Markov decision process,wherein actions are defined as operators and states are extracted from solutions generated by operators.In contrast to other DRLassisted metaheuristics,the proposed method does not rely on any existing metaheuristic.Instead,it innovates by creating totally new operators,able to surpass the performance boundaries of existing metaheuristics.Compared with state-of-the-art metaheuristics and DRL methods,the proposed method shows significantly faster convergence speed on challenging continuous optimization problems.
基金supported by the National Natural Science Foundation of China(Grant 62276047)Sichuan Science and Technology Programme(Grant 2025HJRC0021).
摘要Intrinsic motivation serves as the predominant paradigm of exploration in reinforcement learning.In pursuit of an informative and robust state representation,the behavioural metric groups behaviourally equivalent states together,which share the same single-step reward and transition distribution.However,due to the presence of uninformative rewards and the dynamic nature of procedurally generated environments,these behavioural metric-based approaches could limit the effectiveness of the learnt state representations,potentially leading to a representation collapse and an ineffective exploration.Therefore,a more comprehensive and generalisable behavioural metric is needed to overcome the above issues.In this work,we approach the exploration problem from a novel perspective,extending beyond the conventional single-step assessments to encompass a longterm consideration of the whole trajectory.Specifically,we propose a novel trajectory-level behavioural metric(TBM)that exploits temporal dependencies of the trajectory and captures the underlying sequential information of behaviour patterns.To achieve an effective trajectory representation for exploration,we develop a pivotal state identifier(PSI)and a trajectory return estimator(TRE)to distinguish the diverse contributions of individual states in the trajectory.Moreover,an auxiliary representation regulariser is developed to promote the diversity and informativeness of the trajectory representation,mitigating the risk of representation mode collapse.Extensive experiments and empirical analysis conducted on procedurally generated environments showcase the superior performance of our proposed framework.
基金the National Key Lab-oratory of Air-based Information Perception and Fusion(Grant No.202510)the Key Research and Development Program of Shaanxi Province(Grant No.2023-GHZD-33)+2 种基金the Fundamental Research Funds for the Central Universities(Grant No.H20250607)the Open Project of the State Key Laboratory of Intelligent Game(Grant No.ZBKF-23-05)the National Nature Science Foundation of China(Grant No.62003267)to provide fund for conducting experiments.
摘要With the increasing maturity of multi-UAV technology and its broad applications in scenarios such as UAV roundup tasks,this paper proposes a novel approach to enhance interception efficiency and system robustness by addressing insufficient historical data utilization and inadequate environmental explo-ration.The multi-UAV roundup problem is formulated as a Markov Decision Process(MDP),and an Improved Cross-Entropy Method with Intrinsic Curiosity-enhanced Multi-Agent Twin Delayed Deep Deterministic Policy Gradient(I2C-MATD3)is designed.Specifically,an Improved Cross-Entropy Method(ICEM)based on global elite samples rapidly optimizes training strategies while generating extensive experience for a Multi-Agent Twin Delayed Deep Deterministic Policy Gradient algorithm augmented with intrinsic curiosity rewards(IC-MATD3).In turn,IC-MATD3 guides the optimization direction of ICEM,enabling a synergistic interaction that facilitates effective historical data exploitation and pro-active environmental exploration for UAV agents to accomplish roundup tasks.Experiments in complex scenarios demonstrate that the proposed algorithm achieves superior training efficiency and conver-gence performance compared to state-of-the-art multi-agent reinforcement learning(MARL)methods.Robustness tests and ablation experiments further validate its enhanced generalizability and robustness.
基金supported by the National Science and Technology Major Project,China(No.2022ZD0119703)the National Natural Science Foundation of China(No.62273044)the National Natural Science Foundation of China National Science Fund for Distinguished Young Scholars(No.62025301)。
摘要As a special type of dynamic game,Pursuit-Evasion Games(PEGs)have expanded their application range from initial military confrontations to areas such as navigation control and aerospace,demonstrating broad applicability and significant value in addressing a wide array of modern complex decision-making problems.Traditional optimal control methods based on differential game theory are classic approaches to solve PEG problems.However,these methods often struggle to perform well in complex environments,nonlinear systems,and situations involving highly uncertain participant behaviors.In recent years,rapidly developing Reinforcement Learning(RL)techniques has provided new avenues for PEG research.RL is capable of adapting to environmental changes through efficient online computation and feedback-driven learning,exhibiting strong generalization capabilities.Therefore,this survey presents a detailed and systematic review of PEG research based on RL methods.First,it classifies and discusses key RL algorithms and theoretical foundations in PEGs according to different forms of strategy learning.Then,it summarizes typical application scenarios,including tactical combat,unmanned systems control,and spacecraft interception,demonstrating the potential and effectiveness of RL in addressing real-world challenges.Finally,the survey explores current challenges and future opportunities in applying RL to PEGs,with the aim of promoting further research on more effective and practical solutions.
基金co-supported by the National Natural Science Foundation of China(Nos.62325602,62036010 and 62372416)the National Natural Science Foundation of Henan Province,China(No.242300421215)the Key Scientific Research Project of Higher Education Institutions of Henan Province,China(No.25B520021)。
摘要Carrier-borne aircraft are the primary formidable assets in aircraft carrier combat,and their sortie rate is a pivotal metric for evaluating the carrier's combat capability.Enhancing the efficiency of aircraft support operations scheduling is a significant means to improve the sortie rate,where the central issue is to assign multi-wave aircraft to support stations according to the flight plan,and then obtain support resources for completing support operations.The existing studies primarily focus on considering partial operation processes(e.g.,ammunition transfer,disturbance handling,support personnel deployment,and deck arrangement),and lack modeling of the entire process of multi-wave aircraft support.In this paper,we investigate the multi-wave Aircraft Dispatch and Recovery Planning(ADRP)problem that aims to reasonably plan the operation processes,stations and resources of multi-wave aircraft with the goal of minimizing the total support operation time,and present a three-layer solution framework based on hierarchical reinforcement learning to address it.Specifically,we first abstract the multi-wave aircraft support operation planning process into the process layer,station layer and resource layer,and model it as a Decentralized Partially Observable Markov Decision Process(Dec-POMDP).Then,we propose a three-layer solution framework based on hierarchical reinforcement learning to solve the ADRP problem.To further improve planning results from long-term and global perspective,we design an inter-layer communication mechanism to allow efficient information exchange between stations.Extensive experimental results demonstrate that our proposed approach can achieve a high-quality operating schedule while meeting real-time demands.
基金funded by the U.S.Department of Education under Grant Number ED#P116S210005the National Science Foundation under Grant Numbers 2226936 and 2420405.
摘要Theintegration of human factors into artificial intelligence(AI)systems has emerged as a critical research frontier,particularly in reinforcement learning(RL),where human-AI interaction(HAII)presents both opportunities and challenges.As RL continues to demonstrate remarkable success in model-free and partially observable environments,its real-world deployment increasingly requires effective collaboration with human operators and stakeholders.This article systematically examines HAII techniques in RL through both theoretical analysis and practical case studies.We establish a conceptual framework built upon three fundamental pillars of effective human-AI collaboration:computational trust modeling,system usability,and decision understandability.Our comprehensive review organizes HAII methods into five key categories:(1)learning from human feedback,including various shaping approaches;(2)learning from human demonstration through inverse RL and imitation learning;(3)shared autonomy architectures for dynamic control allocation;(4)human-in-the-loop querying strategies for active learning;and(5)explainable RL techniques for interpretable policy generation.Recent state-of-the-art works are critically reviewed,with particular emphasis on advances incorporating large language models in human-AI interaction research.To illustrate some concepts,we present three detailed case studies:an empirical trust model for farmers adopting AI-driven agricultural management systems,the implementation of ethical constraints in roboticmotion planning through human-guided RL,and an experimental investigation of human trust dynamics using a multi-armed bandit paradigm.These applications demonstrate how HAII principles can enhance RL systems’practical utility while bridging the gap between theoretical RL and real-world human-centered applications,ultimately contributing to more deployable and socially beneficial intelligent systems.
基金the National Natural Science Foundation of China under(Grant No.62501094)in part by the Natural Science Foundation of Chongqing under(Grant Nos.CSTB2025NSCQLZX0152,CSTB2024NSCQ-LZX0134 and CSTB2025NSCQ-LZX0052).
摘要As a government-regulated public service,traffic signal control(TSC)requires reliable and transparent decision-making.However,existing deep reinforcement learning(DRL)methods,despite improvements in control accuracy,still lack explainability and generalisation,severely limiting their applicability in real-world environments.To address the challenges above,this paper proposes GenEx-TSC,a generalisable and explainable TSC method that integrates deep reinforcement learning with large language models(LLMs).First,starting from vehicle-level states,we train a DRL agent incorporating intersection physical heterogeneity and neighbourhood information,which lays the evaluation foundation for constructing a high-quality LLM dataset.Subsequently,the LLM agent is optimised through a two-stage training mechanism.In the distillation stage,a lightweight LLM agent is trained using the reasoning trajectories of a larger-scale LLM agent,inheriting its semantic understanding and decision-generation capabilities and in the alignment stage,the DRL evaluation network is employed to calibrate the outputs of the distilled LLM agent,ensuring that the generated cycle-level signal timing strategies are both efficient and interpretable.We synthesise 10 intersection networks with different physical attributes in SUMO and set traffic flows of varying scales.Experimental results across diverse traffic environments demonstrate that the proposed GenEx-TSC exhibits clear advantages over traditional methods,mainstream DRL methods and LLM baselines in terms of control accuracy,generalisation and explainability.
基金supports from the National Natural Science Foundation of China(No.62535015,62575231)the Fundamental Research Funds for the Central Universities(QTZX23041)Xidian University Specially Funded Project for Interdisciplinary Exploration(TZJH2024009).
摘要Intelligent routing plays a key role in modern communication infrastructure,including data centers,computing networks,and future 6G networks.Although reinforcement learning(RL)has shown great potential for intelligent routing,its practical deployment remains constrained by high energy consumption and decision latency.Here,we propose a photonic spiking RL architecture that implements a proximal policy optimization(PPO)–based intelligent routing algorithm.The performance of the proposed approach is systematically evaluated on a softwaredefined network(SDN)with a fat-tree topology.The results demonstrate that,under various baseline traffic rate conditions,the PPO-based routing strategy significantly outperforms the conventional Dijkstra algorithm in key performance metrics,including throughput,packet loss rate,average latency,and load balance.Furthermore,a hardware-software collaborative framework of the spiking Actor network is realized for three typical baseline traffic rates,utilizing a photonic synapse chip based on a Mach-Zehnder interferometer(MZI)array and a photonic spiking neuron chip based on distributed feedback lasers with a saturable absorber(DFB-SAs).Experimental validation on 640 state–action pairs shows that the inference accuracy of the hardware-software collaborative framework is consistent with that of the pure algorithmic implementation.The impacts of different hidden-layer scales in the spiking Actor network and varying network size of fat-tree topology are further analyzed.The integration of photonic spiking RL with SDN-based routing establishes a novel paradigm for intelligent routing optimization,featuring ultralow latency and high energy efficiency.This approach exhibits broad application prospects in real-time network optimization scenarios,including large-scale data centers,computing networks,satellite Internet systems,and future 6G networks.
摘要With the advent of sixth-generation mobile communications(6G),space-air-ground integrated networks have become mainstream.This paper focuses on collaborative scheduling for mobile edge computing(MEC)under a three-tier heterogeneous architecture composed of mobile devices,unmanned aerial vehicles(UAVs),and macro base stations(BSs).This scenario typically faces fast channel fading,dynamic computational loads,and energy constraints,whereas classical queuing-theoretic or convex-optimization approaches struggle to yield robust solutions in highly dynamic settings.To address this issue,we formulate a multi-agent Markov decision process(MDP)for an air-ground-fused MEC system,unify link selection,bandwidth/power allocation,and task offloading into a continuous action space and propose a joint scheduling strategy that is based on an improved MATD3 algorithm.The improvements include Alternating Layer Normalization(ALN)in the actor to suppress gradient variance,Residual Orthogonalization(RO)in the critic to reduce the correlation between the twin Q-value estimates,and a dynamic-temperature reward to enable adaptive trade-offs during training.On a multi-user,dual-link simulation platform,we conduct ablation and baseline comparisons.The results reveal that the proposed method has better convergence and stability.Compared with MADDPG,TD3,and DSAC,our algorithm achieves more robust performance across key metrics.
基金supported by the National Natural Science Foundation of China(No.62103432)the Open Fund of Key Laboratory of Cross-Domain Flight Interdisciplinary Technology(No.2024-KYKF-4004),China.
摘要To meet the requirement of simultaneous arrival for multiple hypersonic glide vehicles(HGVs),we propose a time control entry guidance(TCEG)method leveraging deep reinforcement learning.First,the entry guidance problem is solved with a reinforcement learning framework based on a designed reference flight profile.By appropriately designing the observation space and training environment,the well-trained agent demonstrates robust guidance performance under varying widths of the heading error corridor.Then,a novel method for predicting the remaining flight time is established,which consists of two main components.The first component estimates the remaining flight time using an analytical formula,while the second component employs a deep neural network(DNN)to predict the residual error between the estimated and the true value.Subsequently,based on the predicted terminal time error,the threshold of the heading error and the observation vector are corrected in real time,thereby guiding the agent to dynamically adjust its output actions.This enables precise control of the terminal time.Since the generation of guidance commands only requires forward computations by the neural network,the proposed method exhibits excellent real-time performance.Finally,the effectiveness and robustness of the method are demonstrated through numerical simulations in various scenarios.
基金supported by the National Natural Science Foundation of China under No.62202247.
摘要This paper presents HealthNet,a novel framework for the dynamic optimisation of healthcare transportation networks using multi-agent reinforcement learning.HealthNet leverages a spatiotemporal dependency module to capture complex spatiotemporal relationships in healthcare demand and resource allocation patterns,combined with centralised training and a decentralised execution approach.The system is modelled as a Markov game and solved using a deep reinforcement learning algorithm.Extensive simulations demonstrate that HealthNet outperforms eight state-of-the-art baseline methods across multiple network configurations and evaluation metrics.In a 4×4 grid network,HealthNet reduces average waiting times by 47.6%compared to model predictive control and 22.1%compared to the best-performing baseline.Traffic congestion rates are reduced to 16.7%compared to 42.3%for the worst baseline and 23.1%for the best baseline.Under irregular network topologies with stochastic disruptions,including demand surges and vehicle unavailability,HealthNet maintains superior performance with 42.1%lower average waiting time and 51.1%improvement in peak response times compared to competing approaches.These findings indicate that HealthNet can enhance both efficiency and resilience in healthcare transportation systems,potentially improving patient outcomes in complex urban environments.
基金funding from the National Key Research and Development Program of China(No.2018YFE0110000)the National Natural Science Foundation of China(No.11274259,No.11574258)the Science and Technology Commission Foundation of Shanghai(21DZ1205500)in support of the present research.
摘要While reinforcement learning-based underwater acoustic adaptive modulation shows promise for enabling environment-adaptive communication as supported by extensive simulation-based research,its practical performance remains underexplored in field investigations.To evaluate the practical applicability of this emerging technique in adverse shallow sea channels,a field experiment was conducted using three communication modes:orthogonal frequency division multiplexing(OFDM),M-ary frequency-shift keying(MFSK),and direct sequence spread spectrum(DSSS)for reinforcement learning-driven adaptive modulation.Specifically,a Q-learning method is used to select the optimal modulation mode according to the channel quality quantified by signal-to-noise ratio,multipath spread length,and Doppler frequency offset.Experimental results demonstrate that the reinforcement learning-based adaptive modulation scheme outperformed fixed threshold detection in terms of total throughput and average bit error rate,surpassing conventional adaptive modulation strategies.
基金supported in part by the National Natural Science Foundation of China(62273180,62403245,62233012)Natural Science Foundation of Jiangsu Province of China(BK20241458,BK20232038)。
摘要Dear Editor,This letter deals with the security control for nonlinear cyber-physical systems(CPSs)under mixed deception attacks.Both sensors and actuators are assumed to be injected deception data during the data transmission via networks.In order to identify the unknown dynamics of the attacked system,a neural network(NN)is adopted,on basis of which an NN-based secure observer is designed to diminish the attack impact on state estimation.Then,by resorting to the reinforcement learning approach,the secure control strategy is presented via actor-critic and zero-sum games.At last,the designed control scheme is proved via a numerical simulation.
基金supported by Key Science and Technology Program of Henan Province,China(Grant Nos.242102210147,242102210027)Fujian Province Young and Middle aged Teacher Education Research Project(Science and Technology Category)(No.JZ240101)(Corresponding author:Dong Yuan).
摘要Vehicle Edge Computing(VEC)and Cloud Computing(CC)significantly enhance the processing efficiency of delay-sensitive and computation-intensive applications by offloading compute-intensive tasks from resource-constrained onboard devices to nearby Roadside Unit(RSU),thereby achieving lower delay and energy consumption.However,due to the limited storage capacity and energy budget of RSUs,it is challenging to meet the demands of the highly dynamic Internet of Vehicles(IoV)environment.Therefore,determining reasonable service caching and computation offloading strategies is crucial.To address this,this paper proposes a joint service caching scheme for cloud-edge collaborative IoV computation offloading.By modeling the dynamic optimization problem using Markov Decision Processes(MDP),the scheme jointly optimizes task delay,energy consumption,load balancing,and privacy entropy to achieve better quality of service.Additionally,a dynamic adaptive multi-objective deep reinforcement learning algorithm is proposed.Each Double Deep Q-Network(DDQN)agent obtains rewards for different objectives based on distinct reward functions and dynamically updates the objective weights by learning the value changes between objectives using Radial Basis Function Networks(RBFN),thereby efficiently approximating the Pareto-optimal decisions for multiple objectives.Extensive experiments demonstrate that the proposed algorithm can better coordinate the three-tier computing resources of cloud,edge,and vehicles.Compared to existing algorithms,the proposed method reduces task delay and energy consumption by 10.64%and 5.1%,respectively.
基金funded by the Beijing Engineering Research Center of Electric Rail Transportation.
摘要Effective partitioning is crucial for enabling parallel restoration of power systems after blackouts.This paper proposes a novel partitioning method based on deep reinforcement learning.First,the partitioning decision process is formulated as a Markov decision process(MDP)model to maximize the modularity.Corresponding key partitioning constraints on parallel restoration are considered.Second,based on the partitioning objective and constraints,the reward function of the partitioning MDP model is set by adopting a relative deviation normalization scheme to reduce mutual interference between the reward and penalty in the reward function.The soft bonus scaling mechanism is introduced to mitigate overestimation caused by abrupt jumps in the reward.Then,the deep Q network method is applied to solve the partitioning MDP model and generate partitioning schemes.Two experience replay buffers are employed to speed up the training process of the method.Finally,case studies on the IEEE 39-bus test system demonstrate that the proposed method can generate a high-modularity partitioning result that meets all key partitioning constraints,thereby improving the parallelism and reliability of the restoration process.Moreover,simulation results demonstrate that an appropriate discount factor is crucial for ensuring both the convergence speed and the stability of the partitioning training.