期刊文献+
共找到10,214篇文章
< 1 2 250 >
每页显示 20 50 100
An Improved Reinforcement Learning-Based 6G UAV Communication for Smart Cities 认领 引用 被引量:1
1
作者 Vi Hoai Nam Chu Thi Minh Hue Dang Van Anh 《Computers, Materials & Continua》 SCIE EI 2026年第1期2030-2044,共15页
Unmanned Aerial Vehicles(UAVs)have become integral components in smart city infrastructures,supporting applications such as emergency response,surveillance,and data collection.However,the high mobility and dynamic top... Unmanned Aerial Vehicles(UAVs)have become integral components in smart city infrastructures,supporting applications such as emergency response,surveillance,and data collection.However,the high mobility and dynamic topology of Flying Ad Hoc Networks(FANETs)present significant challenges for maintaining reliable,low-latency communication.Conventional geographic routing protocols often struggle in situations where link quality varies and mobility patterns are unpredictable.To overcome these limitations,this paper proposes an improved routing protocol based on reinforcement learning.This new approach integrates Q-learning with mechanisms that are both link-aware and mobility-aware.The proposed method optimizes the selection of relay nodes by using an adaptive reward function that takes into account energy consumption,delay,and link quality.Additionally,a Kalman filter is integrated to predict UAV mobility,improving the stability of communication links under dynamic network conditions.Simulation experiments were conducted using realistic scenarios,varying the number of UAVs to assess scalability.An analysis was conducted on key performance metrics,including the packet delivery ratio,end-to-end delay,and total energy consumption.The results demonstrate that the proposed approach significantly improves the packet delivery ratio by 12%–15%and reduces delay by up to 25.5%when compared to conventional GEO and QGEO protocols.However,this improvement comes at the cost of higher energy consumption due to additional computations and control overhead.Despite this trade-off,the proposed solution ensures reliable and efficient communication,making it well-suited for large-scale UAV networks operating in complex urban environments. 展开更多
关键词 UAV FANET smart cities reinforcement learning Q-learning
暂未订购 下载PDF
Path Planning for Unmanned Surface Vehicles in Dynamic Environments Based on Artificial Potential Field and Global Guided Reinforcement Learning 认领 引用 被引量:2
2
作者 Shanqiang Li Chaoxi Li 《哈尔滨工程大学学报(英文版)》 CSCD 2026年第2期575-586,共12页
For unmanned surface vehicles(USVs),how to find an effective,feasible path that substantially improves mission success rates and time efficiency in dynamic marine environments is a critical issue.To address the path p... For unmanned surface vehicles(USVs),how to find an effective,feasible path that substantially improves mission success rates and time efficiency in dynamic marine environments is a critical issue.To address the path planning problem for USVs using deep reinforcement learning(DRL)in dynamic ocean environments,an improved algorithm based on Deep Q-Networks(DQN)is proposed,which is called Fast Guided Deep Q-Network Algorithm(FG-DQN).This algorithm combines DQN with the artificial potential field(APF)method and uses the A*algorithm to initialize a guiding path in a global static environment and to provide prior knowledge for the USVs.Additionally,the configuration of the reward function using APF and the guiding path effectively reduces the frequency of random movements during the early exploration phase of the DQN algorithm,which accelerates convergence,improves the computational efficiency of path planning,and increases path safety.Finally,the performance of the presented algorithm is validated through experiments in a 2D environment.Compared with traditional reinforcement learning methods such as Q-learning and Sarsa,as well as the original DQN algorithm,FG-DQN is more effective for USV path planning. 展开更多
关键词 Deep reinforcement learning Path planning Unmanned surface vehicles Fast guided deep Q-Network algorithm
暂未订购 下载PDF
Control-Communication Co-Optimization for Wireless Cloud Robotic System via Multi-Agent Transfer Reinforcement Learning 认领 引用 被引量:1
3
作者 Chi Xu Junyuan Zhang Haibin Yu 《IEEE/CAA Journal of Automatica Sinica》 SCIE EI CSCD 2026年第2期311-326,共16页
The wireless cloud robotic system(WCRS),which fully integrates sensing,communication,computing,and control capabilities as an intelligent agent,is a promising way to achieve intelligent manufacturing due to easy deplo... The wireless cloud robotic system(WCRS),which fully integrates sensing,communication,computing,and control capabilities as an intelligent agent,is a promising way to achieve intelligent manufacturing due to easy deployment and flexible expansion.However,the high-precision control of WCRS requires deterministic wireless communication,which is always challenging in the complex and dynamic radio space.This paper employs the reconfigurable intelligent surface(RIS)to establish a novel RIS-assisted WCRS architecture,where the radio channel is controlled to achieve ultra-reliable,low-delay,and low-jitter communication for high-precision closed-loop motion control.However,control and communication are strongly coupled and should be co-optimized.Fully considering the constraints of control input threshold,control delay deadline,beam phase,antenna power,and information distortion,we establish a stability maximization problem to jointly optimize control input compensation,RIS phase shift,and beamforming.Herein,a new jitter-oriented system stability objective with respect to control error and communication jitter is defined and the closed-form expression of control delay deadline is derived based on the Jensen Inequality and Lyapunov-Krasovskii functional.Due to the time-varying and partial observability of the channel and robot states,we model the problem as a partially observable Markov decision process(POMDP).To solve this complex problem,we propose a multi-agent transfer reinforcement learning algorithm named LSTM-PPO-MATRL,where the LSTM-enhanced proximal policy optimization(PPO)is designed to approximate an optimal solution and the option-guided policy transfer learning is proposed to facilitate the learning process.By centralized training and decentralized execution,LSTM-PPO-MATRL is validated by extensive experiments on MuJoCo tasks for both low-mobility and high-mobility robotic control scenarios.The results demonstrate that LSTM-PPO-MATRL not only realizes high learning efficiency,but also supports low-delay,low-jitter communication for low error control,where 71.9%control accuracy improvement and 68.7%delay jitter reduction are achieved compared to the PPO-MADRL baseline. 展开更多
关键词 Multi-agent transfer reinforcement learning(MATRL) partially observable Markov decision process(POMDP) reconfigurable intelligent surface(RIS) system stability wireless cloud robotic system(WCRS)
暂未订购 下载PDF
A brief review of evolutionary game dynamics in the reinforcement learning paradigm 认领 引用
4
作者 Guozhong Zheng Xin Ou +2 位作者 Shengfeng Deng Jiqiang Zhang Li Chen 《Communications in Theoretical Physics》 SCIE CAS CSCD 2026年第6期220-232,共13页
Cooperation,fairness,trust,and resource coordination are cornerstones of modern civilization,yet their emergence remains inadequately explained,largely due to persistent discrepancies between theoretical predictions a... Cooperation,fairness,trust,and resource coordination are cornerstones of modern civilization,yet their emergence remains inadequately explained,largely due to persistent discrepancies between theoretical predictions and behavioral experiments.Part of this gap may arise from the imitation learning paradigm commonly used in prior theoretical models,which assumes individuals merely copy successful neighbors according to predetermined,fixed rules.This review examines recent advances in evolutionary game dynamics that employ reinforcement learning(RL)as an alternative paradigm.In RL,individuals learn through trial and error and intro spec tively refine their strategies based on environmental feedback.We begin by introducing key concepts in evolutionary game theory and the two learning paradigms,then synthesize progress in applying RL to elucidate cooperation,trust,fairness,optimal resource coordination,and ecological dynamics.Collectively,these studies indicate that RL offers a promising unified framework for understanding the diverse social and ecological phenomena observed in human and natural systems. 展开更多
关键词 reinforcement learning evolutionary game theory cooperation fairness trust resource allocation biodiversity
暂未订购 下载PDF
Deep Reinforcement Learning Based on Search Space Independent Operators for Black-Box Continuous Optimization 认领 引用
5
作者 Ye Tian Yisai Liu +1 位作者 Shangshang Yang Xingyi Zhang 《IEEE/CAA Journal of Automatica Sinica》 SCIE EI CSCD 2026年第4期913-925,共13页
Deep reinforcement learning(DRL)has demonstrated exceptional capabilities in combinatorial optimization,which automatically devises policies for solution construction and optimizer refinement.DRL is particularly adept... Deep reinforcement learning(DRL)has demonstrated exceptional capabilities in combinatorial optimization,which automatically devises policies for solution construction and optimizer refinement.DRL is particularly adept in generating training samples by itself,thereby providing the flexibility to solve a variety of combinatorial optimization problems without supervision.While DRL takes actions according to states extracted from problem-specific information,it cannot be directly applied to black-box continuous optimization lacking explicit information.To address this issue,this paper proposes a search space independent operator based DRL method for black-box continuous optimization.It conceptualizes the optimization process driven by search space independent operators as a Markov decision process,wherein actions are defined as operators and states are extracted from solutions generated by operators.In contrast to other DRLassisted metaheuristics,the proposed method does not rely on any existing metaheuristic.Instead,it innovates by creating totally new operators,able to surpass the performance boundaries of existing metaheuristics.Compared with state-of-the-art metaheuristics and DRL methods,the proposed method shows significantly faster convergence speed on challenging continuous optimization problems. 展开更多
关键词 Black-box optimization continuous optimization metaheuristic reinforcement learning search operator
暂未订购 下载PDF
Temporal Dependency-Aware Trajectory-Level Behavioural Metric for Exploration in Reinforcement Learning 认领 引用
6
作者 Anjie Zhu Yongjun Yang +1 位作者 Guangyi Zhao Jie Shao 《CAAI Transactions on Intelligence Technology》 SCIE EI CSCD 2026年第2期332-348,共17页
Intrinsic motivation serves as the predominant paradigm of exploration in reinforcement learning.In pursuit of an informative and robust state representation,the behavioural metric groups behaviourally equivalent stat... Intrinsic motivation serves as the predominant paradigm of exploration in reinforcement learning.In pursuit of an informative and robust state representation,the behavioural metric groups behaviourally equivalent states together,which share the same single-step reward and transition distribution.However,due to the presence of uninformative rewards and the dynamic nature of procedurally generated environments,these behavioural metric-based approaches could limit the effectiveness of the learnt state representations,potentially leading to a representation collapse and an ineffective exploration.Therefore,a more comprehensive and generalisable behavioural metric is needed to overcome the above issues.In this work,we approach the exploration problem from a novel perspective,extending beyond the conventional single-step assessments to encompass a longterm consideration of the whole trajectory.Specifically,we propose a novel trajectory-level behavioural metric(TBM)that exploits temporal dependencies of the trajectory and captures the underlying sequential information of behaviour patterns.To achieve an effective trajectory representation for exploration,we develop a pivotal state identifier(PSI)and a trajectory return estimator(TRE)to distinguish the diverse contributions of individual states in the trajectory.Moreover,an auxiliary representation regulariser is developed to promote the diversity and informativeness of the trajectory representation,mitigating the risk of representation mode collapse.Extensive experiments and empirical analysis conducted on procedurally generated environments showcase the superior performance of our proposed framework. 展开更多
关键词 learning(artificial intelligence) machine learning reinforcement learning
暂未订购 下载PDF
Enhanced exploration for multi-UAV cooperative roundup:An I2C-MATD3 reinforcement learning framework 认领 引用
7
作者 Bo Li Jingyi Huang +2 位作者 Haohui Zhang Liangliang Huai Evgeny Neretin 《Defence Technology(防务技术)》 SCIE EI CAS CSCD 2026年第4期374-389,共16页
With the increasing maturity of multi-UAV technology and its broad applications in scenarios such as UAV roundup tasks,this paper proposes a novel approach to enhance interception efficiency and system robustness by a... With the increasing maturity of multi-UAV technology and its broad applications in scenarios such as UAV roundup tasks,this paper proposes a novel approach to enhance interception efficiency and system robustness by addressing insufficient historical data utilization and inadequate environmental explo-ration.The multi-UAV roundup problem is formulated as a Markov Decision Process(MDP),and an Improved Cross-Entropy Method with Intrinsic Curiosity-enhanced Multi-Agent Twin Delayed Deep Deterministic Policy Gradient(I2C-MATD3)is designed.Specifically,an Improved Cross-Entropy Method(ICEM)based on global elite samples rapidly optimizes training strategies while generating extensive experience for a Multi-Agent Twin Delayed Deep Deterministic Policy Gradient algorithm augmented with intrinsic curiosity rewards(IC-MATD3).In turn,IC-MATD3 guides the optimization direction of ICEM,enabling a synergistic interaction that facilitates effective historical data exploitation and pro-active environmental exploration for UAV agents to accomplish roundup tasks.Experiments in complex scenarios demonstrate that the proposed algorithm achieves superior training efficiency and conver-gence performance compared to state-of-the-art multi-agent reinforcement learning(MARL)methods.Robustness tests and ablation experiments further validate its enhanced generalizability and robustness. 展开更多
关键词 Multi-UAV roundup Intrinsic curiosity module Cross-entropy method Multi-agent reinforcement learning
暂未订购 下载PDF
A review of reinforcement learning approaches for pursuit-evasion games 认领 引用
8
作者 Kun YANG Ao SHEN +3 位作者 Nengwei XU Fang DENG Maobin LU Chen CHEN 《Chinese Journal of Aeronautics》 SCIE EI CAS CSCD 2026年第6期423-445,共23页
As a special type of dynamic game,Pursuit-Evasion Games(PEGs)have expanded their application range from initial military confrontations to areas such as navigation control and aerospace,demonstrating broad applicabili... As a special type of dynamic game,Pursuit-Evasion Games(PEGs)have expanded their application range from initial military confrontations to areas such as navigation control and aerospace,demonstrating broad applicability and significant value in addressing a wide array of modern complex decision-making problems.Traditional optimal control methods based on differential game theory are classic approaches to solve PEG problems.However,these methods often struggle to perform well in complex environments,nonlinear systems,and situations involving highly uncertain participant behaviors.In recent years,rapidly developing Reinforcement Learning(RL)techniques has provided new avenues for PEG research.RL is capable of adapting to environmental changes through efficient online computation and feedback-driven learning,exhibiting strong generalization capabilities.Therefore,this survey presents a detailed and systematic review of PEG research based on RL methods.First,it classifies and discusses key RL algorithms and theoretical foundations in PEGs according to different forms of strategy learning.Then,it summarizes typical application scenarios,including tactical combat,unmanned systems control,and spacecraft interception,demonstrating the potential and effectiveness of RL in addressing real-world challenges.Finally,the survey explores current challenges and future opportunities in applying RL to PEGs,with the aim of promoting further research on more effective and practical solutions. 展开更多
关键词 Pursuit-evasion games Reinforcement learning Spacecraft interception Tactical combat Unmanned systems
暂未订购 下载PDF
Efficient collaborative planning for carrier-borne aircraft dispatch and recovery via hierarchical reinforcement learning 认领 引用
9
作者 Shaohui ZHANG Qiuying HAN +3 位作者 Qingshun WU Di ZHANG Yafei LI Mingliang XU 《Chinese Journal of Aeronautics》 SCIE EI CAS CSCD 2026年第4期528-540,共13页
Carrier-borne aircraft are the primary formidable assets in aircraft carrier combat,and their sortie rate is a pivotal metric for evaluating the carrier's combat capability.Enhancing the efficiency of aircraft sup... Carrier-borne aircraft are the primary formidable assets in aircraft carrier combat,and their sortie rate is a pivotal metric for evaluating the carrier's combat capability.Enhancing the efficiency of aircraft support operations scheduling is a significant means to improve the sortie rate,where the central issue is to assign multi-wave aircraft to support stations according to the flight plan,and then obtain support resources for completing support operations.The existing studies primarily focus on considering partial operation processes(e.g.,ammunition transfer,disturbance handling,support personnel deployment,and deck arrangement),and lack modeling of the entire process of multi-wave aircraft support.In this paper,we investigate the multi-wave Aircraft Dispatch and Recovery Planning(ADRP)problem that aims to reasonably plan the operation processes,stations and resources of multi-wave aircraft with the goal of minimizing the total support operation time,and present a three-layer solution framework based on hierarchical reinforcement learning to address it.Specifically,we first abstract the multi-wave aircraft support operation planning process into the process layer,station layer and resource layer,and model it as a Decentralized Partially Observable Markov Decision Process(Dec-POMDP).Then,we propose a three-layer solution framework based on hierarchical reinforcement learning to solve the ADRP problem.To further improve planning results from long-term and global perspective,we design an inter-layer communication mechanism to allow efficient information exchange between stations.Extensive experimental results demonstrate that our proposed approach can achieve a high-quality operating schedule while meeting real-time demands. 展开更多
关键词 Carrier-borne aircraft Multi-wave support operations Multi-agent systems Hierarchical reinforcement learning Markov processes
暂未订购 下载PDF
Implementation of Human-AI Interaction in Reinforcement Learning: Literature Review and Case Studies 认领 引用
10
作者 Shaoping Xiao Zhaoan Wang +3 位作者 Junchao Li Caden Noeller Jiefeng Jiang Jun Wang 《Computers, Materials & Continua》 SCIE EI 2026年第2期1-62,共62页
Theintegration of human factors into artificial intelligence(AI)systems has emerged as a critical research frontier,particularly in reinforcement learning(RL),where human-AI interaction(HAII)presents both opportunitie... Theintegration of human factors into artificial intelligence(AI)systems has emerged as a critical research frontier,particularly in reinforcement learning(RL),where human-AI interaction(HAII)presents both opportunities and challenges.As RL continues to demonstrate remarkable success in model-free and partially observable environments,its real-world deployment increasingly requires effective collaboration with human operators and stakeholders.This article systematically examines HAII techniques in RL through both theoretical analysis and practical case studies.We establish a conceptual framework built upon three fundamental pillars of effective human-AI collaboration:computational trust modeling,system usability,and decision understandability.Our comprehensive review organizes HAII methods into five key categories:(1)learning from human feedback,including various shaping approaches;(2)learning from human demonstration through inverse RL and imitation learning;(3)shared autonomy architectures for dynamic control allocation;(4)human-in-the-loop querying strategies for active learning;and(5)explainable RL techniques for interpretable policy generation.Recent state-of-the-art works are critically reviewed,with particular emphasis on advances incorporating large language models in human-AI interaction research.To illustrate some concepts,we present three detailed case studies:an empirical trust model for farmers adopting AI-driven agricultural management systems,the implementation of ethical constraints in roboticmotion planning through human-guided RL,and an experimental investigation of human trust dynamics using a multi-armed bandit paradigm.These applications demonstrate how HAII principles can enhance RL systems’practical utility while bridging the gap between theoretical RL and real-world human-centered applications,ultimately contributing to more deployable and socially beneficial intelligent systems. 展开更多
关键词 Human-AI interaction reinforcement learning partially observable environments trust model ethical constraints
暂未订购 下载PDF
Towards Generalisable and Explainable Traffic Signal Control via Deep Reinforcement Learning and Large Language Models 认领 引用
11
作者 Hao Huang Wenjie He +6 位作者 Qilie Liu Qian Liu Chao Huang Anwar PPAbdul Majeed Xiangguang Dai Gang Fang Xiaohua Xu 《CAAI Transactions on Intelligence Technology》 SCIE EI CSCD 2026年第2期483-497,共15页
As a government-regulated public service,traffic signal control(TSC)requires reliable and transparent decision-making.However,existing deep reinforcement learning(DRL)methods,despite improvements in control accuracy,s... As a government-regulated public service,traffic signal control(TSC)requires reliable and transparent decision-making.However,existing deep reinforcement learning(DRL)methods,despite improvements in control accuracy,still lack explainability and generalisation,severely limiting their applicability in real-world environments.To address the challenges above,this paper proposes GenEx-TSC,a generalisable and explainable TSC method that integrates deep reinforcement learning with large language models(LLMs).First,starting from vehicle-level states,we train a DRL agent incorporating intersection physical heterogeneity and neighbourhood information,which lays the evaluation foundation for constructing a high-quality LLM dataset.Subsequently,the LLM agent is optimised through a two-stage training mechanism.In the distillation stage,a lightweight LLM agent is trained using the reasoning trajectories of a larger-scale LLM agent,inheriting its semantic understanding and decision-generation capabilities and in the alignment stage,the DRL evaluation network is employed to calibrate the outputs of the distilled LLM agent,ensuring that the generated cycle-level signal timing strategies are both efficient and interpretable.We synthesise 10 intersection networks with different physical attributes in SUMO and set traffic flows of varying scales.Experimental results across diverse traffic environments demonstrate that the proposed GenEx-TSC exhibits clear advantages over traditional methods,mainstream DRL methods and LLM baselines in terms of control accuracy,generalisation and explainability. 展开更多
关键词 deep reinforcement learning explainable intelligence large language model traffic signal control
暂未订购 下载PDF
Photonic spiking reinforcement learning for intelligent routing 认领 引用
12
作者 Shuiying Xiang Yonghang Chen +8 位作者 Ling Zheng Zhicong Tu Xintao Zeng Mengting Yu Shuai Wang Yahui Zhang Xingxing Guo Weitao Pan Yue Hao 《Opto-Electronic Science》 CAS 2026年第5期11-26,共16页
Intelligent routing plays a key role in modern communication infrastructure,including data centers,computing networks,and future 6G networks.Although reinforcement learning(RL)has shown great potential for intelligent... Intelligent routing plays a key role in modern communication infrastructure,including data centers,computing networks,and future 6G networks.Although reinforcement learning(RL)has shown great potential for intelligent routing,its practical deployment remains constrained by high energy consumption and decision latency.Here,we propose a photonic spiking RL architecture that implements a proximal policy optimization(PPO)–based intelligent routing algorithm.The performance of the proposed approach is systematically evaluated on a softwaredefined network(SDN)with a fat-tree topology.The results demonstrate that,under various baseline traffic rate conditions,the PPO-based routing strategy significantly outperforms the conventional Dijkstra algorithm in key performance metrics,including throughput,packet loss rate,average latency,and load balance.Furthermore,a hardware-software collaborative framework of the spiking Actor network is realized for three typical baseline traffic rates,utilizing a photonic synapse chip based on a Mach-Zehnder interferometer(MZI)array and a photonic spiking neuron chip based on distributed feedback lasers with a saturable absorber(DFB-SAs).Experimental validation on 640 state–action pairs shows that the inference accuracy of the hardware-software collaborative framework is consistent with that of the pure algorithmic implementation.The impacts of different hidden-layer scales in the spiking Actor network and varying network size of fat-tree topology are further analyzed.The integration of photonic spiking RL with SDN-based routing establishes a novel paradigm for intelligent routing optimization,featuring ultralow latency and high energy efficiency.This approach exhibits broad application prospects in real-time network optimization scenarios,including large-scale data centers,computing networks,satellite Internet systems,and future 6G networks. 展开更多
关键词 photonic spiking neural network spiking reinforcement learning intelligent routing SDN
暂未订购 下载PDF
Research on UAV-MEC Cooperative Scheduling Algorithms Based on Multi-Agent Deep Reinforcement Learning 认领 引用
13
作者 Yonghua Huo Ying Liu +1 位作者 Anni Jiang Yang Yang 《Computers, Materials & Continua》 SCIE EI 2026年第3期1823-1850,共28页
With the advent of sixth-generation mobile communications(6G),space-air-ground integrated networks have become mainstream.This paper focuses on collaborative scheduling for mobile edge computing(MEC)under a three-tier... With the advent of sixth-generation mobile communications(6G),space-air-ground integrated networks have become mainstream.This paper focuses on collaborative scheduling for mobile edge computing(MEC)under a three-tier heterogeneous architecture composed of mobile devices,unmanned aerial vehicles(UAVs),and macro base stations(BSs).This scenario typically faces fast channel fading,dynamic computational loads,and energy constraints,whereas classical queuing-theoretic or convex-optimization approaches struggle to yield robust solutions in highly dynamic settings.To address this issue,we formulate a multi-agent Markov decision process(MDP)for an air-ground-fused MEC system,unify link selection,bandwidth/power allocation,and task offloading into a continuous action space and propose a joint scheduling strategy that is based on an improved MATD3 algorithm.The improvements include Alternating Layer Normalization(ALN)in the actor to suppress gradient variance,Residual Orthogonalization(RO)in the critic to reduce the correlation between the twin Q-value estimates,and a dynamic-temperature reward to enable adaptive trade-offs during training.On a multi-user,dual-link simulation platform,we conduct ablation and baseline comparisons.The results reveal that the proposed method has better convergence and stability.Compared with MADDPG,TD3,and DSAC,our algorithm achieves more robust performance across key metrics. 展开更多
关键词 UAV-MEC networks multi-agent deep reinforcement learning MATD3 task offloading
暂未订购 下载PDF
Time control entry guidance method for hypersonic glide vehicles based on deep reinforcement learning 认领 引用
14
作者 Zhenyu LIU Gang LEI +3 位作者 Yong XIAN Leliang REN Shaopeng LI Daqiao ZHANG 《Journal of Zhejiang University-SCIENCE A》 SCIE EI CAS CSCD 2026年第4期365-383,I0016-I0025,I0051,共19页
To meet the requirement of simultaneous arrival for multiple hypersonic glide vehicles(HGVs),we propose a time control entry guidance(TCEG)method leveraging deep reinforcement learning.First,the entry guidance problem... To meet the requirement of simultaneous arrival for multiple hypersonic glide vehicles(HGVs),we propose a time control entry guidance(TCEG)method leveraging deep reinforcement learning.First,the entry guidance problem is solved with a reinforcement learning framework based on a designed reference flight profile.By appropriately designing the observation space and training environment,the well-trained agent demonstrates robust guidance performance under varying widths of the heading error corridor.Then,a novel method for predicting the remaining flight time is established,which consists of two main components.The first component estimates the remaining flight time using an analytical formula,while the second component employs a deep neural network(DNN)to predict the residual error between the estimated and the true value.Subsequently,based on the predicted terminal time error,the threshold of the heading error and the observation vector are corrected in real time,thereby guiding the agent to dynamically adjust its output actions.This enables precise control of the terminal time.Since the generation of guidance commands only requires forward computations by the neural network,the proposed method exhibits excellent real-time performance.Finally,the effectiveness and robustness of the method are demonstrated through numerical simulations in various scenarios. 展开更多
关键词 Hypersonic glide vehicles(HGVs) Entry guidance Reinforcement learning Time coordination Deep neural network(DNN)
暂未订购 下载PDF
Multi-Agent Reinforcement Learning Driven Dynamic Resource Optimisation in Healthcare Transportation Networks 认领 引用
15
作者 Jianhui Lv Byung-Gyu Kim +1 位作者 Keqin Li Heng Lu 《CAAI Transactions on Intelligence Technology》 SCIE EI CSCD 2026年第2期316-331,共16页
This paper presents HealthNet,a novel framework for the dynamic optimisation of healthcare transportation networks using multi-agent reinforcement learning.HealthNet leverages a spatiotemporal dependency module to cap... This paper presents HealthNet,a novel framework for the dynamic optimisation of healthcare transportation networks using multi-agent reinforcement learning.HealthNet leverages a spatiotemporal dependency module to capture complex spatiotemporal relationships in healthcare demand and resource allocation patterns,combined with centralised training and a decentralised execution approach.The system is modelled as a Markov game and solved using a deep reinforcement learning algorithm.Extensive simulations demonstrate that HealthNet outperforms eight state-of-the-art baseline methods across multiple network configurations and evaluation metrics.In a 4×4 grid network,HealthNet reduces average waiting times by 47.6%compared to model predictive control and 22.1%compared to the best-performing baseline.Traffic congestion rates are reduced to 16.7%compared to 42.3%for the worst baseline and 23.1%for the best baseline.Under irregular network topologies with stochastic disruptions,including demand surges and vehicle unavailability,HealthNet maintains superior performance with 42.1%lower average waiting time and 51.1%improvement in peak response times compared to competing approaches.These findings indicate that HealthNet can enhance both efficiency and resilience in healthcare transportation systems,potentially improving patient outcomes in complex urban environments. 展开更多
关键词 dynamic resource optimisation healthcare transportation networks reinforcement learning spatiotemporal dependency sustainable cities
暂未订购 下载PDF
Evaluation of Reinforcement Learning-Based Adaptive Modulation in Shallow Sea Acoustic Communication 认领 引用
16
作者 Yifan Qiu Xiaoyu Yang +1 位作者 Feng Tong Dongsheng Chen 《哈尔滨工程大学学报(英文版)》 CSCD 2026年第1期292-299,共8页
While reinforcement learning-based underwater acoustic adaptive modulation shows promise for enabling environment-adaptive communication as supported by extensive simulation-based research,its practical performance re... While reinforcement learning-based underwater acoustic adaptive modulation shows promise for enabling environment-adaptive communication as supported by extensive simulation-based research,its practical performance remains underexplored in field investigations.To evaluate the practical applicability of this emerging technique in adverse shallow sea channels,a field experiment was conducted using three communication modes:orthogonal frequency division multiplexing(OFDM),M-ary frequency-shift keying(MFSK),and direct sequence spread spectrum(DSSS)for reinforcement learning-driven adaptive modulation.Specifically,a Q-learning method is used to select the optimal modulation mode according to the channel quality quantified by signal-to-noise ratio,multipath spread length,and Doppler frequency offset.Experimental results demonstrate that the reinforcement learning-based adaptive modulation scheme outperformed fixed threshold detection in terms of total throughput and average bit error rate,surpassing conventional adaptive modulation strategies. 展开更多
关键词 Adaptive modulation Shallow sea underwater acoustic modulation Reinforcement learning
暂未订购 下载PDF
Security Control of Nonlinear Systems Subject to Deception Attacks:A Reinforcement Learning Approach 认领 引用
17
作者 Lifeng Ma Yongyi Dai Chen Gao 《IEEE/CAA Journal of Automatica Sinica》 SCIE EI CSCD 2026年第7期1764-1766,共3页
Dear Editor,This letter deals with the security control for nonlinear cyber-physical systems(CPSs)under mixed deception attacks.Both sensors and actuators are assumed to be injected deception data during the data tran... Dear Editor,This letter deals with the security control for nonlinear cyber-physical systems(CPSs)under mixed deception attacks.Both sensors and actuators are assumed to be injected deception data during the data transmission via networks.In order to identify the unknown dynamics of the attacked system,a neural network(NN)is adopted,on basis of which an NN-based secure observer is designed to diminish the attack impact on state estimation.Then,by resorting to the reinforcement learning approach,the secure control strategy is presented via actor-critic and zero-sum games.At last,the designed control scheme is proved via a numerical simulation. 展开更多
关键词 nonlinear systems neural network nn reinforcement learning approachthe identify unknown dynamics security control deception attacks mixed deception attacksboth reinforcement learning
暂未订购 下载PDF
A Multi-Objective Deep Reinforcement Learning Algorithm for Computation Offloading in Internet of Vehicles 认领 引用
18
作者 Junjun Ren Guoqiang Chen +1 位作者 Zheng-Yi Chai Dong Yuan 《Computers, Materials & Continua》 SCIE EI 2026年第1期2111-2136,共26页
Vehicle Edge Computing(VEC)and Cloud Computing(CC)significantly enhance the processing efficiency of delay-sensitive and computation-intensive applications by offloading compute-intensive tasks from resource-constrain... Vehicle Edge Computing(VEC)and Cloud Computing(CC)significantly enhance the processing efficiency of delay-sensitive and computation-intensive applications by offloading compute-intensive tasks from resource-constrained onboard devices to nearby Roadside Unit(RSU),thereby achieving lower delay and energy consumption.However,due to the limited storage capacity and energy budget of RSUs,it is challenging to meet the demands of the highly dynamic Internet of Vehicles(IoV)environment.Therefore,determining reasonable service caching and computation offloading strategies is crucial.To address this,this paper proposes a joint service caching scheme for cloud-edge collaborative IoV computation offloading.By modeling the dynamic optimization problem using Markov Decision Processes(MDP),the scheme jointly optimizes task delay,energy consumption,load balancing,and privacy entropy to achieve better quality of service.Additionally,a dynamic adaptive multi-objective deep reinforcement learning algorithm is proposed.Each Double Deep Q-Network(DDQN)agent obtains rewards for different objectives based on distinct reward functions and dynamically updates the objective weights by learning the value changes between objectives using Radial Basis Function Networks(RBFN),thereby efficiently approximating the Pareto-optimal decisions for multiple objectives.Extensive experiments demonstrate that the proposed algorithm can better coordinate the three-tier computing resources of cloud,edge,and vehicles.Compared to existing algorithms,the proposed method reduces task delay and energy consumption by 10.64%and 5.1%,respectively. 展开更多
关键词 Deep reinforcement learning internet of vehicles multi-objective optimization cloud-edge computing computation offloading service caching
暂未订购 下载PDF
A Deep Reinforcement Learning-Based Partitioning Method for Power System Parallel Restoration 认领 引用
19
作者 Changcheng Li Weimeng Chang +1 位作者 Dahai Zhang Jinghan He 《Energy Engineering》 EI 2026年第1期243-264,共22页
Effective partitioning is crucial for enabling parallel restoration of power systems after blackouts.This paper proposes a novel partitioning method based on deep reinforcement learning.First,the partitioning decision... Effective partitioning is crucial for enabling parallel restoration of power systems after blackouts.This paper proposes a novel partitioning method based on deep reinforcement learning.First,the partitioning decision process is formulated as a Markov decision process(MDP)model to maximize the modularity.Corresponding key partitioning constraints on parallel restoration are considered.Second,based on the partitioning objective and constraints,the reward function of the partitioning MDP model is set by adopting a relative deviation normalization scheme to reduce mutual interference between the reward and penalty in the reward function.The soft bonus scaling mechanism is introduced to mitigate overestimation caused by abrupt jumps in the reward.Then,the deep Q network method is applied to solve the partitioning MDP model and generate partitioning schemes.Two experience replay buffers are employed to speed up the training process of the method.Finally,case studies on the IEEE 39-bus test system demonstrate that the proposed method can generate a high-modularity partitioning result that meets all key partitioning constraints,thereby improving the parallelism and reliability of the restoration process.Moreover,simulation results demonstrate that an appropriate discount factor is crucial for ensuring both the convergence speed and the stability of the partitioning training. 展开更多
关键词 Partitioning method parallel restoration deep reinforcement learning experience replay buffer partitioning modularity
暂未订购 下载PDF
上一页 1 2 250 下一页 到第
在线咨询 使用帮助 返回顶部 意见反馈