With the rapid development of power Internet of Things(IoT)scenarios such as smart factories and smart homes,numerous intelligent terminal devices and real-time interactive applications impose higher demands on comput...With the rapid development of power Internet of Things(IoT)scenarios such as smart factories and smart homes,numerous intelligent terminal devices and real-time interactive applications impose higher demands on computing latency and resource supply efficiency.Multi-access edge computing technology deploys cloud computing capabilities at the network edge;constructs distributed computing nodes and multi-access systems and offers infrastructure support for services with low latency and high reliability.Existing research relies on a strong assumption that the environmental state is fully observable and fails to thoroughly consider the continuous time-varying features of edge server load fluctuations,leading to insufficient adaptability of the model in a heterogeneous dynamic environment.Thus,this paper establishes a framework for end-edge collaborative task offloading based on a partially observable Markov decision-making process(POMDP)and proposes a method for end-edge collaborative task offloading in heterogeneous scenarios.It achieves time-series modeling of the historical load characteristics of edge servers and endows the agent with the ability to be aware of the load in dynamic environmental states.Moreover,by dynamically assessing the exploration value of historical trajectories in the central trajectory pool and adjusting the sample weight distribution,directional exploration and strategy optimization of high-value trajectories are realized.Experimental results indicate that the proposed method exhibits distinct advantages compared with existing methods in terms of average delay and task failure rate and also verifies the method’s robustness in a dynamic environment.展开更多
With the dramatic increase in the demand for computing power across various services,the emergence of the computing power network(CPN)becomes inevitable.This paper studies the task scheduling to minimize energy consum...With the dramatic increase in the demand for computing power across various services,the emergence of the computing power network(CPN)becomes inevitable.This paper studies the task scheduling to minimize energy consumption under delay constraints considering the heterogeneity of computing resources.Specifically,we decompose the original problem and alternatively optimize the scheduling strategy and server parameters until convergence.Dynamic Voltage and frequency scaling(DVFS)technology is leveraged to allocate the optimal voltage and frequency for each server based on their task loads.An enhanced projection gradient descent method is utilized to update the scheduling strategy under the given server parameters.Simulation results show that our algorithm achieves significant performance gains compared to the baselines across various CPN scenarios.展开更多
Ambient noise tomography is an established technique in seismology,where calculating single-or ninecomponent noise cross-correlation functions(NCFs)is a fundamental first step.In this study,we introduced a novel CPU-G...Ambient noise tomography is an established technique in seismology,where calculating single-or ninecomponent noise cross-correlation functions(NCFs)is a fundamental first step.In this study,we introduced a novel CPU-GPU heterogeneous computing framework designed to significantly enhance the efficiency of computing 9-component NCFs from seismic ambient noise data.This framework not only accelerated the computational process by leveraging the Compute Unified Device Architecture(CUDA)but also improved the signal-to-noise ratio(SNR)through innovative stacking techniques,such as time-frequency domain phaseweighted stacking(tf-PWS).We validated the program using multiple datasets,confirming its superior computation speed,improved reliability,and higher signal-to-noise ratios for NCFs.Our comprehensive study provides detailed insights into optimizing the computational processes for noise cross-correlation functions,thereby enhancing the precision and efficiency of ambient noise imaging.展开更多
Federated learning is an emerging machine learning techniquethat enables clients to collaboratively train a deep learning model withoutuploading raw data to the aggregation server. Each client may be equippedwith diff...Federated learning is an emerging machine learning techniquethat enables clients to collaboratively train a deep learning model withoutuploading raw data to the aggregation server. Each client may be equippedwith different computing resources for model training. The client equippedwith a lower computing capability requires more time for model training,resulting in a prolonged training time in federated learning. Moreover, it mayfail to train the entire model because of the out-of-memory issue. This studyaims to tackle these problems and propose the federated feature concatenate(FedFC) method for federated learning considering heterogeneous clients.FedFC leverages the model splitting and feature concatenate for offloadinga portion of the training loads from clients to the aggregation server. Eachclient in FedFC can collaboratively train a model with different cutting layers.Therefore, the specific features learned in the deeper layer of the serversidemodel are more identical for the data class classification. Accordingly,FedFC can reduce the computation loading for the resource-constrainedclient and accelerate the convergence time. The performance effectiveness isverified by considering different dataset scenarios, such as data and classimbalance for the participant clients in the experiments. The performanceimpacts of different cutting layers are evaluated during the model training.The experimental results show that the co-adapted features have a criticalimpact on the adequate classification of the deep learning model. Overall,FedFC not only shortens the convergence time, but also improves the bestaccuracy by up to 5.9% and 14.5% when compared to conventional federatedlearning and splitfed, respectively. In conclusion, the proposed approach isfeasible and effective for heterogeneous clients in federated learning.展开更多
Heterogeneous computing (HC) environment utilizes diverse resources with different computational capabilities to solve computing-intensive applications having diverse computational requirements and constraints. The ta...Heterogeneous computing (HC) environment utilizes diverse resources with different computational capabilities to solve computing-intensive applications having diverse computational requirements and constraints. The task assignment problem in HC environment can be formally defined as for a given set of tasks and machines, assigning tasks to machines to achieve the minimum makespan. In this paper we propose a new task scheduling heuristic, high standard deviation first (HSTDF), which considers the standard deviation of the expected execution time of a task as a selection criterion. Standard deviation of the ex- pected execution time of a task represents the amount of variation in task execution time on different machines. Our conclusion is that tasks having high standard deviation must be assigned first for scheduling. A large number of experiments were carried out to check the effectiveness of the proposed heuristic in different scenarios, and the comparison with the existing heuristics (Max-min, Sufferage, Segmented Min-average, Segmented Min-min, and Segmented Max-min) clearly reveals that the proposed heuristic outperforms all existing heuristics in terms of average makespan.展开更多
Low-density parity check(LDPC)decoding is an efficient error correction method in communication systems,especially in 5G networks,which require high performance and low latency;while common general-purpose architectur...Low-density parity check(LDPC)decoding is an efficient error correction method in communication systems,especially in 5G networks,which require high performance and low latency;while common general-purpose architectures cannot meet the requirements.There has been some research on accelerating LDPC decoding,but the current methods still suffer from limitations in performance,flexibility,and communication cost.In this paper,we propose HARLD(Heterogeneous Architecture of RISC-V for LDPC Decoding),a tightly coupled heterogeneous computing architecture based on extended RISC-V for LDPC decoding,consisting of a CPU and a processing array.Compared with a loosely coupled System-on-Chip(SoC)-bus baseline,the tightly coupled design improves throughput by up to 32.4%and reduces average latency by up to 24.7%across evaluated configurations,while also enhancing resource and energy efficiency:processing element utilization up to 93.5%,instruction RAM utilization increased by up to 4.8x,and energy efficiency improved by up to 24.8%.At the system level,area and power are reduced by 17.6%and 10.2%,respectively,versus the loosely coupled design.展开更多
To reduce the running time of network simulation in heterogeneous computing environment,a network simulation task partition method,named LBPHCE,is put forward.In this method,the network simulation task is partitioned ...To reduce the running time of network simulation in heterogeneous computing environment,a network simulation task partition method,named LBPHCE,is put forward.In this method,the network simulation task is partitioned in comprehensive consideration of the load balance of both routing computing simulation and packet forwarding simulation.First,through benchmark experiments,the computation ability and routing simulation ability of each simulation machine are measured in the heterogeneous computing environment.Second,based on the computation ability of each simulation machine,the network simulation task is initially partitioned to meet the load balance of packet forwarding simulation in the heterogeneous computing environment,and then according to the routing computation ability,the scale of each partition is fine-tuned to satisfy the balance of the routing computing simulation,meanwhile the load balance of packet forwarding simulation is guaranteed.Experiments based on PDNS indicate that,compared to traditional uniform partition method,the LBPHCE method can reduce the total simulation running time by 26.3%in average,and compared to the liner partition method,it can reduce the running time by 18.3%in average.展开更多
Non-uniform sampling two-dimensional convolution (NUSC) maps spatially sampling data with irregular distribution to a regular grid by convolution. As the data scale and growth rate continue to increase, accelerating N...Non-uniform sampling two-dimensional convolution (NUSC) maps spatially sampling data with irregular distribution to a regular grid by convolution. As the data scale and growth rate continue to increase, accelerating NUSC with the heterogene-ous computing platform is a feasible way. However, the complex hardware architecture and storage hierarchy of the hetero-geneous computing platform poses a challenge to programming and performance tuning. Therefore, this paper proposes a heterogeneous parallel programming model and runtime framework named AutoNUSC. For the programming difficulties of NUSC in heterogeneous computing environments, AutoNUSC abstracts and encapsulates the parallel execution process of NUSC. Task scheduling, data division, node communication, fault-tolerant recovery, and other parallelization tasks are managed by AutoNUSC. For the performance tuning issues of NUSC, this paper implements performance optimization strategies for AutoNUSC, including vectorization, memory access optimization, data reuse, etc. The experiments show that AutoNUSC effectively reduces the workload of users in developing NUSC applications in heterogeneous computing environments. Performance acceleration of up to 339 times is achieved within a single node compared to the serial program. AutoNUSC can efficiently perform task scheduling and fault-tolerant recovery across multiple nodes, with desirable scal-ability and robustness.展开更多
Applications involving multifarious computational requirements take the advantage of the versatility of heterogeneous computing systems(HCS)with more than one type of parallelism.Efficient scheduling of workflow appli...Applications involving multifarious computational requirements take the advantage of the versatility of heterogeneous computing systems(HCS)with more than one type of parallelism.Efficient scheduling of workflow applications is paramount to harness high performance from HCS.In the present work,a new list-based heuristic strategy namely maximizing parallelism for minimizing earliest finish time(MPEFT)algorithm is proposed with a primary objective of minimizing the makespan.In order to minimize the makespan,the proposed scheduling policy focuses on proliferating the parallelism of the workflows by choosing the globally heaviest task with more number of successors such that more number of successors can be released.Thus,the priority policy maximizes the length of the ready queue by exploring higher degree of parallelism of the workflow.The proposed approach is designed to adapt depth-wise whenever the tasks at subsequent levels are released and continues to be level-wise otherwise.This increases the degree of parallelism and shortens the makespan.To evaluate the proposed scheduling algorithm,experimentations are conducted using randomly generated workflows and scientific workflows namely LIGO,Epigenomics,Cybershake,and Montage.The experimental results show that the proposed MPEFT algorithm surpassed the classical list based heuristic algorithms in terms of metrics viz.,makespan,speedup,efficiency and frequency of best results.展开更多
The problem of joint radio and cloud resources allocation is studied for heterogeneous mobile cloud computing networks. The objective of the proposed joint resource allocation schemes is to maximize the total utility ...The problem of joint radio and cloud resources allocation is studied for heterogeneous mobile cloud computing networks. The objective of the proposed joint resource allocation schemes is to maximize the total utility of users as well as satisfy the required quality of service(QoS) such as the end-to-end response latency experienced by each user. We formulate the problem of joint resource allocation as a combinatorial optimization problem. Three evolutionary approaches are considered to solve the problem: genetic algorithm(GA), ant colony optimization with genetic algorithm(ACO-GA), and quantum genetic algorithm(QGA). To decrease the time complexity, we propose a mapping process between the resource allocation matrix and the chromosome of GA, ACO-GA, and QGA, search the available radio and cloud resource pairs based on the resource availability matrixes for ACOGA, and encode the difference value between the allocated resources and the minimum resource requirement for QGA. Extensive simulation results show that our proposed methods greatly outperform the existing algorithms in terms of running time, the accuracy of final results, the total utility, resource utilization and the end-to-end response latency guaranteeing.展开更多
Many fast pattern-matching mechanisms are used in NIDS(Network Intrusion Detection Systems)to filter higher volumes of network traffic prior to invoking expensive rule verification stages.This filtering phase in signa...Many fast pattern-matching mechanisms are used in NIDS(Network Intrusion Detection Systems)to filter higher volumes of network traffic prior to invoking expensive rule verification stages.This filtering phase in signature-based engines,such as Snort,needs to preserve exact matching semantics while being able to process at high throughput on commodity hardware.Here,we introduce a hybrid CPU–GPU architecture-aware framework for exact multi-pattern matching based on the Weighted Exact Matching Algorithm(WEMA).WEMA performs the most relevant matching based on deterministic ordered indexing of category units,which eliminates chaotic control flow(which occurs with automata learning)and also brings out more regular memory access.An extensive evaluation across CPU-only,GPU-only,and hybrid CPU–GPU execution models is performed to analyze how WEMA interacts with heterogeneous hardware architectures.Informed by these observations,we propose a hybrid design with CPU-based control-intensive tasks—rule parsing,index construction,batching,and rule verification—retained on the CPU while selective data-parallel payload scanning is off-loaded to the GPU.Experimental results show that CPU-only execution on a commodity multicore CPU and integrated GPU platform delivers stable performance across the entire range of payload sizes,while the hybrid model achieves modest but repeatable improvements for small and medium workloads.The architecture evaluated here offers a relatively small advantage for GPU-only execution.Given the very nature of NIDS alongside the results provided in this paper,it is clear that WEMA-based acceleration on NIDS highly depends on hardware features and workload size,while selective,architecture-aware GPU utilization is crucial to maintain deterministic run-time characteristics in security-critical applications.展开更多
In this study,we investigate the ef-ficacy of a hybrid parallel algo-rithm aiming at enhancing the speed of evaluation of two-electron repulsion integrals(ERI)and Fock matrix generation on the Hygon C86/DCU(deep compu...In this study,we investigate the ef-ficacy of a hybrid parallel algo-rithm aiming at enhancing the speed of evaluation of two-electron repulsion integrals(ERI)and Fock matrix generation on the Hygon C86/DCU(deep computing unit)heterogeneous computing platform.Multiple hybrid parallel schemes are assessed using a range of model systems,including those with up to 1200 atoms and 10000 basis func-tions.The findings of our research reveal that,during Hartree-Fock(HF)calculations,a single DCU ex-hibits 33.6 speedups over 32 C86 CPU cores.Compared with the efficiency of Wuhan Electronic Structure Package on Intel X86 and NVIDIA A100 computing platform,the Hygon platform exhibits good cost-effective-ness,showing great potential in quantum chemistry calculation and other high-performance scientific computations.展开更多
Particle-in-cell (PIC) method has got much benefits from GPU-accelerated heterogeneous systems.However,the performance of PIC is constrained by the interpolation operations in the weighting process on GPU (graphic pro...Particle-in-cell (PIC) method has got much benefits from GPU-accelerated heterogeneous systems.However,the performance of PIC is constrained by the interpolation operations in the weighting process on GPU (graphic processing unit).Aiming at this problem,a fast weighting method for PIC simulation on GPU-accelerated systems was proposed to avoid the atomic memory operations during the weighting process.The method was implemented by taking advantage of GPU's thread synchronization mechanism and dividing the problem space properly.Moreover,software managed shared memory on the GPU was employed to buffer the intermediate data.The experimental results show that the method achieves speedups up to 3.5 times compared to previous works,and runs 20.08 times faster on one NVIDIA Tesla M2090 GPU compared to a single core of Intel Xeon X5670 CPU.展开更多
In recent years,with the development of processor architecture,heterogeneous processors including Center processing unit(CPU)and Graphics processing unit(GPU)have become the mainstream.However,due to the differences o...In recent years,with the development of processor architecture,heterogeneous processors including Center processing unit(CPU)and Graphics processing unit(GPU)have become the mainstream.However,due to the differences of heterogeneous core,the heterogeneous system is now facing many problems that need to be solved.In order to solve these problems,this paper try to focus on the utilization and efficiency of heterogeneous core and design some reasonable resource scheduling strategies.To improve the performance of the system,this paper proposes a combination strategy for a single task and a multi-task scheduling strategy for multiple tasks.The combination strategy consists of two sub-strategies,the first strategy improves the execution efficiency of tasks on the GPU by changing the thread organization structure.The second focuses on the working state of the efficient core and develops more reasonable workload balancing schemes to improve resource utilization of heterogeneous systems.The multi-task scheduling strategy obtains the execution efficiency of heterogeneous cores and global task information through the processing of task samples.Based on this information,an improved ant colony algorithm is used to quickly obtain a reasonable task allocation scheme,which fully utilizes the characteristics of heterogeneous cores.The experimental results show that the combination strategy reduces task execution time by 29.13%on average.In the case of processing multiple tasks,the multi-task scheduling strategy reduces the execution time by up to 23.38%based on the combined strategy.Both strategies can make better use of the resources of heterogeneous systems and significantly reduce the execution time of tasks on heterogeneous systems.展开更多
The Monte Carlo(MC)simulation is regarded as the gold standard for dose calculation in brachytherapy,but it consumes a large amount of computing resources.The development of heterogeneous computing makes it possible t...The Monte Carlo(MC)simulation is regarded as the gold standard for dose calculation in brachytherapy,but it consumes a large amount of computing resources.The development of heterogeneous computing makes it possible to substantially accelerate calculations with hardware accelerators.Accordingly,this study develops a fast MC tool,called THUBrachy,which can be accelerated by several types of hardware accelerators.THUBrachy can simulate photons with energy less than 3 MeV and considers all photon interactions in the energy range.It was benchmarked against the American Association of Physicists in Medicine Task Group No.43 Report using a water phantom and validated with Geant4 using a clinical case.A performance test was conducted using the clinical case,showing that a multicore central processing unit,Intel Xeon Phi,and graphics processing unit(GPU)can efficiently accelerate the simulation.GPU-accelerated THUBrachy is the fastest version,which is 200 times faster than the serial version and approximately 500 times faster than Geant4.The proposed tool shows great potential for fast and accurate dose calculations in clinical applications.展开更多
Molecular Dynamics(MD)simulation for computing Interatomic Potential(IAP)is a very important High-Performance Computing(HPC)application.MD simulation on particles of experimental relevance takes huge computation time,...Molecular Dynamics(MD)simulation for computing Interatomic Potential(IAP)is a very important High-Performance Computing(HPC)application.MD simulation on particles of experimental relevance takes huge computation time,despite using an expensive high-end server.Heterogeneous computing,a combination of the Field Programmable Gate Array(FPGA)and a computer,is proposed as a solution to compute MD simulation efficiently.In such heterogeneous computation,communication between FPGA and Computer is necessary.One such MD simulation,explained in the paper,is the(Artificial Neural Network)ANN-based IAP computation of gold(Au147&Au309)nanoparticles.MD simulation calculates the forces between atoms and the total energy of the chemical system.This work proposes the novel design and implementation of an ANN IAP-based MD simulation for Au147&Au309 using communication protocols,such as Universal Asynchronous Receiver-Transmitter(UART)and Ethernet,for communication between the FPGA and the host computer.To improve the latency of MD simulation through heterogeneous computing,Universal Asynchronous Receiver-Transmitter(UART)and Ethernet communication protocols were explored to conduct MD simulation of 50,000 cycles.In this study,computation times of 17.54 and 18.70 h were achieved with UART and Ethernet,respectively,compared to the conventional server time of 29 h for Au147 nanoparticles.The results pave the way for the development of a Lab-on-a-chip application.展开更多
DES (Data Encryption Standard) is one of the most classical algo- rithms of cryptography and its higher security makes it hard to be broke for a very long time. However, along with the constant development of comput...DES (Data Encryption Standard) is one of the most classical algo- rithms of cryptography and its higher security makes it hard to be broke for a very long time. However, along with the constant development of computer technology, especially in the 21st century, DES cannot be applied widely because of its low efficiency. Recently, the novel heterogeneous multi-core architecture represented by APU (Accelerated Processing Unit), provides a new solution for the above problems. APU integrates CPU and GPU in a ground- breaking manner and makes the algorithm to make full use of the performance advantage of heterogeneous multi-core system by realizing the HSA (Hetero- geneous System Architecture) standard. This paper realizes DES on the fresh APU processor. By analyzing the performance, two kinds of improved schemes are proposed. The experimental results show that the running efficiency of algorithm can be greatly improved by using APU with reasonable optimization. In the same way, the other DES-like algorithm would also be optimized on these heterogeneous multi-core architecture.展开更多
The possibility of carrying out a purely heterogeneous Heck reaction in practice without Pd leaching has been previously considered by a number of research groups but no general consent has yet arrived. Here, the reac...The possibility of carrying out a purely heterogeneous Heck reaction in practice without Pd leaching has been previously considered by a number of research groups but no general consent has yet arrived. Here, the reaction was, for the first time, evaluated by a simple computational approach. Modelling experiments were performed on one of the initial catalytic steps: phenyl halides attachment on Pd (111) to (100) and (111) to (111) ridges of a Pd crystal. Three surface structures of resulting were identified as possible reactive intermediates. Following potential energy minimisation calculations based on a universal force field, the relative stabilities of these surface species were then determined. Results showed the most stable species to be one in which a Pd ridge atom is removed from the Pd crystal structure, suggesting Pd leaching induced by phenyl halides is energetically favourable.展开更多
Stereoscopic and multiview rendering are used for virtual reality and the synthetic generation of light fields from three-dimensional scenes.Because rendering multiple views using ray tracing techniques is computation...Stereoscopic and multiview rendering are used for virtual reality and the synthetic generation of light fields from three-dimensional scenes.Because rendering multiple views using ray tracing techniques is computationally expensive,the utilization of multiprocessor machines is necessary to achieve real-time frame rates.In this study,we propose a dynamic load-balancing algorithm for real-time multiview path tracing on multi-compute device platforms.The proposed algorithm was adapted to heterogeneous hardware combinations and dynamic scenes in real time.We show that on a heterogeneous dual-GPU platform,our implementation reduces the rendering time by an average of approximately 30%–50%compared with that of a uniform workload distribution,depending on the scene and number of views.展开更多
Most natural resources are processed as particle-fluid multiphase systems in chemical,mineral and material indus-tries,therefore,discrete particles methods(DPM)are reasonable choices of simulation method for engineeri...Most natural resources are processed as particle-fluid multiphase systems in chemical,mineral and material indus-tries,therefore,discrete particles methods(DPM)are reasonable choices of simulation method for engineering the relevant processes and equipments.However,direct application of these methods is challenged by the complex multiscale behavior of such systems,which leads to enormous computational cost or otherwise qualitatively inac-curate description of the mesoscale structures.The coarse-grained DPM based on the energy-minimization multi-scale(EMMS)model,or EMMS-DPM,was proposed to reduce the computational cost by several orders while main-taining an accurate description of the mesoscale structures,which paves the way for its engineering applications.Further empowered by the high-efficiency multi-scale DEM software DEMms and the corresponding customized heterogeneous supercomputing facilities with graphics processing units(GPUs),it may even approach realtime simulation of industrial reactors.This short review will introduce the principle of DPM,in particular,EMMS-DPM,and the recent developments in modeling,numerical implementation and application of large-scale DPM which aims to reach industrial scale on one hand and resolves mesoscale structures critical to reaction-transport coupling on the other hand.This review finally prospects on the future developments of DPM in this direction.展开更多
基金funded by the State Grid Corporation Science and Technology Project“Research and Application of Key Technologies for Integrated Sensing and Computing for Intelligent Operation of Power Grid”(Grant No.5700-202318596A-3-2-ZN).
摘要With the rapid development of power Internet of Things(IoT)scenarios such as smart factories and smart homes,numerous intelligent terminal devices and real-time interactive applications impose higher demands on computing latency and resource supply efficiency.Multi-access edge computing technology deploys cloud computing capabilities at the network edge;constructs distributed computing nodes and multi-access systems and offers infrastructure support for services with low latency and high reliability.Existing research relies on a strong assumption that the environmental state is fully observable and fails to thoroughly consider the continuous time-varying features of edge server load fluctuations,leading to insufficient adaptability of the model in a heterogeneous dynamic environment.Thus,this paper establishes a framework for end-edge collaborative task offloading based on a partially observable Markov decision-making process(POMDP)and proposes a method for end-edge collaborative task offloading in heterogeneous scenarios.It achieves time-series modeling of the historical load characteristics of edge servers and endows the agent with the ability to be aware of the load in dynamic environmental states.Moreover,by dynamically assessing the exploration value of historical trajectories in the central trajectory pool and adjusting the sample weight distribution,directional exploration and strategy optimization of high-value trajectories are realized.Experimental results indicate that the proposed method exhibits distinct advantages compared with existing methods in terms of average delay and task failure rate and also verifies the method’s robustness in a dynamic environment.
基金supported by National Natural Science Foundation of China(No.U20A20158)Computing Power Foundation Strengthening Project of Ministry of Industry and Information Technology+1 种基金the Proof of Concept Foundation of Xidian University Hangzhou Institute of Technology(No.GNYZ2023GY0205)the significant science and technology project of Xiaoshan District(No.2023111).
摘要With the dramatic increase in the demand for computing power across various services,the emergence of the computing power network(CPN)becomes inevitable.This paper studies the task scheduling to minimize energy consumption under delay constraints considering the heterogeneity of computing resources.Specifically,we decompose the original problem and alternatively optimize the scheduling strategy and server parameters until convergence.Dynamic Voltage and frequency scaling(DVFS)technology is leveraged to allocate the optimal voltage and frequency for each server based on their task loads.An enhanced projection gradient descent method is utilized to update the scheduling strategy under the given server parameters.Simulation results show that our algorithm achieves significant performance gains compared to the baselines across various CPN scenarios.
基金supported by the Key Research and Development Program of China(2021YFC3000704)Institute of Geophysics,China Earthquake Administration Grant DQJB23R18+1 种基金the USTC Research Funds of the Double First-Class Initiative(YD2080002012)NSFC Grant(U2239206)。
摘要Ambient noise tomography is an established technique in seismology,where calculating single-or ninecomponent noise cross-correlation functions(NCFs)is a fundamental first step.In this study,we introduced a novel CPU-GPU heterogeneous computing framework designed to significantly enhance the efficiency of computing 9-component NCFs from seismic ambient noise data.This framework not only accelerated the computational process by leveraging the Compute Unified Device Architecture(CUDA)but also improved the signal-to-noise ratio(SNR)through innovative stacking techniques,such as time-frequency domain phaseweighted stacking(tf-PWS).We validated the program using multiple datasets,confirming its superior computation speed,improved reliability,and higher signal-to-noise ratios for NCFs.Our comprehensive study provides detailed insights into optimizing the computational processes for noise cross-correlation functions,thereby enhancing the precision and efficiency of ambient noise imaging.
基金supported by the National Science and Technology Council (NSTC)of Taiwan under Grants 108-2218-E-033-008-MY3,110-2634-F-A49-005,111-2221-E-033-033the Veterans General Hospitals and University System of Taiwan Joint Research Program under Grant VGHUST111-G6-5-1.
摘要Federated learning is an emerging machine learning techniquethat enables clients to collaboratively train a deep learning model withoutuploading raw data to the aggregation server. Each client may be equippedwith different computing resources for model training. The client equippedwith a lower computing capability requires more time for model training,resulting in a prolonged training time in federated learning. Moreover, it mayfail to train the entire model because of the out-of-memory issue. This studyaims to tackle these problems and propose the federated feature concatenate(FedFC) method for federated learning considering heterogeneous clients.FedFC leverages the model splitting and feature concatenate for offloadinga portion of the training loads from clients to the aggregation server. Eachclient in FedFC can collaboratively train a model with different cutting layers.Therefore, the specific features learned in the deeper layer of the serversidemodel are more identical for the data class classification. Accordingly,FedFC can reduce the computation loading for the resource-constrainedclient and accelerate the convergence time. The performance effectiveness isverified by considering different dataset scenarios, such as data and classimbalance for the participant clients in the experiments. The performanceimpacts of different cutting layers are evaluated during the model training.The experimental results show that the co-adapted features have a criticalimpact on the adequate classification of the deep learning model. Overall,FedFC not only shortens the convergence time, but also improves the bestaccuracy by up to 5.9% and 14.5% when compared to conventional federatedlearning and splitfed, respectively. In conclusion, the proposed approach isfeasible and effective for heterogeneous clients in federated learning.
基金Project supported by the National Natural Science Foundation of China (No. 60703012)the National Basic Research Program (973) of China (No. 2006CB303000)the Heilongjiang Provincial Scientific and Technological Special Fund for Young Scholars (No. QC06C033),China
摘要Heterogeneous computing (HC) environment utilizes diverse resources with different computational capabilities to solve computing-intensive applications having diverse computational requirements and constraints. The task assignment problem in HC environment can be formally defined as for a given set of tasks and machines, assigning tasks to machines to achieve the minimum makespan. In this paper we propose a new task scheduling heuristic, high standard deviation first (HSTDF), which considers the standard deviation of the expected execution time of a task as a selection criterion. Standard deviation of the ex- pected execution time of a task represents the amount of variation in task execution time on different machines. Our conclusion is that tasks having high standard deviation must be assigned first for scheduling. A large number of experiments were carried out to check the effectiveness of the proposed heuristic in different scenarios, and the comparison with the existing heuristics (Max-min, Sufferage, Segmented Min-average, Segmented Min-min, and Segmented Max-min) clearly reveals that the proposed heuristic outperforms all existing heuristics in terms of average makespan.
基金supported by the National Key Research and Development Program of China under Grant No.2022YFB4501400the Institute of Computing Technology,Chinese Academy of Sciences-China Mobile Communications Group Co.,Ltd.Joint Institute,the Beijing Nova Program under Grant Nos.20220484054 and 20230484420+1 种基金the Beijing Natural Science Foundation under Grant No.L234078the State Key Laboratory of Processors SKLP。
摘要Low-density parity check(LDPC)decoding is an efficient error correction method in communication systems,especially in 5G networks,which require high performance and low latency;while common general-purpose architectures cannot meet the requirements.There has been some research on accelerating LDPC decoding,but the current methods still suffer from limitations in performance,flexibility,and communication cost.In this paper,we propose HARLD(Heterogeneous Architecture of RISC-V for LDPC Decoding),a tightly coupled heterogeneous computing architecture based on extended RISC-V for LDPC decoding,consisting of a CPU and a processing array.Compared with a loosely coupled System-on-Chip(SoC)-bus baseline,the tightly coupled design improves throughput by up to 32.4%and reduces average latency by up to 24.7%across evaluated configurations,while also enhancing resource and energy efficiency:processing element utilization up to 93.5%,instruction RAM utilization increased by up to 4.8x,and energy efficiency improved by up to 24.8%.At the system level,area and power are reduced by 17.6%and 10.2%,respectively,versus the loosely coupled design.
基金supported by the National Natural Science Foundation of China(Grant No.61103223)the Natural Science Foundation of Jiangsu Province(No.BK2011003).
摘要To reduce the running time of network simulation in heterogeneous computing environment,a network simulation task partition method,named LBPHCE,is put forward.In this method,the network simulation task is partitioned in comprehensive consideration of the load balance of both routing computing simulation and packet forwarding simulation.First,through benchmark experiments,the computation ability and routing simulation ability of each simulation machine are measured in the heterogeneous computing environment.Second,based on the computation ability of each simulation machine,the network simulation task is initially partitioned to meet the load balance of packet forwarding simulation in the heterogeneous computing environment,and then according to the routing computation ability,the scale of each partition is fine-tuned to satisfy the balance of the routing computing simulation,meanwhile the load balance of packet forwarding simulation is guaranteed.Experiments based on PDNS indicate that,compared to traditional uniform partition method,the LBPHCE method can reduce the total simulation running time by 26.3%in average,and compared to the liner partition method,it can reduce the running time by 18.3%in average.
摘要Non-uniform sampling two-dimensional convolution (NUSC) maps spatially sampling data with irregular distribution to a regular grid by convolution. As the data scale and growth rate continue to increase, accelerating NUSC with the heterogene-ous computing platform is a feasible way. However, the complex hardware architecture and storage hierarchy of the hetero-geneous computing platform poses a challenge to programming and performance tuning. Therefore, this paper proposes a heterogeneous parallel programming model and runtime framework named AutoNUSC. For the programming difficulties of NUSC in heterogeneous computing environments, AutoNUSC abstracts and encapsulates the parallel execution process of NUSC. Task scheduling, data division, node communication, fault-tolerant recovery, and other parallelization tasks are managed by AutoNUSC. For the performance tuning issues of NUSC, this paper implements performance optimization strategies for AutoNUSC, including vectorization, memory access optimization, data reuse, etc. The experiments show that AutoNUSC effectively reduces the workload of users in developing NUSC applications in heterogeneous computing environments. Performance acceleration of up to 339 times is achieved within a single node compared to the serial program. AutoNUSC can efficiently perform task scheduling and fault-tolerant recovery across multiple nodes, with desirable scal-ability and robustness.
摘要Applications involving multifarious computational requirements take the advantage of the versatility of heterogeneous computing systems(HCS)with more than one type of parallelism.Efficient scheduling of workflow applications is paramount to harness high performance from HCS.In the present work,a new list-based heuristic strategy namely maximizing parallelism for minimizing earliest finish time(MPEFT)algorithm is proposed with a primary objective of minimizing the makespan.In order to minimize the makespan,the proposed scheduling policy focuses on proliferating the parallelism of the workflows by choosing the globally heaviest task with more number of successors such that more number of successors can be released.Thus,the priority policy maximizes the length of the ready queue by exploring higher degree of parallelism of the workflow.The proposed approach is designed to adapt depth-wise whenever the tasks at subsequent levels are released and continues to be level-wise otherwise.This increases the degree of parallelism and shortens the makespan.To evaluate the proposed scheduling algorithm,experimentations are conducted using randomly generated workflows and scientific workflows namely LIGO,Epigenomics,Cybershake,and Montage.The experimental results show that the proposed MPEFT algorithm surpassed the classical list based heuristic algorithms in terms of metrics viz.,makespan,speedup,efficiency and frequency of best results.
基金supported by the National Natural Science Foundation of China (No. 61741102, No. 61471164)China Scholarship Council
摘要The problem of joint radio and cloud resources allocation is studied for heterogeneous mobile cloud computing networks. The objective of the proposed joint resource allocation schemes is to maximize the total utility of users as well as satisfy the required quality of service(QoS) such as the end-to-end response latency experienced by each user. We formulate the problem of joint resource allocation as a combinatorial optimization problem. Three evolutionary approaches are considered to solve the problem: genetic algorithm(GA), ant colony optimization with genetic algorithm(ACO-GA), and quantum genetic algorithm(QGA). To decrease the time complexity, we propose a mapping process between the resource allocation matrix and the chromosome of GA, ACO-GA, and QGA, search the available radio and cloud resource pairs based on the resource availability matrixes for ACOGA, and encode the difference value between the allocated resources and the minimum resource requirement for QGA. Extensive simulation results show that our proposed methods greatly outperform the existing algorithms in terms of running time, the accuracy of final results, the total utility, resource utilization and the end-to-end response latency guaranteeing.
摘要Many fast pattern-matching mechanisms are used in NIDS(Network Intrusion Detection Systems)to filter higher volumes of network traffic prior to invoking expensive rule verification stages.This filtering phase in signature-based engines,such as Snort,needs to preserve exact matching semantics while being able to process at high throughput on commodity hardware.Here,we introduce a hybrid CPU–GPU architecture-aware framework for exact multi-pattern matching based on the Weighted Exact Matching Algorithm(WEMA).WEMA performs the most relevant matching based on deterministic ordered indexing of category units,which eliminates chaotic control flow(which occurs with automata learning)and also brings out more regular memory access.An extensive evaluation across CPU-only,GPU-only,and hybrid CPU–GPU execution models is performed to analyze how WEMA interacts with heterogeneous hardware architectures.Informed by these observations,we propose a hybrid design with CPU-based control-intensive tasks—rule parsing,index construction,batching,and rule verification—retained on the CPU while selective data-parallel payload scanning is off-loaded to the GPU.Experimental results show that CPU-only execution on a commodity multicore CPU and integrated GPU platform delivers stable performance across the entire range of payload sizes,while the hybrid model achieves modest but repeatable improvements for small and medium workloads.The architecture evaluated here offers a relatively small advantage for GPU-only execution.Given the very nature of NIDS alongside the results provided in this paper,it is clear that WEMA-based acceleration on NIDS highly depends on hardware features and workload size,while selective,architecture-aware GPU utilization is crucial to maintain deterministic run-time characteristics in security-critical applications.
基金supported by the National Natural Science Foundation of China(No.22373112 to Ji Qi,No.22373111 and 21921004 to Minghui Yang)GH-fund A(No.202107011790)。
摘要In this study,we investigate the ef-ficacy of a hybrid parallel algo-rithm aiming at enhancing the speed of evaluation of two-electron repulsion integrals(ERI)and Fock matrix generation on the Hygon C86/DCU(deep computing unit)heterogeneous computing platform.Multiple hybrid parallel schemes are assessed using a range of model systems,including those with up to 1200 atoms and 10000 basis func-tions.The findings of our research reveal that,during Hartree-Fock(HF)calculations,a single DCU ex-hibits 33.6 speedups over 32 C86 CPU cores.Compared with the efficiency of Wuhan Electronic Structure Package on Intel X86 and NVIDIA A100 computing platform,the Hygon platform exhibits good cost-effective-ness,showing great potential in quantum chemistry calculation and other high-performance scientific computations.
基金Projects(61170049,60903044)supported by National Natural Science Foundation of ChinaProject(2012AA010903)supported by National High Technology Research and Development Program of China
摘要Particle-in-cell (PIC) method has got much benefits from GPU-accelerated heterogeneous systems.However,the performance of PIC is constrained by the interpolation operations in the weighting process on GPU (graphic processing unit).Aiming at this problem,a fast weighting method for PIC simulation on GPU-accelerated systems was proposed to avoid the atomic memory operations during the weighting process.The method was implemented by taking advantage of GPU's thread synchronization mechanism and dividing the problem space properly.Moreover,software managed shared memory on the GPU was employed to buffer the intermediate data.The experimental results show that the method achieves speedups up to 3.5 times compared to previous works,and runs 20.08 times faster on one NVIDIA Tesla M2090 GPU compared to a single core of Intel Xeon X5670 CPU.
基金This work is supported by Beijing Natural Science Foundation[4192007]the National Natural Science Foundation of China[61202076]Beijing University of Technology Project No.2021C02.
摘要In recent years,with the development of processor architecture,heterogeneous processors including Center processing unit(CPU)and Graphics processing unit(GPU)have become the mainstream.However,due to the differences of heterogeneous core,the heterogeneous system is now facing many problems that need to be solved.In order to solve these problems,this paper try to focus on the utilization and efficiency of heterogeneous core and design some reasonable resource scheduling strategies.To improve the performance of the system,this paper proposes a combination strategy for a single task and a multi-task scheduling strategy for multiple tasks.The combination strategy consists of two sub-strategies,the first strategy improves the execution efficiency of tasks on the GPU by changing the thread organization structure.The second focuses on the working state of the efficient core and develops more reasonable workload balancing schemes to improve resource utilization of heterogeneous systems.The multi-task scheduling strategy obtains the execution efficiency of heterogeneous cores and global task information through the processing of task samples.Based on this information,an improved ant colony algorithm is used to quickly obtain a reasonable task allocation scheme,which fully utilizes the characteristics of heterogeneous cores.The experimental results show that the combination strategy reduces task execution time by 29.13%on average.In the case of processing multiple tasks,the multi-task scheduling strategy reduces the execution time by up to 23.38%based on the combined strategy.Both strategies can make better use of the resources of heterogeneous systems and significantly reduce the execution time of tasks on heterogeneous systems.
基金supported by the National Natural Science Foundation of China(No.11875036)。
摘要The Monte Carlo(MC)simulation is regarded as the gold standard for dose calculation in brachytherapy,but it consumes a large amount of computing resources.The development of heterogeneous computing makes it possible to substantially accelerate calculations with hardware accelerators.Accordingly,this study develops a fast MC tool,called THUBrachy,which can be accelerated by several types of hardware accelerators.THUBrachy can simulate photons with energy less than 3 MeV and considers all photon interactions in the energy range.It was benchmarked against the American Association of Physicists in Medicine Task Group No.43 Report using a water phantom and validated with Geant4 using a clinical case.A performance test was conducted using the clinical case,showing that a multicore central processing unit,Intel Xeon Phi,and graphics processing unit(GPU)can efficiently accelerate the simulation.GPU-accelerated THUBrachy is the fastest version,which is 200 times faster than the serial version and approximately 500 times faster than Geant4.The proposed tool shows great potential for fast and accurate dose calculations in clinical applications.
摘要Molecular Dynamics(MD)simulation for computing Interatomic Potential(IAP)is a very important High-Performance Computing(HPC)application.MD simulation on particles of experimental relevance takes huge computation time,despite using an expensive high-end server.Heterogeneous computing,a combination of the Field Programmable Gate Array(FPGA)and a computer,is proposed as a solution to compute MD simulation efficiently.In such heterogeneous computation,communication between FPGA and Computer is necessary.One such MD simulation,explained in the paper,is the(Artificial Neural Network)ANN-based IAP computation of gold(Au147&Au309)nanoparticles.MD simulation calculates the forces between atoms and the total energy of the chemical system.This work proposes the novel design and implementation of an ANN IAP-based MD simulation for Au147&Au309 using communication protocols,such as Universal Asynchronous Receiver-Transmitter(UART)and Ethernet,for communication between the FPGA and the host computer.To improve the latency of MD simulation through heterogeneous computing,Universal Asynchronous Receiver-Transmitter(UART)and Ethernet communication protocols were explored to conduct MD simulation of 50,000 cycles.In this study,computation times of 17.54 and 18.70 h were achieved with UART and Ethernet,respectively,compared to the conventional server time of 29 h for Au147 nanoparticles.The results pave the way for the development of a Lab-on-a-chip application.
摘要DES (Data Encryption Standard) is one of the most classical algo- rithms of cryptography and its higher security makes it hard to be broke for a very long time. However, along with the constant development of computer technology, especially in the 21st century, DES cannot be applied widely because of its low efficiency. Recently, the novel heterogeneous multi-core architecture represented by APU (Accelerated Processing Unit), provides a new solution for the above problems. APU integrates CPU and GPU in a ground- breaking manner and makes the algorithm to make full use of the performance advantage of heterogeneous multi-core system by realizing the HSA (Hetero- geneous System Architecture) standard. This paper realizes DES on the fresh APU processor. By analyzing the performance, two kinds of improved schemes are proposed. The experimental results show that the running efficiency of algorithm can be greatly improved by using APU with reasonable optimization. In the same way, the other DES-like algorithm would also be optimized on these heterogeneous multi-core architecture.
摘要The possibility of carrying out a purely heterogeneous Heck reaction in practice without Pd leaching has been previously considered by a number of research groups but no general consent has yet arrived. Here, the reaction was, for the first time, evaluated by a simple computational approach. Modelling experiments were performed on one of the initial catalytic steps: phenyl halides attachment on Pd (111) to (100) and (111) to (111) ridges of a Pd crystal. Three surface structures of resulting were identified as possible reactive intermediates. Following potential energy minimisation calculations based on a universal force field, the relative stabilities of these surface species were then determined. Results showed the most stable species to be one in which a Pd ridge atom is removed from the Pd crystal structure, suggesting Pd leaching induced by phenyl halides is energetically favourable.
基金Supported by the European Union’s Horizon 2020 Research and Innovation Programme under Marie Skłodowska-Curie grant agreement(No.956770)the Academy of Finland under Grant 325530.
摘要Stereoscopic and multiview rendering are used for virtual reality and the synthetic generation of light fields from three-dimensional scenes.Because rendering multiple views using ray tracing techniques is computationally expensive,the utilization of multiprocessor machines is necessary to achieve real-time frame rates.In this study,we propose a dynamic load-balancing algorithm for real-time multiview path tracing on multi-compute device platforms.The proposed algorithm was adapted to heterogeneous hardware combinations and dynamic scenes in real time.We show that on a heterogeneous dual-GPU platform,our implementation reduces the rendering time by an average of approximately 30%–50%compared with that of a uniform workload distribution,depending on the scene and number of views.
基金supported by the National Natural Sci-ence Foundation of China(Grant Nos.21978295,22078330,92034302 and 91834303)Innovation Academy for Green Manufacture,Chinese Academy of Sciences(Grant Nos.IAGM-2019-A03 and IAGM-2019-A13)+2 种基金Key Research Program of Frontier Sciences,Chinese Academy of Sciences(Grant No.QYZDJ-SSWJSC029)“Transformational Technologies for Clean Energy and Demonstration”Strategic Prior-ity Research Program of the Chinese Academy of Sciences(Grant No.XDA21030700)the Youth Innovation Promotion Association,Chinese Academy of Sciences(Grant No.2019050).
摘要Most natural resources are processed as particle-fluid multiphase systems in chemical,mineral and material indus-tries,therefore,discrete particles methods(DPM)are reasonable choices of simulation method for engineering the relevant processes and equipments.However,direct application of these methods is challenged by the complex multiscale behavior of such systems,which leads to enormous computational cost or otherwise qualitatively inac-curate description of the mesoscale structures.The coarse-grained DPM based on the energy-minimization multi-scale(EMMS)model,or EMMS-DPM,was proposed to reduce the computational cost by several orders while main-taining an accurate description of the mesoscale structures,which paves the way for its engineering applications.Further empowered by the high-efficiency multi-scale DEM software DEMms and the corresponding customized heterogeneous supercomputing facilities with graphics processing units(GPUs),it may even approach realtime simulation of industrial reactors.This short review will introduce the principle of DPM,in particular,EMMS-DPM,and the recent developments in modeling,numerical implementation and application of large-scale DPM which aims to reach industrial scale on one hand and resolves mesoscale structures critical to reaction-transport coupling on the other hand.This review finally prospects on the future developments of DPM in this direction.