期刊文献+
共找到940篇文章
< 1 2 47 >
每页显示 20 50 100
Complex hexagonal close-packed dendritic growth during alloy solidification by graphics processing unit-accelerated three-dimensional phase-field simulations:demo for Mg–Gd alloy 认领 引用 被引量:3
1
作者 Sheng-Lan Yang Jing Zhong +5 位作者 Kai Wang Xun Kang Jian-Bao Gao Jiong Wang Qian Li Li-Jun Zhang 《Rare Metals》 SCIE EI CAS CSCD 2023年第10期3468-3484,共17页
In this study,insights into the effect of interfacial anisotropy on a complex hexagonal close-packed(hcp) dendritic growth during alloy solidification were gained by graphics processing unit(GPU)-accelerated three-dim... In this study,insights into the effect of interfacial anisotropy on a complex hexagonal close-packed(hcp) dendritic growth during alloy solidification were gained by graphics processing unit(GPU)-accelerated three-dimensional(3D) phase-field simulations,as demonstrated for a Mg-Gd alloy.An anisotropic phasefield model with finite interface dissipation was developed by incorporating the contribution of the anisotropy of interfacial energy into the total free energy functional.The modified spherical harmonic anisotropy function was then chosen for the hcp crystal.The GPU parallel computing algorithm was implemented in the present phase-field model,and a corresponding code was developed in the compute unified device architecture parallel computing platform.Benchmark tests indicated that the calculation efficiency of a single TESLA V100 GPU could be~80times that of open multi-processing(OpenMP) with eight central processing unit cores.By coupling the phase-field model with reliable thermodynamic and interfacial energy descriptions,the 3D phase-field simulation of α-Mg dendritic growth in the Mg-6Gd(in wt%) alloy during solidification was performed.Various two-dimensional dendrite morphologies were revealed by cutting the simulated 3D dendrite along different crystallographic planes.Typical sixfold equiaxed and butterflied microstructures observed in experiments were well reproduced. 展开更多
关键词 Interfacial anisotropy Dendrite solidification Phase-field model Graphics processing unit(GPU) Mg–Gd
暂未订购 下载PDF
A graphics processing unit-based robust numerical model for solute transport driven by torrential flow condition 认领 引用 被引量:1
2
作者 Jing-ming HOU Bao-shan SHI +6 位作者 Qiu-hua LIANG Yu TONG Yong-de KANG Zhao-an ZHANG Gang-gang BAI Xu-jun GAO Xiao YANG 《Journal of Zhejiang University-SCIENCE A》 SCIE EI CAS CSCD 2021年第10期835-850,共16页
Solute transport simulations are important in water pollution events.This paper introduces a finite volume Godunovtype model for solving a 4×4 matrix form of the hyperbolic conservation laws consisting of 2D shal... Solute transport simulations are important in water pollution events.This paper introduces a finite volume Godunovtype model for solving a 4×4 matrix form of the hyperbolic conservation laws consisting of 2D shallow water equations and transport equations.The model adopts the Harten-Lax-van Leer-contact(HLLC)-approximate Riemann solution to calculate the cell interface fluxes.It can deal well with the changes in the dry and wet interfaces in an actual complex terrain,and it has a strong shock-wave capturing ability.Using monotonic upstream-centred scheme for conservation laws(MUSCL)linear reconstruction with finite slope and the Runge-Kutta time integration method can achieve second-order accuracy.At the same time,the introduction of graphics processing unit(GPU)-accelerated computing technology greatly increases the computing speed.The model is validated against multiple benchmarks,and the results are in good agreement with analytical solutions and other published numerical predictions.The third test case uses the GPU and central processing unit(CPU)calculation models which take 3.865 s and 13.865 s,respectively,indicating that the GPU calculation model can increase the calculation speed by 3.6 times.In the fourth test case,comparing the numerical model calculated by GPU with the traditional numerical model calculated by CPU,the calculation efficiencies of the numerical model calculated by GPU under different resolution grids are 9.8–44.6 times higher than those by CPU.Therefore,it has better potential than previous models for large-scale simulation of solute transport in water pollution incidents.It can provide a reliable theoretical basis and strong data support in the rapid assessment and early warning of water pollution accidents. 展开更多
关键词 Solute transport Shallow water equations Godunov-type scheme Harten-Lax-van Leer-contact(HLLC)Riemann solver Graphics processing unit(GPU)acceleration technology Torrential flow
暂未订购 下载PDF
Simulation of fluid-structure interaction in a microchannel using the lattice Boltzmann method and size-dependent beam element on a graphics processing unit 认领 引用 被引量:4
3
作者 Vahid Esfahanian Esmaeil Dehdashti Amir Mehdi Dehrouye-Semnani 《Chinese Physics B》 SCIE EI CAS CSCD 2014年第8期389-395,共7页
Fluid-structure interaction (FSI) problems in microchannels play a prominent role in many engineering applications. The present study is an effort toward the simulation of flow in microchannel considering FSI. The b... Fluid-structure interaction (FSI) problems in microchannels play a prominent role in many engineering applications. The present study is an effort toward the simulation of flow in microchannel considering FSI. The bottom boundary of the microchannel is simulated by size-dependent beam elements for the finite element method (FEM) based on a modified cou- ple stress theory. The lattice Boltzmann method (LBM) using the D2Q13 LB model is coupled to the FEM in order to solve the fluid part of the FSI problem. Because of the fact that the LBM generally needs only nearest neighbor information, the algorithm is an ideal candidate for parallel computing. The simulations are carried out on graphics processing units (GPUs) using computed unified device architecture (CUDA). In the present study, the governing equations are non-dimensionalized and the set of dimensionless groups is exhibited to show their effects on micro-beam displacement. The numerical results show that the displacements of the micro-beam predicted by the size-dependent beam element are smaller than those by the classical beam element. 展开更多
关键词 fluid-structure interaction graphics processing unit lattice Boltzmann method size-dependentbeam element
暂未订购 下载PDF
Multi-relaxation-time lattice Boltzmann simulations of lid driven flows using graphics processing unit 认领 引用 被引量:1
4
作者 Chenggong LI J.P.Y.MAA 《Applied Mathematics and Mechanics(English Edition)》 SCIE EI CSCD 2017年第5期707-722,共16页
Large eddy simulation (LES) using the Smagorinsky eddy viscosity model is added to the two-dimensional nine velocity components (D2Q9) lattice Boltzmann equation (LBE) with multi-relaxation-time (MRT) to simul... Large eddy simulation (LES) using the Smagorinsky eddy viscosity model is added to the two-dimensional nine velocity components (D2Q9) lattice Boltzmann equation (LBE) with multi-relaxation-time (MRT) to simulate incompressible turbulent cavity flows with the Reynolds numbers up to 1 × 10^7. To improve the computation efficiency of LBM on the numerical simulations of turbulent flows, the massively parallel computing power from a graphic processing unit (GPU) with a computing unified device architecture (CUDA) is introduced into the MRT-LBE-LES model. The model performs well, compared with the results from others, with an increase of 76 times in computation efficiency. It appears that the higher the Reynolds numbers is, the smaller the Smagorinsky constant should be, if the lattice number is fixed. Also, for a selected high Reynolds number and a selected proper Smagorinsky constant, there is a minimum requirement for the lattice number so that the Smagorinsky eddy viscosity will not be excessively large. 展开更多
关键词 large eddy simulation (LES) multi-relaxation-time (MRT) lattice Boltzmann equation (LBE) two-dimensional nine velocity components (D2Q9) Smagorinskymodel graphic processing unit GPU computing unified device architecture (CUDA)
暂未订购 下载PDF
面向GPU的稀疏对角矩阵自适应SpMV优化方法 认领 引用
5
作者 王宇华 何俊飞 +2 位作者 张宇琪 兰海燕 曹林琳 《计算机工程》 CAS CSCD 北大核心 2026年第3期332-345,共14页
稀疏矩阵向量乘(SpMV)是稀疏线性系统的计算核心和瓶颈,其运算效率会影响迭代求解器的整体性能,其优化研究一直是科学计算和工程应用领域中的研究热点之一。偏微分方程的离散化会产生稀疏对角矩阵,由于其多样的非零元分布,导致没有一种... 稀疏矩阵向量乘(SpMV)是稀疏线性系统的计算核心和瓶颈,其运算效率会影响迭代求解器的整体性能,其优化研究一直是科学计算和工程应用领域中的研究热点之一。偏微分方程的离散化会产生稀疏对角矩阵,由于其多样的非零元分布,导致没有一种方法能够在所有矩阵中取得最优时间性能。针对上述问题,提出一种面向图形处理单元(GPU)的稀疏对角矩阵自适应SpMV优化方法AST(Adaptive SpMV Tuning)。该方法通过设计特征空间,构建特征提取器,提取矩阵结构精细特征,通过深入分析特征和SpMV方法的相关性,建立可扩展的候选方法集合,形成特征和最优方法的映射关系,构建性能预测工具,实现矩阵最优方法的高效预测。实验结果表明,AST能够取得85.8%的预测准确率,平均时间性能损失为0.09,相比于DIA(Diagonal)、HDIA(Hacked DIA)、HDC(Hybrid of DIA and Compressed Sparse Row)、DIA-Adaptive和DRM(Divide-Rearrange and Merge),能够获得平均20.19、1.86、3.06、3.72和1.53倍的内核运行时间加速和1.05、1.28、12.45、1.94和0.97倍的浮点运算性能加速。 展开更多
关键词 稀疏矩阵向量乘 稀疏对角矩阵 图形处理单元 自适应优化方法 矩阵结构特征
暂未订购 下载PDF
非对称容差关系粗糙近似集的GPU加速方法 认领 引用
6
作者 吴正江 王梦松 武星晨 《计算机工程》 CAS CSCD 北大核心 2026年第8期247-259,共13页
大数据时代信息来源纷繁复杂,收集数据形成的信息系统很难保证其完整性。在越来越大的不完备信息系统(IIS)中,使用基于非对称容差关系的粗糙集理论进行知识蒸馏,提升近似集的计算速度成为其应用之前必须要解决的问题。针对非对称容差关... 大数据时代信息来源纷繁复杂,收集数据形成的信息系统很难保证其完整性。在越来越大的不完备信息系统(IIS)中,使用基于非对称容差关系的粗糙集理论进行知识蒸馏,提升近似集的计算速度成为其应用之前必须要解决的问题。针对非对称容差关系的冗余容差类问题,提出非对称容差关系的改进方案,设计不完备信息系统中上、下近似集的布尔矩阵表示方法,计算非对称容差关系的粗糙近似集的矩阵分块算法,并在图像处理单元(GPU)上实现近似集计算过程的加速,提高近似集的计算效率。另外,针对GPU存储空间有限的现状,构建不完备信息系统中对象之间的层次结构及其算法,将全局关系矩阵转换成多个局部关系矩阵,缓解计算最近容差类过程中GPU存储压力。在UCI数据集和生成数据集上的实验结果表明,基于最近容差关系的容差类数量相较于基准情况明显减少,同时,矩阵分块算法在GPU上实现了对近似集计算过程的有效加速,相比CPU串行计算和分布式并行计算,GPU分块算法执行速度平均提高了16.69倍和3.89倍。 展开更多
关键词 图像处理单元 近似集 最近容差关系 矩阵分块 不完备信息系统
暂未订购 下载PDF
基于GPU的大规模新能源电网EMT仿真并行加速方法 认领 引用
7
作者 于智同 杨明皓 +3 位作者 宋炎侃 黄少伟 陈颖 沈沉 《电力自动化设备》 EI CSCD 北大核心 2026年第8期123-131,共9页
高比例新能源并网使含海量电力电子设备的大规模电磁暂态仿真在计算效率与规模扩展性上面临瓶颈。图形处理器具备高吞吐并行优势,但直接应用时面临控制拓扑异构导致的线程束发散、图式计算带来的访存不连续和动态解析引起的指令冗余三... 高比例新能源并网使含海量电力电子设备的大规模电磁暂态仿真在计算效率与规模扩展性上面临瓶颈。图形处理器具备高吞吐并行优势,但直接应用时面临控制拓扑异构导致的线程束发散、图式计算带来的访存不连续和动态解析引起的指令冗余三大难点。为此,提出一种面向大规模新能源电网的异构并行加速方法。引入图同构思想,利用图聚类算法实现控制系统的拓扑级识别与聚合,消除异构控制器导致的指令流分歧;构建面向图形处理器的自动代码生成框架,将聚合后的模型转化为无分支、访存规整的计算内核,提升执行效率。算例测试表明,在包含数千台新能源设备的仿真场景中,所提方法相比传统仿真方法可获得近3倍的加速比,且具备优异的规模可扩展性。 展开更多
关键词 电磁暂态仿真 图形处理器 细粒度并行 图同构 自动代码生成 异构计算
暂未订购 下载PDF
Exploiting Parallelism in the Simulation of General Purpose Graphics Processing Unit Program 认领 引用
8
作者 赵夏 马胜 +1 位作者 陈微 王志英 《Journal of Shanghai Jiaotong university(Science)》 EI 2016年第3期280-288,共9页
The simulation is an important means of performance evaluation of the computer architecture. Nowadays, the serial simulation of general purpose graphics processing unit(GPGPU) architecture is the main bottleneck for t... The simulation is an important means of performance evaluation of the computer architecture. Nowadays, the serial simulation of general purpose graphics processing unit(GPGPU) architecture is the main bottleneck for the simulation speed. To address this issue, we propose the intra-kernel parallelization on a multicore processor and the inter-kernel parallelization on a multiple-machine platform. We apply these two methods to the GPGPU-sim simulator. The intra-kernel parallelization method firstly parallelizes the serial simulation of multiple compute units in one cycle. Then it parallelizes the timing and functional simulation to reduce the performance loss caused by the synchronization between different compute units. The inter-kernel parallelization method divides multiple kernels of a CUDA program into several groups and distributes these groups across multiple simulation hosts to perform the simulation. Experimental results show that the intra-kernel parallelization method achieves a speed-up of up to 12 with a maximum error rate of 0.009 4% on a 32-core machine, and the inter-kernel parallelization method can accelerate the simulation by a factor of up to 3.9 with a maximum error rate of 0.11% on four simulation hosts. The orthogonality between these two methods allows us to combine them together on multiple multi-core hosts to get further performance improvements. 展开更多
关键词 general purpose graphics processing unit(GPGPU) multicore intra-kernel inter-kernel parallel
暂未订购 下载PDF
Compute Unified Device Architecture Implementation of Euler/Navier-Stokes Solver on Graphics Processing Unit Desktop Platform for 2-D Compressible Flows 认领 引用
9
作者 Zhang Jiale Chen Hongquan 《Transactions of Nanjing University of Aeronautics and Astronautics》 EI CSCD 2016年第5期536-545,共10页
Personal desktop platform with teraflops peak performance of thousands of cores is realized at the price of conventional workstations using the programmable graphics processing units(GPUs).A GPU-based parallel Euler/N... Personal desktop platform with teraflops peak performance of thousands of cores is realized at the price of conventional workstations using the programmable graphics processing units(GPUs).A GPU-based parallel Euler/Navier-Stokes solver is developed for 2-D compressible flows by using NVIDIA′s Compute Unified Device Architecture(CUDA)programming model in CUDA Fortran programming language.The techniques of implementation of CUDA kernels,double-layered thread hierarchy and variety memory hierarchy are presented to form the GPU-based algorithm of Euler/Navier-Stokes equations.The resulting parallel solver is validated by a set of typical test flow cases.The numerical results show that dozens of times speedup relative to a serial CPU implementation can be achieved using a single GPU desktop platform,which demonstrates that a GPU desktop can serve as a costeffective parallel computing platform to accelerate computational fluid dynamics(CFD)simulations substantially. 展开更多
关键词 graphics processing unit(GPU) GPU parallel computing compute unified device architecture(CUDA)Fortran finite volume method(FVM) acceleration
暂未订购 下载PDF
面向GPU的低能耗数据传输的组重映射编码方法 认领 引用
10
作者 章铁飞 邢建国 《计算机工程与科学》 CSCD 北大核心 2026年第5期803-809,共7页
现代图形处理单元(GPU)的高性能计算能力,依赖于高带宽的图形DDR(GDDR)接口。高带宽的数据传输速率导致高能耗,特别是GDDR的伪开漏(POD)I/O接口中传输逻辑1值的不对称能耗。通过减少数据传输过程中高能耗的逻辑1值,可以缓解数据传输时... 现代图形处理单元(GPU)的高性能计算能力,依赖于高带宽的图形DDR(GDDR)接口。高带宽的数据传输速率导致高能耗,特别是GDDR的伪开漏(POD)I/O接口中传输逻辑1值的不对称能耗。通过减少数据传输过程中高能耗的逻辑1值,可以缓解数据传输时的高能耗问题。提出一种基于逻辑1值数量的组重映射编码方法。首先,将待传输数据按4位划分为基本单元,根据单元包含的逻辑1值数量再分组,然后将包含逻辑1值较多且数量较多的组映射编码为包含逻辑1值较少且数量较少的组,以最小化全局的逻辑1值数量。在现代GPU架构上评估,结果显示组重映射编码方法可以有效减少各种应用程序在数据传输时的逻辑1值的数量,平均降低比例达到26%,证明了方法的有效性。 展开更多
关键词 数据传输能耗 GDDR I/O接口 组重映射编码 图形处理器
暂未订购 下载PDF
TIME-DOMAIN INTERPOLATION ON GRAPHICS PROCESSING UNIT 认领 引用 被引量:3
11
作者 XIQI LI GUOHUA SHI YUDONG ZHANG 《Journal of Innovative Optical Health Sciences》 SCIE EI 2011年第1期89-95,共7页
The signal processing speed of spectral domain optical coherence tomography(SD-OCT)has become a bottleneck in a lot of medical applications.Recently,a time-domain interpolation method was proposed.This method can get ... The signal processing speed of spectral domain optical coherence tomography(SD-OCT)has become a bottleneck in a lot of medical applications.Recently,a time-domain interpolation method was proposed.This method can get better signal-to-noise ratio(SNR)but much-reduced signal processing time in SD-OCT data processing as compared with the commonly used zeropadding interpolation method.Additionally,the resampled data can be obtained by a few data and coefficients in the cutoff window.Thus,a lot of interpolations can be performed simultaneously.So,this interpolation method is suitable for parallel computing.By using graphics processing unit(GPU)and the compute unified device architecture(CUDA)program model,time-domain interpolation can be accelerated significantly.The computing capability can be achieved more than 250,000 A-lines,200,000 A-lines,and 160,000 A-lines in a second for 2,048 pixel OCT when the cutoff length is L=11,L=21,and L=31,respectively.A frame SD-OCT data(400A-lines×2,048 pixel per line)is acquired and processed on GPU in real time.The results show that signal processing time of SD-OCT can befinished in 6.223 ms when the cutoff length L=21,which is much faster than that on central processing unit(CPU).Real-time signal processing of acquired data can be realized. 展开更多
关键词 Optical coherence tomography real-time signal processing graphics processing unit GPU CUDA
暂未订购 下载PDF
The inversion of density structure by graphic processing unit(GPU) and identification of igneous rocks in Xisha area 认领 引用 被引量:1
12
作者 Lei Yu Jian Zhang +2 位作者 Wei Lin Rongqiang Wei Shiguo Wu 《Earthquake Science》 2014年第1期117-125,共9页
Organic reefs, the targets of deep-water petro- leum exploration, developed widely in Xisha area. However, there are concealed igneous rocks undersea, to which organic rocks have nearly equal wave impedance. So the ig... Organic reefs, the targets of deep-water petro- leum exploration, developed widely in Xisha area. However, there are concealed igneous rocks undersea, to which organic rocks have nearly equal wave impedance. So the igneous rocks have become interference for future explo- ration by having similar seismic reflection characteristics. Yet, the density and magnetism of organic reefs are very different from igneous rocks. It has obvious advantages to identify organic reefs and igneous rocks by gravity and magnetic data. At first, frequency decomposition was applied to the free-air gravity anomaly in Xisha area to obtain the 2D subdivision of the gravity anomaly and magnetic anomaly in the vertical direction. Thus, the dis- tribution of igneous rocks in the horizontal direction can be acquired according to high-frequency field, low-frequency field, and its physical properties. Then, 3D forward model- ing of gravitational field was carried out to establish the density model of this area by reference to physical properties of rocks based on former researches. Furthermore, 3D inversion of gravity anomaly by genetic algorithm method of the graphic processing unit (GPU) parallel processing in Xisha target area was applied, and 3D density structure of this area was obtained. By this way, we can confine the igneous rocks to the certain depth according to the density of the igneous rocks. The frequency decomposition and 3D inversion of gravity anomaly by genetic algorithm method of the GPU parallel processing proved to be a useful method for recognizing igneous rocks to its 3D geological position. So organic reefs and igneous rocks can be identified, which provide a prescient information for further exploration. 展开更多
关键词 Xisha area Organic reefs and igneous rocks -Frequency decomposition of potential field 3D inversionof the graphic processing unit GPU parallel processing
暂未订购 下载PDF
KSSOLV-GPU 2.0:A Lightweight MATLAB Toolbox for GPU-Accelerated Plane-Wave Hybrid Functional and Spin-Polarized Density Functional Theory Calculations 认领 引用
13
作者 Yehao Zeng Zhenlin Zhang +10 位作者 Zhaolong Luo Jun Gao Sheng Chen Wentiao Wu Yexuan Lin Yihong Zhang Xinming Qin Weiyi Wang Zhiwen Zhuo Shizhe Jiao Wei Hu 《Chinese Journal of Chemical Physics》 SCIE EI CAS CSCD 2026年第2期182-196,I0169,共15页
KSSOLV(Kohn−Sham solver)is a MAT-LAB(Matrix Laboratory)toolbox de-signed for solving the Kohn-Sham density functional theory(DFT)equations by us-ing the plane-wave basis set.Leveraging the powerful capabilities of MAT... KSSOLV(Kohn−Sham solver)is a MAT-LAB(Matrix Laboratory)toolbox de-signed for solving the Kohn-Sham density functional theory(DFT)equations by us-ing the plane-wave basis set.Leveraging the powerful capabilities of MATLAB’s parallel computing toolbox and an ad-vanced,optimized calculation workflow,KSSOLV uniquely enables efficient graph-ics processing unit(GPU)acceleration,making DFT calculations accessible on standard personal computing hardware.Here,KSSOLV-GPU 2.0,as the latest release,demonstrates substantial computational gains.In benchmarks,particularly involving calculations such as hybrid functionals and spin-polar-ized systems for complex band structure analysis,KSSOLV-GPU 2.0 achieves a speedup of more than an order of magnitude compared to conventional central processing unit based im-plementations.This significant acceleration marks a pivotal advancement in performing com-plex materials simulations,making KS-DFT increasingly accessible on personal computing platforms. 展开更多
关键词 Kohn-Sham solver Density functional theory Matrix Laboratory(MATLAB) Graphics processing unit(GPU) Hybrid functional k-point sampling Spin polarization
暂未订购 下载PDF
Bypass-Enabled Thread Compaction for Divergent Control Flow in Graphics Processing Units 认领 引用
14
作者 LI Bingchao WEI Jizeng +1 位作者 GUO Wei SUN Jizhou 《Journal of Shanghai Jiaotong university(Science)》 EI 2021年第2期245-256,共12页
Graphics processing units(GPUs)employ the single instruction multiple data(SIMD)hardware to run threads in parallel and allow each thread to maintain an arbitrary control flow.Threads running concurrently within a war... Graphics processing units(GPUs)employ the single instruction multiple data(SIMD)hardware to run threads in parallel and allow each thread to maintain an arbitrary control flow.Threads running concurrently within a warp may jump to different paths after conditional branches.Such divergent control flow makes some lanes idle and hence reduces the SIMD utilization of GPUs.To alleviate the waste of SIMD lanes,threads from multiple warps can be collected together to improve the SIMD lane utilization by compacting threads into idle lanes.However,this mechanism induces extra barrier synchronizations since warps have to be stalled to wait for other warps for compactions,resulting in that no warps are scheduled in some cases.In this paper,we propose an approach to reduce the overhead of barrier synchronizat ions induced by compactions,In our approach,a compaction is bypassed by warps whose threads all jump to the same path after branches.Moreover,warps waiting for a compaction can also bypass this compaction when no warps are ready for issuing.In addition,a compaction is canceled if idle lanes can not be reduced via this compaction.The experimental results demonstrate that our approach provides an average improvement of 21%over the baseline GPU for applications with massive divergent branches,while recovering the performance loss induced by compactions by 13%on average for applications with many non-divergent control flows. 展开更多
关键词 graphics processing unit(GPU) single instruction ultiple data(SIMD) thread warps bypass
暂未订购 下载PDF
基于GPU的DSSS水声通信多普勒快速估计方法 认领 引用
15
作者 彭海源 李宇 +2 位作者 朱雨男 普湛清 王巍 《系统工程与电子技术》 EI CSCD 北大核心 2026年第6期2105-2112,共8页
直接序列扩频(direct sequence spread spectrum,DSSS)水声通信中基于交叉模糊函数(cross ambiguity function,CAF)的多普勒估计方法具有计算密集、求解耗时长的缺点。为了提高多普勒估计效率,提出一种基于图形处理器(graphic processin... 直接序列扩频(direct sequence spread spectrum,DSSS)水声通信中基于交叉模糊函数(cross ambiguity function,CAF)的多普勒估计方法具有计算密集、求解耗时长的缺点。为了提高多普勒估计效率,提出一种基于图形处理器(graphic processing unit,GPU)的多普勒快速估计方法。首先,为了降低CAF多普勒估计方法的计算复杂度,提出一种改进两步式多普勒估计算法。其次,从设备内存预加载、运算单元并行化、GPU算法部署3个方面进行算法实现。最后,使用实际海试数据对方法进行测试并与CPU估计方法进行对比。测试结果表明,提出的GPU多普勒快速估计方法能够实现DSSS水声通信信号多普勒因子的准确估计。与CPU估计方法相比,最大加速比为263.65,极大地提高了多普勒估计算法的执行效率。 展开更多
关键词 直接序列扩频 水声通信 图形处理器 多普勒估计
暂未订购 下载PDF
GPUMDkit:A User‐Friendly Toolkit for GPUMD and NEP 认领 引用
16
作者 Zihan Yan Denan Li +20 位作者 Xin Wu Zhoulin Liu Chen Hua Boyi Situ Hao Yang Shengjie Tang Benrui Tang Ziyang Wang Shangzhao Yi Huan Wang Dian Huang Ke Li Qilin Guo Zherui Chen Ke Xu Yanzhou Wang Ziliang Wang Gang Tang Shi Liu Zheyong Fan Yizhou Zhu 《Materials Genome Engineering Advances》 CAS CSCD 2026年第2期121-133,共13页
Machine‐learned interatomic potentials have revolutionized molecular dynamics simulations by providing quantum-mechanical accuracy at empirical-potential speeds.The graphics processing unit molecular dynamics(GPUMD)p... Machine‐learned interatomic potentials have revolutionized molecular dynamics simulations by providing quantum-mechanical accuracy at empirical-potential speeds.The graphics processing unit molecular dynamics(GPUMD)package,featuring the highly efficient neuroevolution potential(NEP)framework,has emerged as a powerful tool in this domain.However,the complexity of force field development,active learning,and trajectory post‐processing often requires extensive manual scripting,imposing a steep learning curve on new users.To address this,we present GPUMDkit,a comprehensive and user‐friendly toolkit that streamlines the entire simulation workflow for GPUMD and NEP.GPUMDkit integrates a suite of essential functionalities,including format conversion,structure sampling,property calculation,and data visualization,accessible through both interactive and command‐line interfaces.Its modular,extensible architecture ensures accessibility for users of all experience levels while allowing seamless integration of new features.By automating complex tasks and enhancing productivity,GPUMDkit substantially lowers the barrier to using GPUMD and NEP programs.This article describes the program architecture and demonstrates its capabilities through practical applications. 展开更多
关键词 active learning force field developmentactive learningand gpumd machine learned interatomic potentials NEP force field development graphics processing unit molecular dynamics gpumd packagefeaturing molecular dynamics
基于斯托克斯平面近似函数与GPU并行的海洋重力梯度模型计算 认领 引用
17
作者 卜靖宇 叶周润 +3 位作者 梁星辉 刘金钊 柳林涛 王嘉琛 《合肥工业大学学报(自然科学版)》 CAS 北大核心 2026年第2期253-259,共7页
相对于其他重力场元素,扰动重力梯度能更多地反映变化的不规则地球产生的高频信息。在计算扰动重力梯度时,由于斯托克斯积分较为复杂导致被积函数复杂难以直接用牛顿-莱布尼茨公式计算、且计算的数据量过于庞大导致计算耗时过长。为有... 相对于其他重力场元素,扰动重力梯度能更多地反映变化的不规则地球产生的高频信息。在计算扰动重力梯度时,由于斯托克斯积分较为复杂导致被积函数复杂难以直接用牛顿-莱布尼茨公式计算、且计算的数据量过于庞大导致计算耗时过长。为有效解决该问题,文章使用高斯数值积分解决被积函数复杂的问题,同时利用统一计算设备架构(compute unified device architecture,CUDA)在计算过程中实现了在图形处理器(graphics processing unit,GPU)端的并行计算,根据拉普拉斯方程可以检验计算结果的准确性,并且选取了某海域3°×2°范围海平面的重力异常数据进行计算。结果表明,使用高斯数值积分以及CUDA并行计算的方法,提供准确计算结果的同时也提高了计算效率。 展开更多
关键词 扰动重力梯度 重力异常 CUDA并行计算 图形处理器(GPU) 高斯数值积分
暂未订购 下载PDF
容器云环境GPU共享技术研究与实现 认领 引用
18
作者 吴阳阳 吴恒 +1 位作者 唐震 张文博 《广西大学学报(自然科学版)》 CAS 北大核心 2026年第1期177-187,共11页
针对传统基于时间分片的图形处理单元(GPU)共享方案中容器在时间片内独占GPU而导致的任务低负载时资源浪费问题,提出一种GPU共享框架(TQShare)。TQShare整合核函数执行时间预测技术与核函数时间配额管理机制,支持多个容器在同一时间片... 针对传统基于时间分片的图形处理单元(GPU)共享方案中容器在时间片内独占GPU而导致的任务低负载时资源浪费问题,提出一种GPU共享框架(TQShare)。TQShare整合核函数执行时间预测技术与核函数时间配额管理机制,支持多个容器在同一时间片内并发执行深度学习任务,从而提升资源利用率;同时,通过对任务启动的核函数进行动态调度管理,实现资源的有效隔离。实验结果表明,与KubeShare相比,TQShare将GPU平均利用率提高了13.4个百分点,深度学习工作负载完成时间缩短了14.7%,且平均性能开销仅为2.41%。 展开更多
关键词 容器 图形处理单元共享 资源隔离 图形处理单元利用率 深度学习
暂未订购 下载PDF
基于最小熵的地基MIMO-SAR幅相校正GPU加速与嵌入式实现 认领 引用
19
作者 郤伟杰 赖涛 +2 位作者 梁高天 王青松 黄海风 《系统工程与电子技术》 EI CSCD 北大核心 2026年第8期2615-2625,共11页
针对多输入多输出(multiple input multiple output,MIMO)合成孔径雷达(synthetic aperture radar,SAR)中的通道幅相误差问题,最小化图像熵幅相校正算法是一种有效的解决方法,但算法运算量大、估计时间长,难以用于实时校正。对此,提出... 针对多输入多输出(multiple input multiple output,MIMO)合成孔径雷达(synthetic aperture radar,SAR)中的通道幅相误差问题,最小化图像熵幅相校正算法是一种有效的解决方法,但算法运算量大、估计时间长,难以用于实时校正。对此,提出一种最小熵幅相校正的图形处理单元(graphics processing unit,GPU)实现方法,在数据传输和运算方面使用混合精度计算,保证算法精度的同时大幅提升了运算效率;通过异步并行流处理技术,充分利用GPU资源对算法进行加速,并最终将其部署在嵌入式GPU平台上。通过对实测数据的对比分析,桌面级GPU以及嵌入式GPU上的测试结果证明了所提方法的准确性和高效性,该算法可以对幅相误差进行有效估计,加速比最高可达到105%。 展开更多
关键词 合成孔径雷达 多输入多输出雷达 通道幅相校正 最小熵算法 图形处理单元加速
暂未订购 下载PDF
近场三维CZT波束形成算法的GPU实现及性能优化 认领 引用
20
作者 徐浚洋 刘祖延 +2 位作者 于晓阳 周天 陈宝伟 《应用声学》 CSCD 北大核心 2026年第2期434-443,共10页
针对在面阵波束形成过程中运算量大、难以做到实时成像的问题,文章使用图形处理器(GPU)在Visual Studio2019平台上对三维线性调频Z变换(CZT)波束形成算法进行加速,实现了三维CZT波束形成算法的并行化,从存储结构和对数据的访存等方面进... 针对在面阵波束形成过程中运算量大、难以做到实时成像的问题,文章使用图形处理器(GPU)在Visual Studio2019平台上对三维线性调频Z变换(CZT)波束形成算法进行加速,实现了三维CZT波束形成算法的并行化,从存储结构和对数据的访存等方面进行了针对性的设计,有效地利用了GPU的单指令多线程的特性,这些改进提升了算法的运行效率。通过实测数据显示,对于相同的声呐数据,GPU并行处理的计算效率高于CPU串行处理38倍以上,在采样点数量较少的情况下,三维CZT波束形成算法的计算效率明显优于传统的相移波束形成算法。这些发现证实了该方法在小型声呐设备中的应用前景广阔,具有一定的应用价值。 展开更多
关键词 CZT波束形成 图形处理器 平面阵列 并行计算 算法优化
暂未订购 下载PDF
上一页 1 2 47 下一页 到第
在线咨询 使用帮助 返回顶部 意见反馈