Traditional malware detection models rely on a single feature source for detection,resulting in high false positive or false negative rates due to incomplete information.In addition,conventional models depend on manua...Traditional malware detection models rely on a single feature source for detection,resulting in high false positive or false negative rates due to incomplete information.In addition,conventional models depend on manual feature engineering,which is inefficient and hard to adapt to new malware variants.To address these challenges,this paper proposes a malware detection model called WAFDect based on a self-attention mechanism with multi-source feature fusion.The model consists of two key designs.First,we construct a multi-source feature extraction model that analyzes multi-source data such as API call sequences,registry operation logs,file operation logs,and network behavior logs,capturing malware characteristics at multiple abstraction levels and building a global representation of maliciousness,thereby overcoming the problem of single feature sources in traditional models.Second,to address the heterogeneity of multi-source features in terms of dimension,scale,and semantics,we design a feature alignment module based on attention weights.This module can dynamically learn the association strength between different feature modalities and achieve semantic alignment and adaptive fusion of cross-modal features through a weighting allocation mechanism,effectively reducing the reliance on manual feature engineering in traditional methods.The experimental results indicate that WAFDect achieved excellent detection performance on the Speakeasy(trainset)and Avast-CTU_Small datasets,with accuracies of 0.9229 and 0.9878,respectively.Compared with traditional detection models,this method shows significant improvements in key metrics such as accuracy and F1 score,thereby validating its effectiveness.展开更多
Accurate delineation of grape berry boundaries is essential for phenotypic measurement and growth assessment.This study proposes a multi-source feature fusion network(MFFNet)for instance segmentation of grape berries ...Accurate delineation of grape berry boundaries is essential for phenotypic measurement and growth assessment.This study proposes a multi-source feature fusion network(MFFNet)for instance segmentation of grape berries in dense clusters with frequent overlaps and blurred edges.MFFNet employs two parallel branches for feature extraction:a Swin Transformer backbone to capture hierarchical semantic features and an edge-detection branch that predicts an edge probability map to provide boundary cues.To address the substantial scale variation within a single image,the multilevel semantic features were enhanced using Adaptive Spatial Feature Fusion(ASFF).The edge probability map was introduced twice into the ASFFenhanced multi-scale features.First,edge cues were injected into the highest-resolution fused feature map to strengthen global boundary awareness across the cluster.Second,during mask generation,edge cues were reintroduced within each candidate instance region to refine local contours and improve the separation in the adhered areas.Experiments on a custom dataset collected in Yinchuan,Ningxia,showed that MFFNet achieved an mAP50boxof 93.4%and mAP50maskof 93.4%,outperforming representative baselines,including Mask2Former and HTC.The proposed model remained stable on images with severe berry overlap and indistinct edges,supporting practical grape growth monitoring.展开更多
Visible and infrared(RGB-IR)fusion object detection plays an important role in security,disaster relief,etc.In recent years,deep-learning-based RGB-IR fusion detection methods have been developing rapidly,but still st...Visible and infrared(RGB-IR)fusion object detection plays an important role in security,disaster relief,etc.In recent years,deep-learning-based RGB-IR fusion detection methods have been developing rapidly,but still struggle to deal with the complex and changing scenarios captured by drones,mainly due to two reasons:(A)RGB-IR fusion detectors are susceptible to inferior inputs that degrade performance and stability.(B)RGB-IR fusion detectors are susceptible to redundant features that reduce accuracy and efficiency.In this paper,an innovative RGB-IR fusion detection framework based on global-local feature optimization,named GLFDet,is proposed to improve the detection performance and efficiency of drone-captured objects.The key components of GLFDet include a Global Feature Optimization(GFO)module,a Local Feature Optimization(LFO)module and a Channel Separation Fusion(CSF)module.Specifically,GFO calculates the information content of the input image from the frequency domain and optimizes the features holistically.Then,LFO dynamically selects high-value features and filters out low-value features before fusion,which significantly improves the efficiency of fusion.Finally,CSF fuses the RGB and IR features across the corresponding channels,which avoids the rearrangement of the channel relationships and enhances the model stability.Extensive experimental results show that the proposed method achieves the best performance on three popular RGB-IR datasets Drone Vehicle,VEDAI,and LLVIP.In addition,GLFDet is more lightweight than other comparable models,making it more appealing to edge devices such as drones.The code is available at http://gffzz188fe103f8f1460asxf6pnkuuuvx966fk.ffgz.tsg.suse.edu.cn/lao chen330/GLFDet.展开更多
The spatial offset of bridge has a significant impact on the safety,comfort,and durability of high-speed railway(HSR)operations,so it is crucial to rapidly and effectively detect the spatial offset of operational HSR ...The spatial offset of bridge has a significant impact on the safety,comfort,and durability of high-speed railway(HSR)operations,so it is crucial to rapidly and effectively detect the spatial offset of operational HSR bridges.Drive-by monitoring of bridge uneven settlement demonstrates significant potential due to its practicality,cost-effectiveness,and efficiency.However,existing drive-by methods for detecting bridge offset have limitations such as reliance on a single data source,low detection accuracy,and the inability to identify lateral deformations of bridges.This paper proposes a novel drive-by inspection method for spatial offset of HSR bridge based on multi-source data fusion of comprehensive inspection train.Firstly,dung beetle optimizer-variational mode decomposition was employed to achieve adaptive decomposition of non-stationary dynamic signals,and explore the hidden temporal relationships in the data.Subsequently,a long short-term memory neural network was developed to achieve feature fusion of multi-source signal and accurate prediction of spatial settlement of HSR bridge.A dataset of track irregularities and CRH380A high-speed train responses was generated using a 3D train-track-bridge interaction model,and the accuracy and effectiveness of the proposed hybrid deep learning model were numerically validated.Finally,the reliability of the proposed drive-by inspection method was further validated by analyzing the actual measurement data obtained from comprehensive inspection train.The research findings indicate that the proposed approach enables rapid and accurate detection of spatial offset in HSR bridge,ensuring the long-term operational safety of HSR bridges.展开更多
Indoor intrusion detection is essential for various applications,including security systems and smart homes.Recently,WiFi-based detection has gained popularity due to its low cost and non-invasive nature.Current Chann...Indoor intrusion detection is essential for various applications,including security systems and smart homes.Recently,WiFi-based detection has gained popularity due to its low cost and non-invasive nature.Current Channel State Information(CSI)based frameworks primarily use deep learning to extract gait signatures;however,their performance depends heavily on extensive labeled datasets.These methods struggle to differentiate between unlabeled and labeled data that exhibit similar features.To address this challenge,we propose a novel Two-level Feature Fusion model for Indoor Intrusion Detection(TFF-IID)utilizing commercial WiFi CSI.The model adopts a two-level structure to learn rich feature representations and introduces a Transformer with multi-head self-attention alongside a multi-scale convolution module to process sensor data.Additionally,it incorporates a self-supervised learning module to capture general normality patterns.Based on this architecture,TFF-IID achieves accurate intrusion detection using only CSI.Empirical evaluations on a private gait dataset demonstrate that TFF-IID achieves an intrusion detection accuracy of 73.5%and an F1-score of 76.2%across 10 unauthorized subjects.Moreover,cross-scenario assessments verify that the proposed model maintains high efficiency and robustness in environments characterized by diverse spatial layouts and multipath complexities.Furthermore,TFF-IID outperforms the best baseline by 19.7%and 25.7%in accuracy and F1-score,respectively.展开更多
In recent years,with the rapid advancement of artificial intelligence,object detection algorithms have made significant strides in accuracy and computational efficiency.Notably,research and applications of Anchor-Free...In recent years,with the rapid advancement of artificial intelligence,object detection algorithms have made significant strides in accuracy and computational efficiency.Notably,research and applications of Anchor-Free models have opened new avenues for real-time target detection in optical remote sensing images(ORSIs).However,in the realmof adversarial attacks,developing adversarial techniques tailored to Anchor-Freemodels remains challenging.Adversarial examples generated based on Anchor-Based models often exhibit poor transferability to these new model architectures.Furthermore,the growing diversity of Anchor-Free models poses additional hurdles to achieving robust transferability of adversarial attacks.This study presents an improved cross-conv-block feature fusion You Only Look Once(YOLO)architecture,meticulously engineered to facilitate the extraction ofmore comprehensive semantic features during the backpropagation process.To address the asymmetry between densely distributed objects in ORSIs and the corresponding detector outputs,a novel dense bounding box attack strategy is proposed.This approach leverages dense target bounding boxes loss in the calculation of adversarial loss functions.Furthermore,by integrating translation-invariant(TI)and momentum-iteration(MI)adversarial methodologies,the proposed framework significantly improves the transferability of adversarial attacks.Experimental results demonstrate that our method achieves superior adversarial attack performance,with adversarial transferability rates(ATR)of 67.53%on the NWPU VHR-10 dataset and 90.71%on the HRSC2016 dataset.Compared to ensemble adversarial attack and cascaded adversarial attack approaches,our method generates adversarial examples in an average of 0.64 s,representing an approximately 14.5%improvement in efficiency under equivalent conditions.展开更多
Laser powder bed fusion is a key metal additive manufacturing technology capable of fabricating geometrically complex parts,yet its reliable industrial adoption is hindered by the inherent complexity and stochastic de...Laser powder bed fusion is a key metal additive manufacturing technology capable of fabricating geometrically complex parts,yet its reliable industrial adoption is hindered by the inherent complexity and stochastic defect formation of the process.Current quality assessment is constrained by the inherent latency of offline methods and the diagnostic limitations of single-sensor monitoring.To address these challenges,this study developed a multi-source optical signal monitoring system integrating coaxial photodiodes and an off-axis industrial camera to achieve simultaneous powder spreading detection and radiation signal monitoring during LPBF layer-wise process quality monitoring.Based on the successful identification and analysis of typical detectable features,the YOLOv5s deep learning model was employed to achieve rapid and accurate detection of lack-of-powder defects during the printing process.The training results indicated that the model exhibited good performance metrics.The relationships between process parameters,typical defects,and multi-channel monitoring data were also investigated.The monitoring system achieved a spatial resolution of 300μm for in-process monitoring and demonstrated high accuracy in detecting various defect types,including lack of powder,pores,warping,stitching seams,and printing failures.Furthermore,the algorithm-detected signal anomalies exhibited good spatial correlation with the actual surface defects.Simultaneously,wavelet time-frequency analysis was employed to evaluate molten pool dynamic stability under different process parameters and to analyze energy distribution for different defects.Furthermore,3D model reconstruction from signals enabled effective correlation with actual part defects.Based on the signal-driven process optimization,complex conformal cooling molds were successfully fabricated with a grafting accuracy error of less than 0.12 mm on high-performance substrates,demonstrating the practical efficacy of the developed monitoring methodology.This study provides both a technological and a theoretical foundation for intelligent quality control in LPBF and its practical implementation in industry.展开更多
Federated semi-supervised learning(FSSL)has garnered substantial attention for enabling collaborative global model training across multiple clients to address the scarcity of labeled data and to preserve data privacy....Federated semi-supervised learning(FSSL)has garnered substantial attention for enabling collaborative global model training across multiple clients to address the scarcity of labeled data and to preserve data privacy.However,FSSL is plagued by formidable challenges stemming fromcross-client data heterogeneity,as existing methods fail to achieve effective fusion of feature subspaces across distinct clients.To address this issue,we propose a novel FSSL framework,named FedSPQR,which is explicitly tailored for the label-at-server scenario.On the server side,FedSPQR adopts subspace clustering and fusion method based on the Grassmann manifold to construct a unified global feature space,which is further leveraged to refine the global model.On the client side,the pre-established global feature space acts as a benchmark for aligning the local feature subspaces.Based on the aligned local feature subspaces,integrating self-supervised learning with knowledge distillation facilitates effective local learning to alleviate local bias caused by data heterogeneity.Extensive experiments on two standard public benchmarks confirm that FedSPQR outperforms state-of-the-art(SOTA)baselines by a significant margin.展开更多
In recent years,data-driven approaches for online defect monitoring in metal laser additive manufacturing(LAM)have achieved remarkable progress.However,most existing studies primarily rely on spatial features extracte...In recent years,data-driven approaches for online defect monitoring in metal laser additive manufacturing(LAM)have achieved remarkable progress.However,most existing studies primarily rely on spatial features extracted from single-modal transient images,which are insufficient to capture the temporal evolution characteristics of the melt pool and the associated variations in local thermal history during the laser metal deposition(LMD)process.Moreover,the complementary information provided by multi-sensor data has often been overlooked.To address these limitations,this study proposes a multimodal feature-level spatiotemporal network(MFST-Net),which enables joint modeling and deep fusion of melt pool image sequences and in-situ process temperature signals.Specifically,a spatiotemporal feature fusion neural network(STFNN)is constructed to extract spatial distribution patterns from melt pool images while capturing multi-scale temporal dependencies at both the intra-layer and inter-layer levels.In parallel,a self-attention convolutional long short-term memory(SAConvLSTM)network is employed to model the dynamic evolution of thermal signals.Finally,cross-modal feature fusion is performed at the feature level to characterize the relationship between thermal–morphological evolution and pore formation mechanisms.Experimental results demonstrate the effectiveness and superiority of MFST-Net in online monitoring of local porosity,achieving an accuracy of 95.8%.These findings provide a promising reference for the integration of multimodal spatiotemporal feature fusion in complex manufacturing process monitoring.展开更多
Hydraulic presses are indispensable in automotive and aerospace manufacturing,with hydraulic cylinders serving as key components for operational safety and product quality.Internal leakage faults in hydraulic cylinder...Hydraulic presses are indispensable in automotive and aerospace manufacturing,with hydraulic cylinders serving as key components for operational safety and product quality.Internal leakage faults in hydraulic cylinders are difficult to diagnose due to the scarcity of labeled data,the complexity of fault mechanisms,and the limited representation capability of single-signal methods under variable operating conditions.To address these issues,a hybrid deep learning feature fusion model based on displacement error and pressure signal,including convolutional autoencoder,multi-head attention mechanism,residual network and bidirectional long short time series neural network(CAEMRAB),is proposed for the diagnosis and classification of leakage faults in hydraulic cylinders.A hydraulic cylinder test system simulates heavy load,variable speed,and nonlinear motion under actual operating conditions.Through the all-round deep feature decoupling of the proposed model,the multi-source signal representation ability in complex and multi-noise environments is enhanced,effectively extracting the local and global features of displacement error and pressure signal fault data and achieving efficient classification.Experimental results indicate that the proposed model achieves at least a 3.95%improvement in diagnostic accuracy compared with ablation models.In addition,it exhibits high diagnostic stability across other models,single-signal diagnosis,varying sample sizes,and complex noise conditions.These experiments fully validate the superior performance of the proposed method in terms of diagnostic accuracy,reliability,and robustness.展开更多
Cross-domain feature fusion offers an approach to weak target recognition in complex sea environments.This paper proposes a distance metric learning-based method for weak target classification.The method first extract...Cross-domain feature fusion offers an approach to weak target recognition in complex sea environments.This paper proposes a distance metric learning-based method for weak target classification.The method first extracts three timedomain features and three frequency-domain features from radar echo signals.Then,the features are partitioned and mapped to low-dimensional subspaces using linear projection matrices.The squared Euclidean distance is used as a metric function to measure the similarity between samples,and supervised optimization is performed by introducing information from similar and dissimilar sample pairs.Next,the projection matrices of each group are jointly updated iteratively using the gradient descent method to achieve supervised feature fusion.Finally,the fused feature is input into an ensemble one-class support vector machine(EOCSVM)for classification.Verified by IPIX measured data,the proposed method can effectively improve the separability of targets and sea clutter and improve the classification ability of sea clutter and weak targets under short-time observation.The proposed method enhances the features correlation from different domains through metric learning and EOCSVM,which can effectively alleviate the sample imbalance problem between sea clutter and targets.展开更多
The era of big data has profoundly transformed mechanics research,with data-driven approaches playing a vital role in modeling and optimization.This study focuses on tunnel boring machine(TBM),where the thrust-torque ...The era of big data has profoundly transformed mechanics research,with data-driven approaches playing a vital role in modeling and optimization.This study focuses on tunnel boring machine(TBM),where the thrust-torque ratio is a key determinant of their tunneling energy efficiency.However,due to the complexity of experiments and the testing requirements,obtaining sufficient high-quality data under varying geological conditions remains a major challenge in optimizing the tunneling energy efficiency of TBM.To address this,multi-cutter rotary cutting machine experiments and numerical simulations were conducted on 22 different rock types.Comprehensive datasets of normal and rolling forces were systematically collected.Using specific energy(SE)as the rock-breaking efficiency metric,we integrated physical and numerical data through a CatBoost-based fusion framework.The predictive model was initially trained on simulation data to capture the relationships among penetration,uniaxial compressive strength,tensile strength,and SE,and was subsequently fine-tuned with experimental data to develop the final fused model.Compared to models trained solely on experimental or simulated data,the fused model reduced RMSE by 37.1%and 58.6%,respectively,and improved R2by 19.0%and 44.6%,thereby enhancing both prediction accuracy and generalization capability.Furthermore,Bayesian optimization was employed to minimize SE and identify the optimal penetration.The results indicate that as rock strength increases,the optimal penetration decreases,while the corresponding minimal SE increases.These findings provide theoretical and engineering insights for improving TBM energy efficiency and parameter optimization,while establishing a robust data fusion framework for mechanical data analysis.展开更多
This research centers on structural health monitoring of bridges,a critical transportation infrastructure.Owing to the cumulative action of heavy vehicle loads,environmental variations,and material aging,bridge compon...This research centers on structural health monitoring of bridges,a critical transportation infrastructure.Owing to the cumulative action of heavy vehicle loads,environmental variations,and material aging,bridge components are prone to cracks and other defects,severely compromising structural safety and service life.Traditional inspection methods relying on manual visual assessment or vehicle-mounted sensors suffer from low efficiency,strong subjectivity,and high costs,while conventional image processing techniques and early deep learning models(e.g.,UNet,Faster R-CNN)still performinadequately in complex environments(e.g.,varying illumination,noise,false cracks)due to poor perception of fine cracks andmulti-scale features,limiting practical application.To address these challenges,this paper proposes CACNN-Net(CBAM-Augmented CNN),a novel dual-encoder architecture that innovatively couples a CNN for local detail extraction with a CBAM-Transformer for global context modeling.A key contribution is the dedicated Feature FusionModule(FFM),which strategically integratesmulti-scale features and focuses attention on crack regions while suppressing irrelevant noise.Experiments on bridge crack datasets demonstrate that CACNNNet achieves a precision of 77.6%,a recall of 79.4%,and an mIoU of 62.7%.These results significantly outperform several typical models(e.g.,UNet-ResNet34,Deeplabv3),confirming their superior accuracy and robust generalization,providing a high-precision automated solution for bridge crack detection and a novel network design paradigm for structural surface defect identification in complex scenarios,while future research may integrate physical features like depth information to advance intelligent infrastructure maintenance and digital twin management.展开更多
With the rapid development of Industrial 4.0 and Industrial Internet of Things,the data collection with multisource has significantly improved.How to effectively fuse these data for various engineering applications is...With the rapid development of Industrial 4.0 and Industrial Internet of Things,the data collection with multisource has significantly improved.How to effectively fuse these data for various engineering applications is still an open and challenge issue.To this end,we propose the canonical correlation guided deep neural network(CCDNN),a novel deep learning architecture,to learn a correlated representation for multi-source data fusion.Unlike the linear canonical correlation analysis(CCA),kernel CCA and deep CCA,in the proposed method,the optimization formulation is not restricted to maximize correlation,instead we make canonical correlation as a constraint,which preserves the correlated representation learning ability and focuses more on the engineering tasks endowed by optimization formulation,such as reconstruction,classification and prediction.Furthermore,to reduce the redundancy induced by correlation,a redundancy filter is designed.We illustrate its data fusion ability via correlated representation learning and superior performance on various engineering tasks.In experiments on MNIST dataset,the results show that CCDNN has better reconstruction performance in terms of mean squared error and mean absolute error than deep CCA and deep canonically correlated autoencoders(DCCAE).Also,we present the application of the proposed network to industrial fault diagnosis and remaining useful life cases for the classification and prediction tasks accordingly.The proposed method demonstrates approving performance in both tasks when compared to existing methods.Extension of CCDNN to much more deeper with the aid of residual connection is also presented in Appendix.展开更多
As the core propulsion system of supersonic vehicles,the scramjet engine experiences unstable combustion phenomena in the combustor under high-speed operating conditions,which can lead to performance degradation and s...As the core propulsion system of supersonic vehicles,the scramjet engine experiences unstable combustion phenomena in the combustor under high-speed operating conditions,which can lead to performance degradation and structural damage.Therefore,the development of supersonic flame stabilization structure identification technology is urgently needed.A Heterogeneous Feature Fusion Module(HFFM)is proposed to achieve nonlinear and organic fusion of heterogeneous data.Flame Structure Data(FSD)characterize key flame features,while Combustor Wall Pressure Data(CWPD)supplement the missing flame structure features in FSD,generating Flame Heterogeneous Feature Fusion Data(FHFFD).Additionally,a Supersonic Flame Stabilization Identification Module(SFSIM)is proposed,which combines a horn-shaped convolutional neural network with a Simplified Low Latent Transformer(SLLT)to enable dynamic and adaptive multi-scale integration of flame stabilization structure features.Experimental results indicate that HFFM effectively extracts and consolidates stable flame structure features within FHFFD during the training phase,demonstrating the ability to generalize key physical principles from FHFFD.SFSIM achieves a recognition accuracy of 97.09%through parameter optimization and attention-based dimensionality reduction.The low latent space improves training efficiency by 10.3%,while its parameter count accounts for only 29.73%.While maintaining high accuracy,this approach provides efficient and robust technical support for real-time monitoring of supersonic combustion.展开更多
Colorectal cancer(CRC)is a prevalent disease,with polyps serving as its precursors.Accurate polyp segmentation is crucial for early CRC prevention.However,due to different sizes of the polyps,the boundaries are not cl...Colorectal cancer(CRC)is a prevalent disease,with polyps serving as its precursors.Accurate polyp segmentation is crucial for early CRC prevention.However,due to different sizes of the polyps,the boundaries are not clear.Therefore,accurate segmentation of polyps is a challenging task.This paper proposes vision Mamba attention feature fusion UNet(VMA-UNet),a U-shaped asymmetric codec structure model grounded in the state space model(SSM).The VMA-UNet incorporates attention feature fusion(AFF)in order to enhance the feature representation of small polyps.A new IUD loss function,namely combining intersection over union(IoU)loss function and Dice loss function,is proposed to address both large polyps and small polyps,and to mitigate the issue of data imbalance.When applied to multiple datasets,VMA-UNet demonstrates robust performance,particularly in small polyp segmentation,showcasing its practical value.The network proposed in this paper overcomes the inherent shortcomings of convolutional neural network(CNN)and transformers,not only performing well in remote interaction modeling,but also maintaining linear computational complexity.Our study introduces a new method for polyp segmentation based on SSM and advances the field.展开更多
Medical image segmentation is an essential method for computer-aided diagnosis.Although image segmentation models based on convolutional neural networks(CNNs)and vision transformers(ViTs)have achieved significant adva...Medical image segmentation is an essential method for computer-aided diagnosis.Although image segmentation models based on convolutional neural networks(CNNs)and vision transformers(ViTs)have achieved significant advancements,CNNs struggle to effectively capture long-range dependencies,while ViTs face limitations in local information extraction and are hindered by quadratic computational complexity.Recently,their inherent issues have been successfully addressed by the state-space models in Mamba and 2D-selective-scan in Vision Mamba.However,the presence of noise and excessive redundant information in medical images limits the practicality of these methods.To address these challenges,we propose a highly effective and accurate high-low-order feature fusion visual state space module,named HL-VSS.This module primarily consists of two core components:multi-scale spatial convolution and highlow-order feature fusion(HLFF).The former component preliminarily suppresses noise and captures multi-scale feature information from medical images,accurately extracting edge and detail features for the fusion component.The latter processes these features,further reducing redundant information through high-order interaction with 2D-selective-scan,and fuses the local features obtained by low-order parallel Mamba,ultimately extracting deeper medical image features.We incorporate HLVSS into a U-shaped architecture,named high-low-order feature fusion visual Mamba UNet(V-UNet).Comparison experiments and ablation studies are conducted on four publicly available medical image datasets to validate the strong competitiveness of V-UNet in medical image segmentation tasks.The code is available at http://gffzz188fe103f8f1460asxf6pnkuuuvx966fk.ffgz.tsg.suse.edu.cn/ai-dqh0106/V-UNet Code.展开更多
To improve the accuracy of small object feature detection in complex backgrounds for Unmanned Aerial Vehicle(UAV)aerial photography and reduce computational complexity,we propose the lightweight UAV aerial photography...To improve the accuracy of small object feature detection in complex backgrounds for Unmanned Aerial Vehicle(UAV)aerial photography and reduce computational complexity,we propose the lightweight UAV aerial photography small object detection method based on multi-scale feature fusion and contextual information.Firstly,by introducing the grouped content-aware reassembly(GCA)operator and designing lightweight pinwheel context convolution(LPConv),we extend the feature fusion path to the P2 layer,constructing a lightweight multi-scale feature fusion network(SG-PANet).Through the decoupling of fine-grained small object features and background interference features by the GCA operator,combined with the anisotropic receptive field constructed by LPConv,our proposed method can effectively preserve the geometric details of small objects.Furthermore,we introduce the cross-stage dense feature refinement(CSPStage)module as the pre-refining unit of the detection head,and use the full history state awareness mechanism to strengthen feature reuse and gradient propagation to solve the problem of feature degradation across layers.We utilize the Wise-IoU v3 loss function to dynamically optimize the gradient gains of high-quality and low-quality samples,thereby enhancing the detection accuracy and convergence speed of the proposed method in complex scenarios.Finally,we verified the superiority and generalization of the proposed method on the VisDrone2019 dataset and DOTAv1.5 dataset.The results show that compared with YOLOv11n,MFCI-YOLO’s detection mAP50-95 increased by 11.1%,small object mAP50 increased by 16.1%,and mAP50 reached 80.3%.It provides a practical solution for detecting small objects in dense scenes.展开更多
Camouflaged Object Detection(COD)aims to identify objects that share highly similar patterns—such as texture,intensity,and color—with their surrounding environment.Due to their intrinsic resemblance to the backgroun...Camouflaged Object Detection(COD)aims to identify objects that share highly similar patterns—such as texture,intensity,and color—with their surrounding environment.Due to their intrinsic resemblance to the background,camouflaged objects often exhibit vague boundaries and varying scales,making it challenging to accurately locate targets and delineate their indistinct edges.To address this,we propose a novel camouflaged object detection network called Edge-Guided and Multi-scale Fusion Network(EGMFNet),which leverages edge-guided multi-scale integration for enhanced performance.The model incorporates two innovative components:a Multi-scale Fusion Module(MSFM)and an Edge-Guided Attention Module(EGA).These designs exploit multi-scale features to uncover subtle cues between candidate objects and the background while emphasizing camouflaged object boundaries.Moreover,recognizing the rich contextual information in fused features,we introduce a Dual-Branch Global Context Module(DGCM)to refine features using extensive global context,thereby generatingmore informative representations.Experimental results on four benchmark datasets demonstrate that EGMFNet outperforms state-of-the-art methods across five evaluation metrics.Specifically,on COD10K,our EGMFNet-P improves Fβby 4.8 points and reduces mean absolute error(MAE)by 0.006 compared with ZoomNeXt;on NC4K,it achieves a 3.6-point increase in Fβ.OnCAMO and CHAMELEON,it obtains 4.5-point increases in Fβ,respectively.These consistent gains substantiate the superiority and robustness of EGMFNet.展开更多
In recent years,with the advancement of computational hardware performance,machine learning algorithms have achieved significant development and widespread application across various fields,and have become deeply embe...In recent years,with the advancement of computational hardware performance,machine learning algorithms have achieved significant development and widespread application across various fields,and have become deeply embedded in smart grids and communication systems.However,it is important to note that despite the widespread deployment of smart meters in the power system,the lack of reliable intelligent diagnostic,a large number of such electricity meters experiencing communication failures caused by internal topological defects every year.To address this issue,we propose a machine learning-based monitoring and early warning model using multidimensional feature fusion.By integrating more than twenty key features in four categories,including attribute features,operational load features,communication behavior features,and derived combined features-an XGBoost classification algorithm framework is constructed to implement risk early warning for electricity meter communication faults.Validated with data from millions of users,the proposed model achieves an accuracy of approximately 90%,the annual average reduction in power outages caused by communication faults is more than 10,000 hours,and significantly enhances the grid’s safety and operational stability.展开更多
基金funded by Name of the National Nature Science Foundation of China,grant number 62262004.
摘要Traditional malware detection models rely on a single feature source for detection,resulting in high false positive or false negative rates due to incomplete information.In addition,conventional models depend on manual feature engineering,which is inefficient and hard to adapt to new malware variants.To address these challenges,this paper proposes a malware detection model called WAFDect based on a self-attention mechanism with multi-source feature fusion.The model consists of two key designs.First,we construct a multi-source feature extraction model that analyzes multi-source data such as API call sequences,registry operation logs,file operation logs,and network behavior logs,capturing malware characteristics at multiple abstraction levels and building a global representation of maliciousness,thereby overcoming the problem of single feature sources in traditional models.Second,to address the heterogeneity of multi-source features in terms of dimension,scale,and semantics,we design a feature alignment module based on attention weights.This module can dynamically learn the association strength between different feature modalities and achieve semantic alignment and adaptive fusion of cross-modal features through a weighting allocation mechanism,effectively reducing the reliance on manual feature engineering in traditional methods.The experimental results indicate that WAFDect achieved excellent detection performance on the Speakeasy(trainset)and Avast-CTU_Small datasets,with accuracies of 0.9229 and 0.9878,respectively.Compared with traditional detection models,this method shows significant improvements in key metrics such as accuracy and F1 score,thereby validating its effectiveness.
基金funded by the Fengyun Satellite Application Pioneer Program(PhaseⅢ)(Grant No.FY-APP-2024.0301)the Science Foundation of Shandong(Grant No.ZR2021MD097)+1 种基金the Natural Science Foundation of Ningxia(Grant No.2024AAC03419)Shandong Provincial Meteorological Bureau Innovation Team Special Project(Grant No.2024sdcxtd04).
摘要Accurate delineation of grape berry boundaries is essential for phenotypic measurement and growth assessment.This study proposes a multi-source feature fusion network(MFFNet)for instance segmentation of grape berries in dense clusters with frequent overlaps and blurred edges.MFFNet employs two parallel branches for feature extraction:a Swin Transformer backbone to capture hierarchical semantic features and an edge-detection branch that predicts an edge probability map to provide boundary cues.To address the substantial scale variation within a single image,the multilevel semantic features were enhanced using Adaptive Spatial Feature Fusion(ASFF).The edge probability map was introduced twice into the ASFFenhanced multi-scale features.First,edge cues were injected into the highest-resolution fused feature map to strengthen global boundary awareness across the cluster.Second,during mask generation,edge cues were reintroduced within each candidate instance region to refine local contours and improve the separation in the adhered areas.Experiments on a custom dataset collected in Yinchuan,Ningxia,showed that MFFNet achieved an mAP50boxof 93.4%and mAP50maskof 93.4%,outperforming representative baselines,including Mask2Former and HTC.The proposed model remained stable on images with severe berry overlap and indistinct edges,supporting practical grape growth monitoring.
基金supported by the National Natural Science Foundation of China(No.62276204)the Fundamental Research Funds for the Central Universities,China(No.YJSJ24011)+1 种基金the Natural Science Basic Research Program of Shaanxi,China(Nos.2022JM-340 and 2023-JC-QN-0710)the China Postdoctoral Science Foundation(Nos.2020T130494 and 2018M633470)。
摘要Visible and infrared(RGB-IR)fusion object detection plays an important role in security,disaster relief,etc.In recent years,deep-learning-based RGB-IR fusion detection methods have been developing rapidly,but still struggle to deal with the complex and changing scenarios captured by drones,mainly due to two reasons:(A)RGB-IR fusion detectors are susceptible to inferior inputs that degrade performance and stability.(B)RGB-IR fusion detectors are susceptible to redundant features that reduce accuracy and efficiency.In this paper,an innovative RGB-IR fusion detection framework based on global-local feature optimization,named GLFDet,is proposed to improve the detection performance and efficiency of drone-captured objects.The key components of GLFDet include a Global Feature Optimization(GFO)module,a Local Feature Optimization(LFO)module and a Channel Separation Fusion(CSF)module.Specifically,GFO calculates the information content of the input image from the frequency domain and optimizes the features holistically.Then,LFO dynamically selects high-value features and filters out low-value features before fusion,which significantly improves the efficiency of fusion.Finally,CSF fuses the RGB and IR features across the corresponding channels,which avoids the rearrangement of the channel relationships and enhances the model stability.Extensive experimental results show that the proposed method achieves the best performance on three popular RGB-IR datasets Drone Vehicle,VEDAI,and LLVIP.In addition,GLFDet is more lightweight than other comparable models,making it more appealing to edge devices such as drones.The code is available at http://gffzz188fe103f8f1460asxf6pnkuuuvx966fk.ffgz.tsg.suse.edu.cn/lao chen330/GLFDet.
基金sponsored by the National Natural Science Foundation of China(Grant No.52178100).
摘要The spatial offset of bridge has a significant impact on the safety,comfort,and durability of high-speed railway(HSR)operations,so it is crucial to rapidly and effectively detect the spatial offset of operational HSR bridges.Drive-by monitoring of bridge uneven settlement demonstrates significant potential due to its practicality,cost-effectiveness,and efficiency.However,existing drive-by methods for detecting bridge offset have limitations such as reliance on a single data source,low detection accuracy,and the inability to identify lateral deformations of bridges.This paper proposes a novel drive-by inspection method for spatial offset of HSR bridge based on multi-source data fusion of comprehensive inspection train.Firstly,dung beetle optimizer-variational mode decomposition was employed to achieve adaptive decomposition of non-stationary dynamic signals,and explore the hidden temporal relationships in the data.Subsequently,a long short-term memory neural network was developed to achieve feature fusion of multi-source signal and accurate prediction of spatial settlement of HSR bridge.A dataset of track irregularities and CRH380A high-speed train responses was generated using a 3D train-track-bridge interaction model,and the accuracy and effectiveness of the proposed hybrid deep learning model were numerically validated.Finally,the reliability of the proposed drive-by inspection method was further validated by analyzing the actual measurement data obtained from comprehensive inspection train.The research findings indicate that the proposed approach enables rapid and accurate detection of spatial offset in HSR bridge,ensuring the long-term operational safety of HSR bridges.
基金supported by the Shaanxi Province Outstanding Youth Science Foundation Project(2025JC-JCQN-074).
摘要Indoor intrusion detection is essential for various applications,including security systems and smart homes.Recently,WiFi-based detection has gained popularity due to its low cost and non-invasive nature.Current Channel State Information(CSI)based frameworks primarily use deep learning to extract gait signatures;however,their performance depends heavily on extensive labeled datasets.These methods struggle to differentiate between unlabeled and labeled data that exhibit similar features.To address this challenge,we propose a novel Two-level Feature Fusion model for Indoor Intrusion Detection(TFF-IID)utilizing commercial WiFi CSI.The model adopts a two-level structure to learn rich feature representations and introduces a Transformer with multi-head self-attention alongside a multi-scale convolution module to process sensor data.Additionally,it incorporates a self-supervised learning module to capture general normality patterns.Based on this architecture,TFF-IID achieves accurate intrusion detection using only CSI.Empirical evaluations on a private gait dataset demonstrate that TFF-IID achieves an intrusion detection accuracy of 73.5%and an F1-score of 76.2%across 10 unauthorized subjects.Moreover,cross-scenario assessments verify that the proposed model maintains high efficiency and robustness in environments characterized by diverse spatial layouts and multipath complexities.Furthermore,TFF-IID outperforms the best baseline by 19.7%and 25.7%in accuracy and F1-score,respectively.
摘要In recent years,with the rapid advancement of artificial intelligence,object detection algorithms have made significant strides in accuracy and computational efficiency.Notably,research and applications of Anchor-Free models have opened new avenues for real-time target detection in optical remote sensing images(ORSIs).However,in the realmof adversarial attacks,developing adversarial techniques tailored to Anchor-Freemodels remains challenging.Adversarial examples generated based on Anchor-Based models often exhibit poor transferability to these new model architectures.Furthermore,the growing diversity of Anchor-Free models poses additional hurdles to achieving robust transferability of adversarial attacks.This study presents an improved cross-conv-block feature fusion You Only Look Once(YOLO)architecture,meticulously engineered to facilitate the extraction ofmore comprehensive semantic features during the backpropagation process.To address the asymmetry between densely distributed objects in ORSIs and the corresponding detector outputs,a novel dense bounding box attack strategy is proposed.This approach leverages dense target bounding boxes loss in the calculation of adversarial loss functions.Furthermore,by integrating translation-invariant(TI)and momentum-iteration(MI)adversarial methodologies,the proposed framework significantly improves the transferability of adversarial attacks.Experimental results demonstrate that our method achieves superior adversarial attack performance,with adversarial transferability rates(ATR)of 67.53%on the NWPU VHR-10 dataset and 90.71%on the HRSC2016 dataset.Compared to ensemble adversarial attack and cascaded adversarial attack approaches,our method generates adversarial examples in an average of 0.64 s,representing an approximately 14.5%improvement in efficiency under equivalent conditions.
基金supported by National Natural Science Foundation of China(Grant No.52475349)Guangdong Basic and Applied Basic Research Foundation(Grant No.2022B1515020064)National Key R&D program of China(Grant No.2022YFF0606000).
摘要Laser powder bed fusion is a key metal additive manufacturing technology capable of fabricating geometrically complex parts,yet its reliable industrial adoption is hindered by the inherent complexity and stochastic defect formation of the process.Current quality assessment is constrained by the inherent latency of offline methods and the diagnostic limitations of single-sensor monitoring.To address these challenges,this study developed a multi-source optical signal monitoring system integrating coaxial photodiodes and an off-axis industrial camera to achieve simultaneous powder spreading detection and radiation signal monitoring during LPBF layer-wise process quality monitoring.Based on the successful identification and analysis of typical detectable features,the YOLOv5s deep learning model was employed to achieve rapid and accurate detection of lack-of-powder defects during the printing process.The training results indicated that the model exhibited good performance metrics.The relationships between process parameters,typical defects,and multi-channel monitoring data were also investigated.The monitoring system achieved a spatial resolution of 300μm for in-process monitoring and demonstrated high accuracy in detecting various defect types,including lack of powder,pores,warping,stitching seams,and printing failures.Furthermore,the algorithm-detected signal anomalies exhibited good spatial correlation with the actual surface defects.Simultaneously,wavelet time-frequency analysis was employed to evaluate molten pool dynamic stability under different process parameters and to analyze energy distribution for different defects.Furthermore,3D model reconstruction from signals enabled effective correlation with actual part defects.Based on the signal-driven process optimization,complex conformal cooling molds were successfully fabricated with a grafting accuracy error of less than 0.12 mm on high-performance substrates,demonstrating the practical efficacy of the developed monitoring methodology.This study provides both a technological and a theoretical foundation for intelligent quality control in LPBF and its practical implementation in industry.
基金supported by the Scientific Research Foundation of CUIT(No.KYTZ2022108)Sichuan Science and Technology Program(No.2025ZNSFSC0494,No.2024NSFJQ0030).
摘要Federated semi-supervised learning(FSSL)has garnered substantial attention for enabling collaborative global model training across multiple clients to address the scarcity of labeled data and to preserve data privacy.However,FSSL is plagued by formidable challenges stemming fromcross-client data heterogeneity,as existing methods fail to achieve effective fusion of feature subspaces across distinct clients.To address this issue,we propose a novel FSSL framework,named FedSPQR,which is explicitly tailored for the label-at-server scenario.On the server side,FedSPQR adopts subspace clustering and fusion method based on the Grassmann manifold to construct a unified global feature space,which is further leveraged to refine the global model.On the client side,the pre-established global feature space acts as a benchmark for aligning the local feature subspaces.Based on the aligned local feature subspaces,integrating self-supervised learning with knowledge distillation facilitates effective local learning to alleviate local bias caused by data heterogeneity.Extensive experiments on two standard public benchmarks confirm that FedSPQR outperforms state-of-the-art(SOTA)baselines by a significant margin.
基金supported by Ministry of Industry and Information Technology of the People's Republic of China(Grant No.2540STC62584)Science and Technology Department of Sichuan Province(Grant No.2025ZDZX0050)National Natural Science Foundation of China(Grant No.52075352).
摘要In recent years,data-driven approaches for online defect monitoring in metal laser additive manufacturing(LAM)have achieved remarkable progress.However,most existing studies primarily rely on spatial features extracted from single-modal transient images,which are insufficient to capture the temporal evolution characteristics of the melt pool and the associated variations in local thermal history during the laser metal deposition(LMD)process.Moreover,the complementary information provided by multi-sensor data has often been overlooked.To address these limitations,this study proposes a multimodal feature-level spatiotemporal network(MFST-Net),which enables joint modeling and deep fusion of melt pool image sequences and in-situ process temperature signals.Specifically,a spatiotemporal feature fusion neural network(STFNN)is constructed to extract spatial distribution patterns from melt pool images while capturing multi-scale temporal dependencies at both the intra-layer and inter-layer levels.In parallel,a self-attention convolutional long short-term memory(SAConvLSTM)network is employed to model the dynamic evolution of thermal signals.Finally,cross-modal feature fusion is performed at the feature level to characterize the relationship between thermal–morphological evolution and pore formation mechanisms.Experimental results demonstrate the effectiveness and superiority of MFST-Net in online monitoring of local porosity,achieving an accuracy of 95.8%.These findings provide a promising reference for the integration of multimodal spatiotemporal feature fusion in complex manufacturing process monitoring.
基金the Scientific Research Foundation for High-level Talents of Anhui University of Science and Technology(2024yjrc73)R&D and industrialization of high-precision intelligent forging equipment for forming large-size light alloy components(202423i08050024)a large die forging press operation condition monitoring sensor and system application(2023YFB3210805)。
摘要Hydraulic presses are indispensable in automotive and aerospace manufacturing,with hydraulic cylinders serving as key components for operational safety and product quality.Internal leakage faults in hydraulic cylinders are difficult to diagnose due to the scarcity of labeled data,the complexity of fault mechanisms,and the limited representation capability of single-signal methods under variable operating conditions.To address these issues,a hybrid deep learning feature fusion model based on displacement error and pressure signal,including convolutional autoencoder,multi-head attention mechanism,residual network and bidirectional long short time series neural network(CAEMRAB),is proposed for the diagnosis and classification of leakage faults in hydraulic cylinders.A hydraulic cylinder test system simulates heavy load,variable speed,and nonlinear motion under actual operating conditions.Through the all-round deep feature decoupling of the proposed model,the multi-source signal representation ability in complex and multi-noise environments is enhanced,effectively extracting the local and global features of displacement error and pressure signal fault data and achieving efficient classification.Experimental results indicate that the proposed model achieves at least a 3.95%improvement in diagnostic accuracy compared with ablation models.In addition,it exhibits high diagnostic stability across other models,single-signal diagnosis,varying sample sizes,and complex noise conditions.These experiments fully validate the superior performance of the proposed method in terms of diagnostic accuracy,reliability,and robustness.
基金supported by the National Natural Science Foundation of China(62201251)the Open Fund for the Hangzhou Institute of Technology Academician Workstation at Xidian University(XH-KY-202306-0291)。
摘要Cross-domain feature fusion offers an approach to weak target recognition in complex sea environments.This paper proposes a distance metric learning-based method for weak target classification.The method first extracts three timedomain features and three frequency-domain features from radar echo signals.Then,the features are partitioned and mapped to low-dimensional subspaces using linear projection matrices.The squared Euclidean distance is used as a metric function to measure the similarity between samples,and supervised optimization is performed by introducing information from similar and dissimilar sample pairs.Next,the projection matrices of each group are jointly updated iteratively using the gradient descent method to achieve supervised feature fusion.Finally,the fused feature is input into an ensemble one-class support vector machine(EOCSVM)for classification.Verified by IPIX measured data,the proposed method can effectively improve the separability of targets and sea clutter and improve the classification ability of sea clutter and weak targets under short-time observation.The proposed method enhances the features correlation from different domains through metric learning and EOCSVM,which can effectively alleviate the sample imbalance problem between sea clutter and targets.
基金National Natural Science Foundation of China,12021002,Qian Zhang,12372186,QianZhang,Emerging Frontiers Cultivation Program of Tianjin University Interdisciplinary Center.
摘要The era of big data has profoundly transformed mechanics research,with data-driven approaches playing a vital role in modeling and optimization.This study focuses on tunnel boring machine(TBM),where the thrust-torque ratio is a key determinant of their tunneling energy efficiency.However,due to the complexity of experiments and the testing requirements,obtaining sufficient high-quality data under varying geological conditions remains a major challenge in optimizing the tunneling energy efficiency of TBM.To address this,multi-cutter rotary cutting machine experiments and numerical simulations were conducted on 22 different rock types.Comprehensive datasets of normal and rolling forces were systematically collected.Using specific energy(SE)as the rock-breaking efficiency metric,we integrated physical and numerical data through a CatBoost-based fusion framework.The predictive model was initially trained on simulation data to capture the relationships among penetration,uniaxial compressive strength,tensile strength,and SE,and was subsequently fine-tuned with experimental data to develop the final fused model.Compared to models trained solely on experimental or simulated data,the fused model reduced RMSE by 37.1%and 58.6%,respectively,and improved R2by 19.0%and 44.6%,thereby enhancing both prediction accuracy and generalization capability.Furthermore,Bayesian optimization was employed to minimize SE and identify the optimal penetration.The results indicate that as rock strength increases,the optimal penetration decreases,while the corresponding minimal SE increases.These findings provide theoretical and engineering insights for improving TBM energy efficiency and parameter optimization,while establishing a robust data fusion framework for mechanical data analysis.
基金supported by the National Natural Science Foundation of China(No.52308332)the General Scientific Research Project of the Education Department of Zhejiang Province(No.Y202455824).
摘要This research centers on structural health monitoring of bridges,a critical transportation infrastructure.Owing to the cumulative action of heavy vehicle loads,environmental variations,and material aging,bridge components are prone to cracks and other defects,severely compromising structural safety and service life.Traditional inspection methods relying on manual visual assessment or vehicle-mounted sensors suffer from low efficiency,strong subjectivity,and high costs,while conventional image processing techniques and early deep learning models(e.g.,UNet,Faster R-CNN)still performinadequately in complex environments(e.g.,varying illumination,noise,false cracks)due to poor perception of fine cracks andmulti-scale features,limiting practical application.To address these challenges,this paper proposes CACNN-Net(CBAM-Augmented CNN),a novel dual-encoder architecture that innovatively couples a CNN for local detail extraction with a CBAM-Transformer for global context modeling.A key contribution is the dedicated Feature FusionModule(FFM),which strategically integratesmulti-scale features and focuses attention on crack regions while suppressing irrelevant noise.Experiments on bridge crack datasets demonstrate that CACNNNet achieves a precision of 77.6%,a recall of 79.4%,and an mIoU of 62.7%.These results significantly outperform several typical models(e.g.,UNet-ResNet34,Deeplabv3),confirming their superior accuracy and robust generalization,providing a high-precision automated solution for bridge crack detection and a novel network design paradigm for structural surface defect identification in complex scenarios,while future research may integrate physical features like depth information to advance intelligent infrastructure maintenance and digital twin management.
基金supported in part by the National Natural Science Foundation of China(62173349)the Natural Science Foundation of Hunan Province(2025JJ10007)+1 种基金the Natural Science Foundation of Hunan Province(2022JJ20076)the Science and Technology Innovation Program of Hunan Province(2022RC1090)。
摘要With the rapid development of Industrial 4.0 and Industrial Internet of Things,the data collection with multisource has significantly improved.How to effectively fuse these data for various engineering applications is still an open and challenge issue.To this end,we propose the canonical correlation guided deep neural network(CCDNN),a novel deep learning architecture,to learn a correlated representation for multi-source data fusion.Unlike the linear canonical correlation analysis(CCA),kernel CCA and deep CCA,in the proposed method,the optimization formulation is not restricted to maximize correlation,instead we make canonical correlation as a constraint,which preserves the correlated representation learning ability and focuses more on the engineering tasks endowed by optimization formulation,such as reconstruction,classification and prediction.Furthermore,to reduce the redundancy induced by correlation,a redundancy filter is designed.We illustrate its data fusion ability via correlated representation learning and superior performance on various engineering tasks.In experiments on MNIST dataset,the results show that CCDNN has better reconstruction performance in terms of mean squared error and mean absolute error than deep CCA and deep canonically correlated autoencoders(DCCAE).Also,we present the application of the proposed network to industrial fault diagnosis and remaining useful life cases for the classification and prediction tasks accordingly.The proposed method demonstrates approving performance in both tasks when compared to existing methods.Extension of CCDNN to much more deeper with the aid of residual connection is also presented in Appendix.
基金supported by Outstanding Star Program of Zhiqiang Fund,China(Ye TIAN)。
摘要As the core propulsion system of supersonic vehicles,the scramjet engine experiences unstable combustion phenomena in the combustor under high-speed operating conditions,which can lead to performance degradation and structural damage.Therefore,the development of supersonic flame stabilization structure identification technology is urgently needed.A Heterogeneous Feature Fusion Module(HFFM)is proposed to achieve nonlinear and organic fusion of heterogeneous data.Flame Structure Data(FSD)characterize key flame features,while Combustor Wall Pressure Data(CWPD)supplement the missing flame structure features in FSD,generating Flame Heterogeneous Feature Fusion Data(FHFFD).Additionally,a Supersonic Flame Stabilization Identification Module(SFSIM)is proposed,which combines a horn-shaped convolutional neural network with a Simplified Low Latent Transformer(SLLT)to enable dynamic and adaptive multi-scale integration of flame stabilization structure features.Experimental results indicate that HFFM effectively extracts and consolidates stable flame structure features within FHFFD during the training phase,demonstrating the ability to generalize key physical principles from FHFFD.SFSIM achieves a recognition accuracy of 97.09%through parameter optimization and attention-based dimensionality reduction.The low latent space improves training efficiency by 10.3%,while its parameter count accounts for only 29.73%.While maintaining high accuracy,this approach provides efficient and robust technical support for real-time monitoring of supersonic combustion.
基金supported by the Natural Science Research Project of Tianjin Education Commission(No.2020KJ124)the National Natural Science Foundation of China(No.11601372)the National Key Research and Development Program of China(No.2022YFF0706003)。
摘要Colorectal cancer(CRC)is a prevalent disease,with polyps serving as its precursors.Accurate polyp segmentation is crucial for early CRC prevention.However,due to different sizes of the polyps,the boundaries are not clear.Therefore,accurate segmentation of polyps is a challenging task.This paper proposes vision Mamba attention feature fusion UNet(VMA-UNet),a U-shaped asymmetric codec structure model grounded in the state space model(SSM).The VMA-UNet incorporates attention feature fusion(AFF)in order to enhance the feature representation of small polyps.A new IUD loss function,namely combining intersection over union(IoU)loss function and Dice loss function,is proposed to address both large polyps and small polyps,and to mitigate the issue of data imbalance.When applied to multiple datasets,VMA-UNet demonstrates robust performance,particularly in small polyp segmentation,showcasing its practical value.The network proposed in this paper overcomes the inherent shortcomings of convolutional neural network(CNN)and transformers,not only performing well in remote interaction modeling,but also maintaining linear computational complexity.Our study introduces a new method for polyp segmentation based on SSM and advances the field.
基金partially supported by the Japan Society for the Promotion of Science(JSPS)KAKENHI(JP25K21298,JP25K03179)Japan Science and Technology Agency(JST)Support for Pioneering Research Initiated by the Next Generation(SPRING)(JPMJSP2145)。
摘要Medical image segmentation is an essential method for computer-aided diagnosis.Although image segmentation models based on convolutional neural networks(CNNs)and vision transformers(ViTs)have achieved significant advancements,CNNs struggle to effectively capture long-range dependencies,while ViTs face limitations in local information extraction and are hindered by quadratic computational complexity.Recently,their inherent issues have been successfully addressed by the state-space models in Mamba and 2D-selective-scan in Vision Mamba.However,the presence of noise and excessive redundant information in medical images limits the practicality of these methods.To address these challenges,we propose a highly effective and accurate high-low-order feature fusion visual state space module,named HL-VSS.This module primarily consists of two core components:multi-scale spatial convolution and highlow-order feature fusion(HLFF).The former component preliminarily suppresses noise and captures multi-scale feature information from medical images,accurately extracting edge and detail features for the fusion component.The latter processes these features,further reducing redundant information through high-order interaction with 2D-selective-scan,and fuses the local features obtained by low-order parallel Mamba,ultimately extracting deeper medical image features.We incorporate HLVSS into a U-shaped architecture,named high-low-order feature fusion visual Mamba UNet(V-UNet).Comparison experiments and ablation studies are conducted on four publicly available medical image datasets to validate the strong competitiveness of V-UNet in medical image segmentation tasks.The code is available at http://gffzz188fe103f8f1460asxf6pnkuuuvx966fk.ffgz.tsg.suse.edu.cn/ai-dqh0106/V-UNet Code.
基金supported in part by the Natural Science Foundation of Henan Province under Grant 252300423317the Science and Technology Research Project of Henan Province under Grant 262102211081+1 种基金the Key Scientific Research Projects of Colleges and Universities in Henan Province under Grant 25B510012“Pioneer”and“Leading Goose”R&DProgram of Zhejiang under grant 2026LDC01003(JT).
摘要To improve the accuracy of small object feature detection in complex backgrounds for Unmanned Aerial Vehicle(UAV)aerial photography and reduce computational complexity,we propose the lightweight UAV aerial photography small object detection method based on multi-scale feature fusion and contextual information.Firstly,by introducing the grouped content-aware reassembly(GCA)operator and designing lightweight pinwheel context convolution(LPConv),we extend the feature fusion path to the P2 layer,constructing a lightweight multi-scale feature fusion network(SG-PANet).Through the decoupling of fine-grained small object features and background interference features by the GCA operator,combined with the anisotropic receptive field constructed by LPConv,our proposed method can effectively preserve the geometric details of small objects.Furthermore,we introduce the cross-stage dense feature refinement(CSPStage)module as the pre-refining unit of the detection head,and use the full history state awareness mechanism to strengthen feature reuse and gradient propagation to solve the problem of feature degradation across layers.We utilize the Wise-IoU v3 loss function to dynamically optimize the gradient gains of high-quality and low-quality samples,thereby enhancing the detection accuracy and convergence speed of the proposed method in complex scenarios.Finally,we verified the superiority and generalization of the proposed method on the VisDrone2019 dataset and DOTAv1.5 dataset.The results show that compared with YOLOv11n,MFCI-YOLO’s detection mAP50-95 increased by 11.1%,small object mAP50 increased by 16.1%,and mAP50 reached 80.3%.It provides a practical solution for detecting small objects in dense scenes.
基金financially supported byChongqingUniversity of Technology Graduate Innovation Foundation(Grant No.gzlcx20253267).
摘要Camouflaged Object Detection(COD)aims to identify objects that share highly similar patterns—such as texture,intensity,and color—with their surrounding environment.Due to their intrinsic resemblance to the background,camouflaged objects often exhibit vague boundaries and varying scales,making it challenging to accurately locate targets and delineate their indistinct edges.To address this,we propose a novel camouflaged object detection network called Edge-Guided and Multi-scale Fusion Network(EGMFNet),which leverages edge-guided multi-scale integration for enhanced performance.The model incorporates two innovative components:a Multi-scale Fusion Module(MSFM)and an Edge-Guided Attention Module(EGA).These designs exploit multi-scale features to uncover subtle cues between candidate objects and the background while emphasizing camouflaged object boundaries.Moreover,recognizing the rich contextual information in fused features,we introduce a Dual-Branch Global Context Module(DGCM)to refine features using extensive global context,thereby generatingmore informative representations.Experimental results on four benchmark datasets demonstrate that EGMFNet outperforms state-of-the-art methods across five evaluation metrics.Specifically,on COD10K,our EGMFNet-P improves Fβby 4.8 points and reduces mean absolute error(MAE)by 0.006 compared with ZoomNeXt;on NC4K,it achieves a 3.6-point increase in Fβ.OnCAMO and CHAMELEON,it obtains 4.5-point increases in Fβ,respectively.These consistent gains substantiate the superiority and robustness of EGMFNet.
基金supported by the State Grid Sichuan Electric Power Company Employee Technological Innovation Project”Research on Internal Topological Structure Perception and Health Assessment in Electricity Meters”(Grant No.:B319Q4250008).
摘要In recent years,with the advancement of computational hardware performance,machine learning algorithms have achieved significant development and widespread application across various fields,and have become deeply embedded in smart grids and communication systems.However,it is important to note that despite the widespread deployment of smart meters in the power system,the lack of reliable intelligent diagnostic,a large number of such electricity meters experiencing communication failures caused by internal topological defects every year.To address this issue,we propose a machine learning-based monitoring and early warning model using multidimensional feature fusion.By integrating more than twenty key features in four categories,including attribute features,operational load features,communication behavior features,and derived combined features-an XGBoost classification algorithm framework is constructed to implement risk early warning for electricity meter communication faults.Validated with data from millions of users,the proposed model achieves an accuracy of approximately 90%,the annual average reduction in power outages caused by communication faults is more than 10,000 hours,and significantly enhances the grid’s safety and operational stability.