Gait recognition is a key biometric for long-distance identification,yet its performance is severely degraded by real-world challenges such as varying clothing,carrying conditions,and changing viewpoints.While combini...Gait recognition is a key biometric for long-distance identification,yet its performance is severely degraded by real-world challenges such as varying clothing,carrying conditions,and changing viewpoints.While combining silhouette and skeleton data is a promising direction,effectively fusing these heterogeneous modalities and adaptively weighting their contributions in response to diverse conditions remains a central problem.This paper introduces GaitMAFF,a novelMulti-modal Adaptive Feature Fusion Network,to address this challenge.Our approach first transforms discrete skeleton joints into a dense SkeletonMap representation to align with silhouettes,then employs an attention-based module to dynamically learn the fusion weights between the two modalities.These fused features are processed by a powerful spatio-temporal backbone withWeighted Global-Local Feature FusionModules(WFFM)to learn a discriminative representation.Extensive experiments on the challenging CCPG and Gait3D datasets show that GaitMAFF achieves state-of-the-art performance,with an average Rank-1 accuracy of 84.6%on CCPG and 58.7%on Gait3D.These results demonstrate that our adaptive fusion strategy effectively integrates complementary multimodal information,significantly enhancing gait recognition robustness and accuracy in complex scenes and providing a practical solution for real-world applications.展开更多
In multi-modal emotion recognition,excessive reliance on historical context often impedes the detection of emotional shifts,while modality heterogeneity and unimodal noise limit recognition performance.Existing method...In multi-modal emotion recognition,excessive reliance on historical context often impedes the detection of emotional shifts,while modality heterogeneity and unimodal noise limit recognition performance.Existing methods struggle to dynamically adjust cross-modal complementary strength to optimize fusion quality and lack effective mechanisms to model the dynamic evolution of emotions.To address these issues,we propose a multi-level dynamic gating and emotion transfer framework for multi-modal emotion recognition.A dynamic gating mechanism is applied across unimodal encoding,cross-modal alignment,and emotion transfer modeling,substantially improving noise robustness and feature alignment.First,we construct a unimodal encoder based on gated recurrent units and feature-selection gating to suppress intra-modal noise and enhance contextual representation.Second,we design a gated-attention crossmodal encoder that dynamically calibrates the complementary contributions of visual and audio modalities to the dominant textual features and eliminates redundant information.Finally,we introduce a gated enhanced emotion transfer module that explicitly models the temporal dependence of emotional evolution in dialogues via transfer gating and optimizes continuity modeling with a comparative learning loss.Experimental results demonstrate that the proposed method outperforms state-of-the-art models on the public MELD and IEMOCAP datasets.展开更多
The fasteners employed in the railway tracks are susceptible to defects arising from their intricate composition.Foreign objects are frequently observed on the track bed in an open environment.These two types of defec...The fasteners employed in the railway tracks are susceptible to defects arising from their intricate composition.Foreign objects are frequently observed on the track bed in an open environment.These two types of defects pose potential threats to high-speed trains,thus necessitating timely and accurate track inspection.The majority of extant automatic inspection methods are predicated on the utilization of single visible light data,and the efficacy of the algorithmic processes is influenced by complex environments.Furthermore,due to the single information dimension,the detection accuracy of defects in similar,occluded,and small object categories is low.To address the aforementioned issues,this paper proposes a track defect detectionmethod based on dynamicmulti-modal fusion and challenging object enhanced perception.First,in light of the variances in the representation dimensions ofmultimodal information,this paper proposes a dynamic weighted multi-modal feature fusion module.The fused multi-modal features are assigned weights,and thenmultiplied with the extracted single-modal features atmultiple levels,achieving adaptive adjustment of the response degree of fusion features.Second,a novel stepwise multi-scale convolution feature aggregation module is proposed for challenging objects.The proposed method employs depth separable convolution and cross-scale aggregation operations of different receptive fields to enhance feature extraction and reuse,thereby reducing the degree of progressive loss of effective information.The experimental results demonstrate the efficacy of the proposed method in comparison to eight established methods,encompassing both single-modal and multi-modal methods,as evidenced by the extensive findings within the constructed RGBD dataset.展开更多
To address the challenge of achieving decentralized,scalable,and adaptive control for large-scale multiple unmanned aerial vehicle(multi-UAV)swarms in dynamic urban environments with obstacles and wind perturbations,w...To address the challenge of achieving decentralized,scalable,and adaptive control for large-scale multiple unmanned aerial vehicle(multi-UAV)swarms in dynamic urban environments with obstacles and wind perturbations,we proposed a hybrid framework integrating adaptive reinforcement learning(RL),multi-modal perception fusion,and enhanced pigeon flock optimization(PFO)with curiosity-driven exploration to enable robust autonomous and formation control.The framework leverages meta-learning to optimize RL policies for real-time adaptation,fuses sensor data for precise state estimation,and enhances PFO with learned leader-follower dynamics and exploration rewards to maintain cohesive formations and explore uncertain areas.For swarms of 10–30 UAVs,it achieves 34%faster convergence,61%reduced stability root mean square error(RMSE),88%fewer collisions and 85.6%–92.3%success rates in target detection and encirclement,outperforming standard multi-agent RL,pure PFO,and single-modality RL.Three-dimensional trajectory visualizations confirm cohesive formations,collision-free maneuvers,and efficient exploration in urban search-and-rescue scenarios.Innovations include meta-RL for rapid adaptation,multi-modal fusion for robust perception,and curiosity-driven PFO for scalable,decentralized control,advancing real-world multi-UAV swarm autonomy and coordination.展开更多
This paper exploits multi-modal Physical(PHY)-layer features in terms of artificial fingerprint,In-phase/Quadrature(IQ)imbalance and Angle of Arrival(AoA)to propose a novel PHY-layer authentication framework for a Mil...This paper exploits multi-modal Physical(PHY)-layer features in terms of artificial fingerprint,In-phase/Quadrature(IQ)imbalance and Angle of Arrival(AoA)to propose a novel PHY-layer authentication framework for a Millimeter Wave(mmWave)Multiple-Input Multiple-Output(MIMO)Unmanned Aerial Vehicle(UAV)-enabled communication system.First,we resort to the AoA-based spatial fingerprint to effectively address the challenge of channel fingerprint instability induced by high-speed UAV mobility.To further enhance the low discriminability of hardware fingerprints caused by refined manufacturing techniques,artificial Gaussian noise is injected into the transmission signals to assist the receiver in better distinguishing between legitimate and illegitimate UAVs.Then,we jointly combine with inherent IQ imbalance and AoA features to design a hybrid authentication scheme and thus construct a multi-dimensional fingerprint space for a comprehensive characterization of UAV identities.To theoretically evaluate the effectiveness of the proposed authentication framework,the analytical closed-form expressions of performance metrics like false alarm and detection probabilities are also exactly derived based on the statistical signal processing technology and composite hypothesis testing.Finally,we provide large simulation results to validate the correctness and feasibility of the proposed theoretical models,and also discuss the relation between system security and communication service quality under different artificial fingerprint level.展开更多
Background:In recent years,deep convolutional neural networks(CNNs)have achieved great successes in medical imaging.However,it is difficult to obtain accurate pathological information for clinical diagnosis and treatm...Background:In recent years,deep convolutional neural networks(CNNs)have achieved great successes in medical imaging.However,it is difficult to obtain accurate pathological information for clinical diagnosis and treatment by leveraging single-modality medical images.This study aims to provide an efficient multimodality whole heart segmentation method for the diagnosis of coronary heart disease.Methods:We propose SFAM-TransUnet for multimodality whole heart segmentation,a novel deep learning framework combining CNNs and transformers.Primarily,the method integrates CNNs and visual transformers(Vits)into a unified fusion framework.Specifically,the shallow feature fusion module is designed to connect MRI and CT images,thereby providing a powerful and efficient multimodality fusion backbone for semantic segmentation.Furthermore,we propose a fusion ViT(FViT)module including self-attention(SA)and adaptive mutual boost attention(Ada-MBA)to enhance contextual information within and across modalities.The Ada-MBA module assigns attention to semantic perception regions by calculating SA and cross-attention,which improves the ability to understand context from the different modalities.Extensive experiments are con-ducted on the clinical Multi-Modality Whole Heart Segmentation datasets.Results:We successfully improved the whole heart segmentation DSCs to 0.902(AA),0.920(LV-blood),0.863(LA-blood),and 0.837(LV-myo),the HDs to 9.886(AA),9.947(LV-blood),11.911(LA-blood),and 13.599(LV-myo),the PSNR values to 33.577(AA),30.091(LV-blood),32.055(LA-blood),and 29.837(LV-myo),SSMI values to 0.901(AA),0.818(LV-blood),0.765(LA-blood),and 0.743(LV-myo).This demonstrate SFAM-TransUnet outperforms various alternative methods.Conclusions:We propose SFAM-TransUnet,an efficient framework tailored for whole heart segmentation that combines CNNs and transformers.It provides a powerful multimodality fusion network to improve the performance of whole heart semantic segmentation.These results demonstrate the efficacy of SFAM-TransUnet in integrating relevant information between different modalities in multimodal tasks.展开更多
Modern malware is increasingly employing polymorphism,packing,and metamorphism to evade traditional signature-based detection.Because of this,there is an urgency to have more reliable classification systems.Visual mal...Modern malware is increasingly employing polymorphism,packing,and metamorphism to evade traditional signature-based detection.Because of this,there is an urgency to have more reliable classification systems.Visual malware analysis,where binaries are converted into grayscale images,has demonstrated potential in revealing structural patterns of malware family classification.However,recent methods mostly rely on single-stream,lightweight Convolutional Neural Networks(CNNs).These models have a major blind spot.The visual representation textures can be heavily obscured without changing the underlying malicious code,causing severe performance drops on newer or even rare malware classes.This paper presents a Hybrid Multi-Modal Deep Learning framework to fix this vulnerability.The proposed dual-stream architecture uses image recognition via EfficientNetB0 alongside metadata analysis using 1D-convolutional byte embeddings.This paper evaluated the framework on the modern MalwareVision-2025 dataset(approximately 125000 samples)and the legacy Malimg dataset(9339 samples).On MalwareVision-2025,the model reached a weighted accuracy of 88.44%and achieved 100%benign recall on the evaluated split.The testing across both datasets shows that combining visual and structural features reduces modality collapse.This creates a much stronger system compared to using either input type on its own.In particular,the hybrid approach improves detection performance on deeply hidden contemporary threats,including the AveMariaRAT and CobaltStrike families.展开更多
Aiming at the problems of data sparsity,uneven behavior weight allocation,and insufficient timeliness modeling existing in traditional recommendation systems in the scenario of personalized fashion recommendation,this...Aiming at the problems of data sparsity,uneven behavior weight allocation,and insufficient timeliness modeling existing in traditional recommendation systems in the scenario of personalized fashion recommendation,this paper proposes a personalized recommendation method that integrates multi-behavior weights and multi-modal features.A dynamic weighted collaborative filtering algorithm is designed,which comprehensively considers the multi-dimensional behaviors of users,and introduces a time attenuation factor to construct a time-sensitive user-item scoring matrix,so as to more accurately depict the dynamic changes of user interests.A multi-modal deep fusion framework is built:ResNet-50 is used to extract commodity image features,and the pre-trained BERT model is combined to extract text features;meanwhile,the multi-head self-attention mechanism is adopted to realize semantic-level interaction and adaptive fusion of cross-modal features,thereby enhancing the expressive ability of commodity representation.Then,user preference score prediction is carried out based on the deep predictive network to generate a personalized recommendation list.Experimental results on real e-commerce datasets show that the method in this paper achieves 0.703 and 0.491 on HR@5 and NDCG@5,respectively,which is significantly superior to other baseline models.Ablation experiments further verify the effectiveness of each module including time attenuation,multi-behavior weights and multi-modal features.This study provides a more accurate,dynamic and transparently interpretable personalized recommendation solution for e-commerce platforms,and has certain theoretical value and practical significance.展开更多
The urgent need for advanced fire detection methods stems from the increased intensity of fire incidents,which cause massive property loss and irreversible damage.To overcome the limitations of traditional fire detect...The urgent need for advanced fire detection methods stems from the increased intensity of fire incidents,which cause massive property loss and irreversible damage.To overcome the limitations of traditional fire detection methods,such as those of smoke detectors,fire detection based on computer vision(CV)algorithms has been adopted to improve detection accuracy.Compared to single-modal fire detection,multi-modal fire detection has gained attention because it leverages the richer information present in both RGB and thermal images.However,prevalent multi-modal fire detection methods significantly increase model complexity by requiring two separate streams in the backbone to process RGB and thermal images independently.To address this issue,this paper proposes a four-channel single-stream fire detection method based on YOLOv5,which concatenates RGB and thermal images to form the required four-channel input.Comparison experiments with dual-stream YOLOv5 models using add fusion and transformer fusion demonstrate that the four-channel single-stream model reduces model complexity while improving detection accuracy.To further enhance detection accuracy and reduce model complexity,this study redesigned YOLOv5’s C3 module by integrating the convolutional block attention module(CBAM)to form the C3CBAM module and introduced the SCYLLA-Intersection over Union(SIoU)loss function.By comparing its performance with that of state-of-the-art(SOTA)models in multi-modal object detection,such as the YOLOv5-based dual-stream model,this study shows that the proposed approach improves detection in the diverse conditions presented in the selected dataset.展开更多
This paper proposes an efficient algorithm for real-time multi-modal image matching based on a lightweight feature fusion network,targeting the challenges of multi-modal image matching in multi-source data analysis.Th...This paper proposes an efficient algorithm for real-time multi-modal image matching based on a lightweight feature fusion network,targeting the challenges of multi-modal image matching in multi-source data analysis.The algorithm addresses significant multi-modal feature differences and real-time processing limitations by incorporating key technologies including reparameterization in convolutional neural networks,multi-scale image pyramids,and feature fusion modules.The matching process employs a coarse-to-fine strategy,ensuring robust performance in complex environments.Experimental results using multi-modal datasets demonstrate that the proposed algorithm achieves superior accuracy and speed,with a success rate of 98.3%and an average matching time of 30.51 ms per 500×500 image pair.These results highlight the practical value and strong generalization capability of the algorithm in real-time applications.展开更多
To address the challenges of dusty,foggy and other complex construction site environments leading to the failure of visible light imaging and difficulties in small target detection,as well as the high resource consump...To address the challenges of dusty,foggy and other complex construction site environments leading to the failure of visible light imaging and difficulties in small target detection,as well as the high resource consumption hindering model deployment,an enhanced and lightweight algorithm is proposed.This algorithm employs a hybrid architecture,integrating red green blue(RGB)(visible light)and thermal infrared(RGBT)multi-modal images through a fusion framework based on you only look once(YOLO)version 8 and Mamba-Transformer(MT).We refer to this integrated model as YOLOv8-RGBT-MT.In terms of network improvements,a frequency enhancement module is first employed to enhance visible light and infrared images.And then,a module integrating Mamba and Transformer components is designed to replace base convolutional blocks in the backbone network,thereby expanding the receptive field of the model and improving feature extraction in complex backgrounds.Finally,a multi-modal feature fusion mechanism is introduced,through which complementary information from visible and infrared images is effectively integrated via an adaptive weighting strategy,so that both the detection accuracy and robustness for small targets are enhanced.Experimental results demonstrate that,compared to YOLOv8-RGBT,the enhanced algorithm achieves an improvement of 18.7%in mAP50,while reducing the number of inference time by 79.7%.展开更多
To improve the efficiency of power grid emergency response after disasters,this study proposes a multi-modal risk profiling-driven power grid disaster emergency response strategy and dynamic resource synergy optimizat...To improve the efficiency of power grid emergency response after disasters,this study proposes a multi-modal risk profiling-driven power grid disaster emergency response strategy and dynamic resource synergy optimization model.A risk assessment model is constructed by integrating equipment health status,real-time failure rate,and power grid topology importance to generate equipment risk profiles for identifying key nodes.A two-stage optimization mechanism is then designed,the first stage achieves priority coverage of high-risk equipment and minimization of inspection costs through multi-objective path planning.The second stage adopts a mixed-integer programming model to coordinate personnel scheduling and material allocation under resource constraints.A rolling optimization framework is introduced to dynamically respond to sudden failures and resource changes,ensuring the adaptability of scheduling schemes.To verify the model’s effectiveness,three typical scenarios,”no sudden failures”,“equipment risk escalation”,and“personnel working hour constraints”,are simulated.Compared with traditional strategies,the model significantly improves the rationality and dynamic adaptability of resource scheduling,providing new ideas and engineering practice support for enhancing the resilience of smart grid disaster emergency response.展开更多
Multi-modal knowledge graph completion(MMKGC)aims to complete missing entities or relations in multi-modal knowledge graphs,thereby discovering more previously unknown triples.Due to the continuous growth of data and ...Multi-modal knowledge graph completion(MMKGC)aims to complete missing entities or relations in multi-modal knowledge graphs,thereby discovering more previously unknown triples.Due to the continuous growth of data and knowledge and the limitations of data sources,the visual knowledge within the knowledge graphs is generally of low quality,and some entities suffer from the issue of missing visual modality.Nevertheless,previous studies of MMKGC have primarily focused on how to facilitate modality interaction and fusion while neglecting the problems of low modality quality and modality missing.In this case,mainstream MMKGC models only use pre-trained visual encoders to extract features and transfer the semantic information to the joint embeddings through modal fusion,which inevitably suffers from problems such as error propagation and increased uncertainty.To address these problems,we propose a Multi-modal knowledge graph Completion model based on Super-resolution and Detailed Description Generation(MMCSD).Specifically,we leverage a pre-trained residual network to enhance the resolution and improve the quality of the visual modality.Moreover,we design multi-level visual semantic extraction and entity description generation,thereby further extracting entity semantics from structural triples and visual images.Meanwhile,we train a variational multi-modal auto-encoder and utilize a pre-trained multi-modal language model to complement the missing visual features.We conducted experiments on FB15K-237 and DB13K,and the results showed that MMCSD can effectively perform MMKGC and achieve state-of-the-art performance.展开更多
Multi-modal Named Entity Recognition(MNER)aims to better identify meaningful textual entities by integrating information from images.Previous work has focused on extracting visual semantics at a fine-grained level,or ...Multi-modal Named Entity Recognition(MNER)aims to better identify meaningful textual entities by integrating information from images.Previous work has focused on extracting visual semantics at a fine-grained level,or obtaining entity related external knowledge from knowledge bases or Large Language Models(LLMs).However,these approaches ignore the poor semantic correlation between visual and textual modalities in MNER datasets and do not explore different multi-modal fusion approaches.In this paper,we present MMAVK,a multi-modal named entity recognition model with auxiliary visual knowledge and word-level fusion,which aims to leverage the Multi-modal Large Language Model(MLLM)as an implicit knowledge base.It also extracts vision-based auxiliary knowledge from the image formore accurate and effective recognition.Specifically,we propose vision-based auxiliary knowledge generation,which guides the MLLM to extract external knowledge exclusively derived from images to aid entity recognition by designing target-specific prompts,thus avoiding redundant recognition and cognitive confusion caused by the simultaneous processing of image-text pairs.Furthermore,we employ a word-level multi-modal fusion mechanism to fuse the extracted external knowledge with each word-embedding embedded from the transformerbased encoder.Extensive experimental results demonstrate that MMAVK outperforms or equals the state-of-the-art methods on the two classical MNER datasets,even when the largemodels employed have significantly fewer parameters than other baselines.展开更多
Integrating multiple medical imaging techniques,including Magnetic Resonance Imaging(MRI),Computed Tomography,Positron Emission Tomography(PET),and ultrasound,provides a comprehensive view of the patient health status...Integrating multiple medical imaging techniques,including Magnetic Resonance Imaging(MRI),Computed Tomography,Positron Emission Tomography(PET),and ultrasound,provides a comprehensive view of the patient health status.Each of these methods contributes unique diagnostic insights,enhancing the overall assessment of patient condition.Nevertheless,the amalgamation of data from multiple modalities presents difficulties due to disparities in resolution,data collection methods,and noise levels.While traditional models like Convolutional Neural Networks(CNNs)excel in single-modality tasks,they struggle to handle multi-modal complexities,lacking the capacity to model global relationships.This research presents a novel approach for examining multi-modal medical imagery using a transformer-based system.The framework employs self-attention and cross-attention mechanisms to synchronize and integrate features across various modalities.Additionally,it shows resilience to variations in noise and image quality,making it adaptable for real-time clinical use.To address the computational hurdles linked to transformer models,particularly in real-time clinical applications in resource-constrained environments,several optimization techniques have been integrated to boost scalability and efficiency.Initially,a streamlined transformer architecture was adopted to minimize the computational load while maintaining model effectiveness.Methods such as model pruning,quantization,and knowledge distillation have been applied to reduce the parameter count and enhance the inference speed.Furthermore,efficient attention mechanisms such as linear or sparse attention were employed to alleviate the substantial memory and processing requirements of traditional self-attention operations.For further deployment optimization,researchers have implemented hardware-aware acceleration strategies,including the use of TensorRT and ONNX-based model compression,to ensure efficient execution on edge devices.These optimizations allow the approach to function effectively in real-time clinical settings,ensuring viability even in environments with limited resources.Future research directions include integrating non-imaging data to facilitate personalized treatment and enhancing computational efficiency for implementation in resource-limited environments.This study highlights the transformative potential of transformer models in multi-modal medical imaging,offering improvements in diagnostic accuracy and patient care outcomes.展开更多
With the advent of the next-generation Air Traffic Control(ATC)system,there is growing interest in using Artificial Intelligence(AI)techniques to enhance Situation Awareness(SA)for ATC Controllers(ATCOs),i.e.,Intellig...With the advent of the next-generation Air Traffic Control(ATC)system,there is growing interest in using Artificial Intelligence(AI)techniques to enhance Situation Awareness(SA)for ATC Controllers(ATCOs),i.e.,Intelligent SA(ISA).However,the existing AI-based SA approaches often rely on unimodal data and lack a comprehensive description and benchmark of the ISA tasks utilizing multi-modal data for real-time ATC environments.To address this gap,by analyzing the situation awareness procedure of the ATCOs,the ISA task is refined to the processing of the two primary elements,i.e.,spoken instructions and flight trajectories.Subsequently,the ISA is further formulated into Controlling Intent Understanding(CIU)and Flight Trajectory Prediction(FTP)tasks.For the CIU task,an innovative automatic speech recognition and understanding framework is designed to extract the controlling intent from unstructured and continuous ATC communications.For the FTP task,the single-and multi-horizon FTP approaches are investigated to support the high-precision prediction of the situation evolution.A total of 32 unimodal/multi-modal advanced methods with extensive evaluation metrics are introduced to conduct the benchmarks on the real-world multi-modal ATC situation dataset.Experimental results demonstrate the effectiveness of AI-based techniques in enhancing ISA for the ATC environment.展开更多
Traditional Chinese medicine(TCM)demonstrates distinctive advantages in disease prevention and treatment.However,analyzing its biological mechanisms through the modern medical research paradigm of“single drug,single ...Traditional Chinese medicine(TCM)demonstrates distinctive advantages in disease prevention and treatment.However,analyzing its biological mechanisms through the modern medical research paradigm of“single drug,single target”presents significant challenges due to its holistic approach.Network pharmacology and its core theory of network targets connect drugs and diseases from a holistic and systematic perspective based on biological networks,overcoming the limitations of reductionist research models and showing considerable value in TCM research.Recent integration of network target computational and experimental methods with artificial intelligence(AI)and multi-modal multi-omics technologies has substantially enhanced network pharmacology methodology.The advancement in computational and experimental techniques provides complementary support for network target theory in decoding TCM principles.This review,centered on network targets,examines the progress of network target methods combined with AI in predicting disease molecular mechanisms and drug-target relationships,alongside the application of multi-modal multi-omics technologies in analyzing TCM formulae,syndromes,and toxicity.Looking forward,network target theory is expected to incorporate emerging technologies while developing novel approaches aligned with its unique characteristics,potentially leading to significant breakthroughs in TCM research and advancing scientific understanding and innovation in TCM.展开更多
The multi-modal characteristics of mineral particles play a pivotal role in enhancing the classification accuracy,which is critical for obtaining a profound understanding of the Earth's composition and ensuring ef...The multi-modal characteristics of mineral particles play a pivotal role in enhancing the classification accuracy,which is critical for obtaining a profound understanding of the Earth's composition and ensuring effective exploitation utilization of its resources.However,the existing methods for classifying mineral particles do not fully utilize these multi-modal features,thereby limiting the classification accuracy.Furthermore,when conventional multi-modal image classification methods are applied to planepolarized and cross-polarized sequence images of mineral particles,they encounter issues such as information loss,misaligned features,and challenges in spatiotemporal feature extraction.To address these challenges,we propose a multi-modal mineral particle polarization image classification network(MMGC-Net)for precise mineral particle classification.Initially,MMGC-Net employs a two-dimensional(2D)backbone network with shared parameters to extract features from two types of polarized images to ensure feature alignment.Subsequently,a cross-polarized intra-modal feature fusion module is designed to refine the spatiotemporal features from the extracted features of the cross-polarized sequence images.Ultimately,the inter-modal feature fusion module integrates the two types of modal features to enhance the classification precision.Quantitative and qualitative experimental results indicate that when compared with the current state-of-the-art multi-modal image classification methods,MMGC-Net demonstrates marked superiority in terms of mineral particle multi-modal feature learning and four classification evaluation metrics.It also demonstrates better stability than the existing models.展开更多
To address the challenge of missing modal information in entity alignment and to mitigate information loss or bias arising frommodal heterogeneity during fusion,while also capturing shared information acrossmodalities...To address the challenge of missing modal information in entity alignment and to mitigate information loss or bias arising frommodal heterogeneity during fusion,while also capturing shared information acrossmodalities,this paper proposes a Multi-modal Pre-synergistic Entity Alignmentmodel based on Cross-modalMutual Information Strategy Optimization(MPSEA).The model first employs independent encoders to process multi-modal features,including text,images,and numerical values.Next,a multi-modal pre-synergistic fusion mechanism integrates graph structural and visual modal features into the textual modality as preparatory information.This pre-fusion strategy enables unified perception of heterogeneous modalities at the model’s initial stage,reducing discrepancies during the fusion process.Finally,using cross-modal deep perception reinforcement learning,the model achieves adaptive multilevel feature fusion between modalities,supporting learningmore effective alignment strategies.Extensive experiments on multiple public datasets show that the MPSEA method achieves gains of up to 7% in Hits@1 and 8.2% in MRR on the FBDB15K dataset,and up to 9.1% in Hits@1 and 7.7% in MRR on the FBYG15K dataset,compared to existing state-of-the-art methods.These results confirm the effectiveness of the proposed model.展开更多
基金funded by the Natural Science Foundation of Chongqing Municipality,grant number CSTB2022NSCQ-MSX0503.
摘要Gait recognition is a key biometric for long-distance identification,yet its performance is severely degraded by real-world challenges such as varying clothing,carrying conditions,and changing viewpoints.While combining silhouette and skeleton data is a promising direction,effectively fusing these heterogeneous modalities and adaptively weighting their contributions in response to diverse conditions remains a central problem.This paper introduces GaitMAFF,a novelMulti-modal Adaptive Feature Fusion Network,to address this challenge.Our approach first transforms discrete skeleton joints into a dense SkeletonMap representation to align with silhouettes,then employs an attention-based module to dynamically learn the fusion weights between the two modalities.These fused features are processed by a powerful spatio-temporal backbone withWeighted Global-Local Feature FusionModules(WFFM)to learn a discriminative representation.Extensive experiments on the challenging CCPG and Gait3D datasets show that GaitMAFF achieves state-of-the-art performance,with an average Rank-1 accuracy of 84.6%on CCPG and 58.7%on Gait3D.These results demonstrate that our adaptive fusion strategy effectively integrates complementary multimodal information,significantly enhancing gait recognition robustness and accuracy in complex scenes and providing a practical solution for real-world applications.
基金funded by“the Fanying Special Program of the National Natural Science Foundation of China,grant number 62341307”“the Scientific research project of Jiangxi Provincial Department of Education,grant number GJJ200839”“the Doctoral startup fund of Jiangxi University of Technology,grant number 205200100402”.
摘要In multi-modal emotion recognition,excessive reliance on historical context often impedes the detection of emotional shifts,while modality heterogeneity and unimodal noise limit recognition performance.Existing methods struggle to dynamically adjust cross-modal complementary strength to optimize fusion quality and lack effective mechanisms to model the dynamic evolution of emotions.To address these issues,we propose a multi-level dynamic gating and emotion transfer framework for multi-modal emotion recognition.A dynamic gating mechanism is applied across unimodal encoding,cross-modal alignment,and emotion transfer modeling,substantially improving noise robustness and feature alignment.First,we construct a unimodal encoder based on gated recurrent units and feature-selection gating to suppress intra-modal noise and enhance contextual representation.Second,we design a gated-attention crossmodal encoder that dynamically calibrates the complementary contributions of visual and audio modalities to the dominant textual features and eliminates redundant information.Finally,we introduce a gated enhanced emotion transfer module that explicitly models the temporal dependence of emotional evolution in dialogues via transfer gating and optimizes continuity modeling with a comparative learning loss.Experimental results demonstrate that the proposed method outperforms state-of-the-art models on the public MELD and IEMOCAP datasets.
基金funded by Beijing Natural Science Foundation,grant number L241078.
摘要The fasteners employed in the railway tracks are susceptible to defects arising from their intricate composition.Foreign objects are frequently observed on the track bed in an open environment.These two types of defects pose potential threats to high-speed trains,thus necessitating timely and accurate track inspection.The majority of extant automatic inspection methods are predicated on the utilization of single visible light data,and the efficacy of the algorithmic processes is influenced by complex environments.Furthermore,due to the single information dimension,the detection accuracy of defects in similar,occluded,and small object categories is low.To address the aforementioned issues,this paper proposes a track defect detectionmethod based on dynamicmulti-modal fusion and challenging object enhanced perception.First,in light of the variances in the representation dimensions ofmultimodal information,this paper proposes a dynamic weighted multi-modal feature fusion module.The fused multi-modal features are assigned weights,and thenmultiplied with the extracted single-modal features atmultiple levels,achieving adaptive adjustment of the response degree of fusion features.Second,a novel stepwise multi-scale convolution feature aggregation module is proposed for challenging objects.The proposed method employs depth separable convolution and cross-scale aggregation operations of different receptive fields to enhance feature extraction and reuse,thereby reducing the degree of progressive loss of effective information.The experimental results demonstrate the efficacy of the proposed method in comparison to eight established methods,encompassing both single-modal and multi-modal methods,as evidenced by the extensive findings within the constructed RGBD dataset.
基金supported by the National Natural Science Foundation of China(No.62350048)。
摘要To address the challenge of achieving decentralized,scalable,and adaptive control for large-scale multiple unmanned aerial vehicle(multi-UAV)swarms in dynamic urban environments with obstacles and wind perturbations,we proposed a hybrid framework integrating adaptive reinforcement learning(RL),multi-modal perception fusion,and enhanced pigeon flock optimization(PFO)with curiosity-driven exploration to enable robust autonomous and formation control.The framework leverages meta-learning to optimize RL policies for real-time adaptation,fuses sensor data for precise state estimation,and enhances PFO with learned leader-follower dynamics and exploration rewards to maintain cohesive formations and explore uncertain areas.For swarms of 10–30 UAVs,it achieves 34%faster convergence,61%reduced stability root mean square error(RMSE),88%fewer collisions and 85.6%–92.3%success rates in target detection and encirclement,outperforming standard multi-agent RL,pure PFO,and single-modality RL.Three-dimensional trajectory visualizations confirm cohesive formations,collision-free maneuvers,and efficient exploration in urban search-and-rescue scenarios.Innovations include meta-RL for rapid adaptation,multi-modal fusion for robust perception,and curiosity-driven PFO for scalable,decentralized control,advancing real-world multi-UAV swarm autonomy and coordination.
基金supported in part by the National Key R&D Program of China under Grant 2023YFB3107500in part by the National Natural Science Foundation of China under Grant 62272241+1 种基金in part by the State Key Laboratory of Integrated Services Networks(Xidian University),under Grant ISN24-18in part by the Nanjing University of Posts and Telecommunications Scientific Research Foundation under Grant NY221122。
摘要This paper exploits multi-modal Physical(PHY)-layer features in terms of artificial fingerprint,In-phase/Quadrature(IQ)imbalance and Angle of Arrival(AoA)to propose a novel PHY-layer authentication framework for a Millimeter Wave(mmWave)Multiple-Input Multiple-Output(MIMO)Unmanned Aerial Vehicle(UAV)-enabled communication system.First,we resort to the AoA-based spatial fingerprint to effectively address the challenge of channel fingerprint instability induced by high-speed UAV mobility.To further enhance the low discriminability of hardware fingerprints caused by refined manufacturing techniques,artificial Gaussian noise is injected into the transmission signals to assist the receiver in better distinguishing between legitimate and illegitimate UAVs.Then,we jointly combine with inherent IQ imbalance and AoA features to design a hybrid authentication scheme and thus construct a multi-dimensional fingerprint space for a comprehensive characterization of UAV identities.To theoretically evaluate the effectiveness of the proposed authentication framework,the analytical closed-form expressions of performance metrics like false alarm and detection probabilities are also exactly derived based on the statistical signal processing technology and composite hypothesis testing.Finally,we provide large simulation results to validate the correctness and feasibility of the proposed theoretical models,and also discuss the relation between system security and communication service quality under different artificial fingerprint level.
基金supported by the Henan Province Science and Technology Research Project(Grant 252102311276)Henan Province Key Scientific Research Projects of Universities(Grant 25B520002)+1 种基金the Fund of the Institute of Complexity Science from Henan University of Technology(Grant CSKFJJ-2025-13)the 2023 Research Nursery Engineering Project of Henan University of Chinese Medicine(Grant MP2023-10).
摘要Background:In recent years,deep convolutional neural networks(CNNs)have achieved great successes in medical imaging.However,it is difficult to obtain accurate pathological information for clinical diagnosis and treatment by leveraging single-modality medical images.This study aims to provide an efficient multimodality whole heart segmentation method for the diagnosis of coronary heart disease.Methods:We propose SFAM-TransUnet for multimodality whole heart segmentation,a novel deep learning framework combining CNNs and transformers.Primarily,the method integrates CNNs and visual transformers(Vits)into a unified fusion framework.Specifically,the shallow feature fusion module is designed to connect MRI and CT images,thereby providing a powerful and efficient multimodality fusion backbone for semantic segmentation.Furthermore,we propose a fusion ViT(FViT)module including self-attention(SA)and adaptive mutual boost attention(Ada-MBA)to enhance contextual information within and across modalities.The Ada-MBA module assigns attention to semantic perception regions by calculating SA and cross-attention,which improves the ability to understand context from the different modalities.Extensive experiments are con-ducted on the clinical Multi-Modality Whole Heart Segmentation datasets.Results:We successfully improved the whole heart segmentation DSCs to 0.902(AA),0.920(LV-blood),0.863(LA-blood),and 0.837(LV-myo),the HDs to 9.886(AA),9.947(LV-blood),11.911(LA-blood),and 13.599(LV-myo),the PSNR values to 33.577(AA),30.091(LV-blood),32.055(LA-blood),and 29.837(LV-myo),SSMI values to 0.901(AA),0.818(LV-blood),0.765(LA-blood),and 0.743(LV-myo).This demonstrate SFAM-TransUnet outperforms various alternative methods.Conclusions:We propose SFAM-TransUnet,an efficient framework tailored for whole heart segmentation that combines CNNs and transformers.It provides a powerful multimodality fusion network to improve the performance of whole heart semantic segmentation.These results demonstrate the efficacy of SFAM-TransUnet in integrating relevant information between different modalities in multimodal tasks.
摘要Modern malware is increasingly employing polymorphism,packing,and metamorphism to evade traditional signature-based detection.Because of this,there is an urgency to have more reliable classification systems.Visual malware analysis,where binaries are converted into grayscale images,has demonstrated potential in revealing structural patterns of malware family classification.However,recent methods mostly rely on single-stream,lightweight Convolutional Neural Networks(CNNs).These models have a major blind spot.The visual representation textures can be heavily obscured without changing the underlying malicious code,causing severe performance drops on newer or even rare malware classes.This paper presents a Hybrid Multi-Modal Deep Learning framework to fix this vulnerability.The proposed dual-stream architecture uses image recognition via EfficientNetB0 alongside metadata analysis using 1D-convolutional byte embeddings.This paper evaluated the framework on the modern MalwareVision-2025 dataset(approximately 125000 samples)and the legacy Malimg dataset(9339 samples).On MalwareVision-2025,the model reached a weighted accuracy of 88.44%and achieved 100%benign recall on the evaluated split.The testing across both datasets shows that combining visual and structural features reduces modality collapse.This creates a much stronger system compared to using either input type on its own.In particular,the hybrid approach improves detection performance on deeply hidden contemporary threats,including the AveMariaRAT and CobaltStrike families.
基金funded by the Shandong University of Technology Science and Technology Doctor Startup Fund(Project:UsingThe Real-Time Services of Variable Rate and Pauseable Non-Real-Time ServicesGrant number:4041/422022)+2 种基金the Research Fund(Project:Patent Assignment for User Identification System of a Neural Network-Based VR DeviceGrant number:9101/22502551)the Research Fund(Project:Patent Rights Transfer for the Generating 3D Facial Images with a Deep Learning Method,Grant number:9101/22502561).
摘要Aiming at the problems of data sparsity,uneven behavior weight allocation,and insufficient timeliness modeling existing in traditional recommendation systems in the scenario of personalized fashion recommendation,this paper proposes a personalized recommendation method that integrates multi-behavior weights and multi-modal features.A dynamic weighted collaborative filtering algorithm is designed,which comprehensively considers the multi-dimensional behaviors of users,and introduces a time attenuation factor to construct a time-sensitive user-item scoring matrix,so as to more accurately depict the dynamic changes of user interests.A multi-modal deep fusion framework is built:ResNet-50 is used to extract commodity image features,and the pre-trained BERT model is combined to extract text features;meanwhile,the multi-head self-attention mechanism is adopted to realize semantic-level interaction and adaptive fusion of cross-modal features,thereby enhancing the expressive ability of commodity representation.Then,user preference score prediction is carried out based on the deep predictive network to generate a personalized recommendation list.Experimental results on real e-commerce datasets show that the method in this paper achieves 0.703 and 0.491 on HR@5 and NDCG@5,respectively,which is significantly superior to other baseline models.Ablation experiments further verify the effectiveness of each module including time attenuation,multi-behavior weights and multi-modal features.This study provides a more accurate,dynamic and transparently interpretable personalized recommendation solution for e-commerce platforms,and has certain theoretical value and practical significance.
基金supported by the Campus as a Living Lab Program Grant from the Office of the Vice-President Research and Innovation at the University of British Columbia(Grant No.AWD-027461 UBCVPFIO 2023).
摘要The urgent need for advanced fire detection methods stems from the increased intensity of fire incidents,which cause massive property loss and irreversible damage.To overcome the limitations of traditional fire detection methods,such as those of smoke detectors,fire detection based on computer vision(CV)algorithms has been adopted to improve detection accuracy.Compared to single-modal fire detection,multi-modal fire detection has gained attention because it leverages the richer information present in both RGB and thermal images.However,prevalent multi-modal fire detection methods significantly increase model complexity by requiring two separate streams in the backbone to process RGB and thermal images independently.To address this issue,this paper proposes a four-channel single-stream fire detection method based on YOLOv5,which concatenates RGB and thermal images to form the required four-channel input.Comparison experiments with dual-stream YOLOv5 models using add fusion and transformer fusion demonstrate that the four-channel single-stream model reduces model complexity while improving detection accuracy.To further enhance detection accuracy and reduce model complexity,this study redesigned YOLOv5’s C3 module by integrating the convolutional block attention module(CBAM)to form the C3CBAM module and introduced the SCYLLA-Intersection over Union(SIoU)loss function.By comparing its performance with that of state-of-the-art(SOTA)models in multi-modal object detection,such as the YOLOv5-based dual-stream model,this study shows that the proposed approach improves detection in the diverse conditions presented in the selected dataset.
基金supported by National Natural Science Foundation of China Projects of International Cooperation and Exchanges(No.W2411055)。
摘要This paper proposes an efficient algorithm for real-time multi-modal image matching based on a lightweight feature fusion network,targeting the challenges of multi-modal image matching in multi-source data analysis.The algorithm addresses significant multi-modal feature differences and real-time processing limitations by incorporating key technologies including reparameterization in convolutional neural networks,multi-scale image pyramids,and feature fusion modules.The matching process employs a coarse-to-fine strategy,ensuring robust performance in complex environments.Experimental results using multi-modal datasets demonstrate that the proposed algorithm achieves superior accuracy and speed,with a success rate of 98.3%and an average matching time of 30.51 ms per 500×500 image pair.These results highlight the practical value and strong generalization capability of the algorithm in real-time applications.
摘要To address the challenges of dusty,foggy and other complex construction site environments leading to the failure of visible light imaging and difficulties in small target detection,as well as the high resource consumption hindering model deployment,an enhanced and lightweight algorithm is proposed.This algorithm employs a hybrid architecture,integrating red green blue(RGB)(visible light)and thermal infrared(RGBT)multi-modal images through a fusion framework based on you only look once(YOLO)version 8 and Mamba-Transformer(MT).We refer to this integrated model as YOLOv8-RGBT-MT.In terms of network improvements,a frequency enhancement module is first employed to enhance visible light and infrared images.And then,a module integrating Mamba and Transformer components is designed to replace base convolutional blocks in the backbone network,thereby expanding the receptive field of the model and improving feature extraction in complex backgrounds.Finally,a multi-modal feature fusion mechanism is introduced,through which complementary information from visible and infrared images is effectively integrated via an adaptive weighting strategy,so that both the detection accuracy and robustness for small targets are enhanced.Experimental results demonstrate that,compared to YOLOv8-RGBT,the enhanced algorithm achieves an improvement of 18.7%in mAP50,while reducing the number of inference time by 79.7%.
摘要To improve the efficiency of power grid emergency response after disasters,this study proposes a multi-modal risk profiling-driven power grid disaster emergency response strategy and dynamic resource synergy optimization model.A risk assessment model is constructed by integrating equipment health status,real-time failure rate,and power grid topology importance to generate equipment risk profiles for identifying key nodes.A two-stage optimization mechanism is then designed,the first stage achieves priority coverage of high-risk equipment and minimization of inspection costs through multi-objective path planning.The second stage adopts a mixed-integer programming model to coordinate personnel scheduling and material allocation under resource constraints.A rolling optimization framework is introduced to dynamically respond to sudden failures and resource changes,ensuring the adaptability of scheduling schemes.To verify the model’s effectiveness,three typical scenarios,”no sudden failures”,“equipment risk escalation”,and“personnel working hour constraints”,are simulated.Compared with traditional strategies,the model significantly improves the rationality and dynamic adaptability of resource scheduling,providing new ideas and engineering practice support for enhancing the resilience of smart grid disaster emergency response.
基金funded by Research Project,grant number BHQ090003000X03。
摘要Multi-modal knowledge graph completion(MMKGC)aims to complete missing entities or relations in multi-modal knowledge graphs,thereby discovering more previously unknown triples.Due to the continuous growth of data and knowledge and the limitations of data sources,the visual knowledge within the knowledge graphs is generally of low quality,and some entities suffer from the issue of missing visual modality.Nevertheless,previous studies of MMKGC have primarily focused on how to facilitate modality interaction and fusion while neglecting the problems of low modality quality and modality missing.In this case,mainstream MMKGC models only use pre-trained visual encoders to extract features and transfer the semantic information to the joint embeddings through modal fusion,which inevitably suffers from problems such as error propagation and increased uncertainty.To address these problems,we propose a Multi-modal knowledge graph Completion model based on Super-resolution and Detailed Description Generation(MMCSD).Specifically,we leverage a pre-trained residual network to enhance the resolution and improve the quality of the visual modality.Moreover,we design multi-level visual semantic extraction and entity description generation,thereby further extracting entity semantics from structural triples and visual images.Meanwhile,we train a variational multi-modal auto-encoder and utilize a pre-trained multi-modal language model to complement the missing visual features.We conducted experiments on FB15K-237 and DB13K,and the results showed that MMCSD can effectively perform MMKGC and achieve state-of-the-art performance.
基金funded by Research Project,grant number BHQ090003000X03.
摘要Multi-modal Named Entity Recognition(MNER)aims to better identify meaningful textual entities by integrating information from images.Previous work has focused on extracting visual semantics at a fine-grained level,or obtaining entity related external knowledge from knowledge bases or Large Language Models(LLMs).However,these approaches ignore the poor semantic correlation between visual and textual modalities in MNER datasets and do not explore different multi-modal fusion approaches.In this paper,we present MMAVK,a multi-modal named entity recognition model with auxiliary visual knowledge and word-level fusion,which aims to leverage the Multi-modal Large Language Model(MLLM)as an implicit knowledge base.It also extracts vision-based auxiliary knowledge from the image formore accurate and effective recognition.Specifically,we propose vision-based auxiliary knowledge generation,which guides the MLLM to extract external knowledge exclusively derived from images to aid entity recognition by designing target-specific prompts,thus avoiding redundant recognition and cognitive confusion caused by the simultaneous processing of image-text pairs.Furthermore,we employ a word-level multi-modal fusion mechanism to fuse the extracted external knowledge with each word-embedding embedded from the transformerbased encoder.Extensive experimental results demonstrate that MMAVK outperforms or equals the state-of-the-art methods on the two classical MNER datasets,even when the largemodels employed have significantly fewer parameters than other baselines.
基金supported by the Deanship of Research and Graduate Studies at King Khalid University under Small Research Project grant number RGP1/139/45.
摘要Integrating multiple medical imaging techniques,including Magnetic Resonance Imaging(MRI),Computed Tomography,Positron Emission Tomography(PET),and ultrasound,provides a comprehensive view of the patient health status.Each of these methods contributes unique diagnostic insights,enhancing the overall assessment of patient condition.Nevertheless,the amalgamation of data from multiple modalities presents difficulties due to disparities in resolution,data collection methods,and noise levels.While traditional models like Convolutional Neural Networks(CNNs)excel in single-modality tasks,they struggle to handle multi-modal complexities,lacking the capacity to model global relationships.This research presents a novel approach for examining multi-modal medical imagery using a transformer-based system.The framework employs self-attention and cross-attention mechanisms to synchronize and integrate features across various modalities.Additionally,it shows resilience to variations in noise and image quality,making it adaptable for real-time clinical use.To address the computational hurdles linked to transformer models,particularly in real-time clinical applications in resource-constrained environments,several optimization techniques have been integrated to boost scalability and efficiency.Initially,a streamlined transformer architecture was adopted to minimize the computational load while maintaining model effectiveness.Methods such as model pruning,quantization,and knowledge distillation have been applied to reduce the parameter count and enhance the inference speed.Furthermore,efficient attention mechanisms such as linear or sparse attention were employed to alleviate the substantial memory and processing requirements of traditional self-attention operations.For further deployment optimization,researchers have implemented hardware-aware acceleration strategies,including the use of TensorRT and ONNX-based model compression,to ensure efficient execution on edge devices.These optimizations allow the approach to function effectively in real-time clinical settings,ensuring viability even in environments with limited resources.Future research directions include integrating non-imaging data to facilitate personalized treatment and enhancing computational efficiency for implementation in resource-limited environments.This study highlights the transformative potential of transformer models in multi-modal medical imaging,offering improvements in diagnostic accuracy and patient care outcomes.
基金supported by the National Natural Science Foundation of China(Nos.62371323,62401380,U2433217,U2333209,and U20A20161)Natural Science Foundation of Sichuan Province,China(Nos.2025ZNSFSC1476)+2 种基金Sichuan Science and Technology Program,China(Nos.2024YFG0010 and 2024ZDZX0046)the Institutional Research Fund from Sichuan University(Nos.2024SCUQJTX030)the Open Fund of Key Laboratory of Flight Techniques and Flight Safety,CAAC(Nos.GY2024-01A).
摘要With the advent of the next-generation Air Traffic Control(ATC)system,there is growing interest in using Artificial Intelligence(AI)techniques to enhance Situation Awareness(SA)for ATC Controllers(ATCOs),i.e.,Intelligent SA(ISA).However,the existing AI-based SA approaches often rely on unimodal data and lack a comprehensive description and benchmark of the ISA tasks utilizing multi-modal data for real-time ATC environments.To address this gap,by analyzing the situation awareness procedure of the ATCOs,the ISA task is refined to the processing of the two primary elements,i.e.,spoken instructions and flight trajectories.Subsequently,the ISA is further formulated into Controlling Intent Understanding(CIU)and Flight Trajectory Prediction(FTP)tasks.For the CIU task,an innovative automatic speech recognition and understanding framework is designed to extract the controlling intent from unstructured and continuous ATC communications.For the FTP task,the single-and multi-horizon FTP approaches are investigated to support the high-precision prediction of the situation evolution.A total of 32 unimodal/multi-modal advanced methods with extensive evaluation metrics are introduced to conduct the benchmarks on the real-world multi-modal ATC situation dataset.Experimental results demonstrate the effectiveness of AI-based techniques in enhancing ISA for the ATC environment.
摘要Traditional Chinese medicine(TCM)demonstrates distinctive advantages in disease prevention and treatment.However,analyzing its biological mechanisms through the modern medical research paradigm of“single drug,single target”presents significant challenges due to its holistic approach.Network pharmacology and its core theory of network targets connect drugs and diseases from a holistic and systematic perspective based on biological networks,overcoming the limitations of reductionist research models and showing considerable value in TCM research.Recent integration of network target computational and experimental methods with artificial intelligence(AI)and multi-modal multi-omics technologies has substantially enhanced network pharmacology methodology.The advancement in computational and experimental techniques provides complementary support for network target theory in decoding TCM principles.This review,centered on network targets,examines the progress of network target methods combined with AI in predicting disease molecular mechanisms and drug-target relationships,alongside the application of multi-modal multi-omics technologies in analyzing TCM formulae,syndromes,and toxicity.Looking forward,network target theory is expected to incorporate emerging technologies while developing novel approaches aligned with its unique characteristics,potentially leading to significant breakthroughs in TCM research and advancing scientific understanding and innovation in TCM.
基金supported by the National Natural Science Foundation of China(Grant Nos.62071315 and 62271336).
摘要The multi-modal characteristics of mineral particles play a pivotal role in enhancing the classification accuracy,which is critical for obtaining a profound understanding of the Earth's composition and ensuring effective exploitation utilization of its resources.However,the existing methods for classifying mineral particles do not fully utilize these multi-modal features,thereby limiting the classification accuracy.Furthermore,when conventional multi-modal image classification methods are applied to planepolarized and cross-polarized sequence images of mineral particles,they encounter issues such as information loss,misaligned features,and challenges in spatiotemporal feature extraction.To address these challenges,we propose a multi-modal mineral particle polarization image classification network(MMGC-Net)for precise mineral particle classification.Initially,MMGC-Net employs a two-dimensional(2D)backbone network with shared parameters to extract features from two types of polarized images to ensure feature alignment.Subsequently,a cross-polarized intra-modal feature fusion module is designed to refine the spatiotemporal features from the extracted features of the cross-polarized sequence images.Ultimately,the inter-modal feature fusion module integrates the two types of modal features to enhance the classification precision.Quantitative and qualitative experimental results indicate that when compared with the current state-of-the-art multi-modal image classification methods,MMGC-Net demonstrates marked superiority in terms of mineral particle multi-modal feature learning and four classification evaluation metrics.It also demonstrates better stability than the existing models.
基金partially supported by the National Natural Science Foundation of China under Grants 62471493 and 62402257(for conceptualization and investigation)partially supported by the Natural Science Foundation of Shandong Province,China under Grants ZR2023LZH017,ZR2024MF066,and 2023QF025(for formal analysis and validation)+1 种基金partially supported by the Open Foundation of Key Laboratory of Computing Power Network and Information Security,Ministry of Education,Qilu University of Technology(Shandong Academy of Sciences)under Grant 2023ZD010(for methodology and model design)partially supported by the Russian Science Foundation(RSF)Project under Grant 22-71-10095-P(for validation and results verification).
摘要To address the challenge of missing modal information in entity alignment and to mitigate information loss or bias arising frommodal heterogeneity during fusion,while also capturing shared information acrossmodalities,this paper proposes a Multi-modal Pre-synergistic Entity Alignmentmodel based on Cross-modalMutual Information Strategy Optimization(MPSEA).The model first employs independent encoders to process multi-modal features,including text,images,and numerical values.Next,a multi-modal pre-synergistic fusion mechanism integrates graph structural and visual modal features into the textual modality as preparatory information.This pre-fusion strategy enables unified perception of heterogeneous modalities at the model’s initial stage,reducing discrepancies during the fusion process.Finally,using cross-modal deep perception reinforcement learning,the model achieves adaptive multilevel feature fusion between modalities,supporting learningmore effective alignment strategies.Extensive experiments on multiple public datasets show that the MPSEA method achieves gains of up to 7% in Hits@1 and 8.2% in MRR on the FBDB15K dataset,and up to 9.1% in Hits@1 and 7.7% in MRR on the FBYG15K dataset,compared to existing state-of-the-art methods.These results confirm the effectiveness of the proposed model.