In this paper,we propose a new privacy-aware transmission scheduling algorithm for 6G ad hoc networks.This system enables end nodes to select the optimum time and scheme to transmit private data safely.In 6G dynamic h...In this paper,we propose a new privacy-aware transmission scheduling algorithm for 6G ad hoc networks.This system enables end nodes to select the optimum time and scheme to transmit private data safely.In 6G dynamic heterogeneous infrastructures,unstable links and non-uniform hardware capabilities create critical issues regarding security and privacy.Traditional protocols are often too computationally heavy to allow 6G services to achieve their expected Quality-of-Service(QoS).As the transport network is built of ad hoc nodes,there is no guarantee about their trustworthiness or behavior,and transversal functionalities are delegated to the extreme nodes.However,while security can be guaranteed in extreme-to-extreme solutions,privacy cannot,as all intermediate nodes still have to handle the data packets they are transporting.Besides,traditional schemes for private anonymous ad hoc communications are vulnerable against modern intelligent attacks based on learning models.The proposed scheme fulfills this gap.Findings show the probability of a successful intelligent attack reduces by up to 65%compared to ad hoc networks with no privacy protection strategy when used the proposed technology.While congestion probability can remain below 0.001%,as required in 6G services.展开更多
The collection and annotation of lar ge-scale bird datasets are resource-intensive and time-consuming processes that significantly limit the scalability and accuracy of biodiversity monitoring systems.While self-super...The collection and annotation of lar ge-scale bird datasets are resource-intensive and time-consuming processes that significantly limit the scalability and accuracy of biodiversity monitoring systems.While self-supervised learning(SSL)has emerged as a promising approach for leveraging unannotated data,current SSL methods face two critical challenges in bird species recognition:(1)long-tailed data distributions that result in poor performance on underrepresented species;and(2)domain shift issues caused by data augmentation strategies designed to mitigate class imbalance.Here we present SDNet,a novel SSL-based bird recognition framework that integrates diffusion models with large language models(LLMs)to overcome these limitations.SDNet employs LLMs to generate semantically rich textual descriptions for tail-class species by prompting the models with species taxonomy,morphological attributes,and habitat information,producing detailed natural language priors that capture fine-grained visual characteristics(e.g.,plumage patterns,body proportions,and distinctive markings).These textual descriptions are subsequently used by a conditional diffusion model to synthesize new bird image samples through cross-attention mechanisms that fuse textual embeddings with intermediate visual feature representations during the denoising process,ensuring generated images preserve species-specific morphological details while maintaining photorealistic quality.Additionally,we incorporate a Swin Transformer as the feature extraction backbone whose hierarchical window-based attention mechanism and shifted windowing scheme enable multi-scale local feature extraction that proves particularly effective at capturing finegrained discriminative patterns(such as beak shape and feather texture)while mitigating domain shift between synthetic and original images through consistent feature representations across both data sources.SDNet is validated on both a self-constructed dataset(Bird_BXS)an d a publicly available benchmark(Birds_25),demonstrating substantial improvements over conventional SSL approaches.Our results indicate that the synergistic integration of LLMs,diffusion models,and the Swin Transformer architecture contributes significantly to recognition accuracy,particularly for rare and morphologically similar species.These findings highlight the potential of SDNet for addressing fundamental limitations of existing SSL methods in avian recognition tasks and establishing a new paradigm for efficient self-supervised learning in large-scale ornithological vision applications.展开更多
Diffusion models have emerged as powerful generative tools.Recently,numerous image watermarking techniques have been developed to ensure copyright protection and content traceability.These techniques also play a vital...Diffusion models have emerged as powerful generative tools.Recently,numerous image watermarking techniques have been developed to ensure copyright protection and content traceability.These techniques also play a vital role in multimedia forensics,where digital watermarks are used to verify authenticity,detect tampering,and support legal attribution.Conventional watermark removal techniques operate at the image level,employing various post-processing strategies to degrade embedded watermarks.However,these methods often incur substantial computational costs,limiting their practical applicability.In this work,we propose a novel watermark removal approach targeting the diffusion model itself,specifically focusing on fine-tuning the Variational Autoencoder(VAE)decoder.Utilising Low-Rank Adaptation(LoRA),we efficiently optimise a subset of parameters to suppress watermark signals while preserving image fidelity.An adversarial classifier is incorporated to guide the training process,ensuring effective watermark removal.Furthermore,perceptual and structural similarity losses are integrated to mitigate degradation in image quality.Experimental results show that our approach effectively eliminates embedded watermarks,providing a robust and lightweight solution for watermark removal in diffusion-based generative models.This work also provides a technical perspective for analysing the reliability and resilience of watermark-based mechanisms in generative AI systems,offering insight into their potential limitations in forensic identification.展开更多
Image colorization has attracted considerable research interest over the past few decades.However,current methodologies frequently struggle with limited local colorization flexibility and produce unnatural color outpu...Image colorization has attracted considerable research interest over the past few decades.However,current methodologies frequently struggle with limited local colorization flexibility and produce unnatural color outputs,primarily due to the absence of comprehensive understanding of color perception.In this work,we propose an expressive diffusion network(EDN)that leverages a robust diffusion network to significantly enhance both colorization accuracy and diversity.The EDN consists of two main components:a pre-trained latent diffusion model and a perceptual luminance model based on VQ-Diffusion.These components work together to generate rich and vibrant colors while maintaining high fidelity to the structural features of the original grayscale image.The EDN incorporates controllable creative diffusion(CCD)to direct the color generation process toward more realistic outcomes.Extensive experiments demonstrate that the EDN outperforms existing methods in perceptual quality,offering notable improvements in visual realism and vibrancy across various scenes.The proposed EDN showcases significant improvements over ChromaGAN and InstColor,confirming its robustness in both simple and complex scenarios.展开更多
Traditional steganography conceals information by modifying cover data,but steganalysis tools easily detect such alterations.While deep learning-based steganography often involves high training costs and complex deplo...Traditional steganography conceals information by modifying cover data,but steganalysis tools easily detect such alterations.While deep learning-based steganography often involves high training costs and complex deployment.Diffusion model-based methods face security vulnerabilities,particularly due to potential information leakage during generation.We propose a fixed neural network image steganography framework based on secure diffu-sion models to address these challenges.Unlike conventional approaches,our method minimizes cover modifications through neural network optimization,achieving superior steganographic performance in human visual perception and computer vision analyses.The cover images are generated in an anime style using state-of-the-art diffusion models,ensuring the transmitted images appear more natural.This study introduces fixed neural network technology that allows senders to transmit only minimal critical information alongside stego-images.Recipients can accurately reconstruct secret images using this compact data,significantly reducing transmission overhead compared to conventional deep steganography.Furthermore,our framework innovatively integrates ElGamal,a cryptographic algorithm,to protect critical information during transmission,enhancing overall system security and ensuring end-to-end information protection.This dual optimization of payload reduction and cryptographic reinforcement establishes a new paradigm for secure and efficient image steganography.展开更多
Air target intent recognition holds significant importance in aiding commanders to assess battlefield situations and secure a competitive edge in decision-making.Progress in this domain has been hindered by challenges...Air target intent recognition holds significant importance in aiding commanders to assess battlefield situations and secure a competitive edge in decision-making.Progress in this domain has been hindered by challenges posed by imbalanced battlefield data and the limited robustness of traditional recognition models.Inspired by the success of diffusion models in addressing visual domain sample imbalances,this paper introduces a new approach that utilizes the Markov Transfer Field(MTF)method for time series data visualization.This visualization,when combined with the Denoising Diffusion Probabilistic Model(DDPM),effectively enhances sample data and mitigates noise within the original dataset.Additionally,a transformer-based model tailored for time series visualization and air target intent recognition is developed.Comprehensive experimental results,encompassing comparative,ablation,and denoising validations,reveal that the proposed method achieves a notable 98.86%accuracy in air target intent recognition while demonstrating exceptional robustness and generalization capabilities.This approach represents a promising avenue for advancing air target intent recognition.展开更多
High-Resolution(HR)data on flow fields are critical for accurately evaluating the aerodynamic performance of aircraft.However,acquiring such data through large-scale numerical simulations or wind tunnel experiments is...High-Resolution(HR)data on flow fields are critical for accurately evaluating the aerodynamic performance of aircraft.However,acquiring such data through large-scale numerical simulations or wind tunnel experiments is highly resource intensive.This paper proposes a FlowViT-Diff framework that integrates a Vision Transformer(ViT)with an enhanced denoising diffusion probabilistic model for the Super-Resolution(SR)reconstruction of HR flow fields based on low-resolution inputs.It provides a quick initial prediction of the HR flow field by optimizing the ViT architecture,and incorporates this preliminary output as guidance within an enhanced diffusion model.The latter captures the Gaussian noise distribution during forward diffusion and progressively removes it during backward diffusion to generate the flow field.Experiments on various supercritical airfoils under different flow conditions show that FlowViT-Diff can robustly reconstruct the flow field across multiple levels of downsampling.It obtains more consistent global and local features than traditional SR methods,and yields a 3.6-fold increase in its training speed via transfer learning.Its accuracy of reconstruction of the flow field is 99.7%under ultra-low downsampling.The results demonstrate that Flow Vi T-Diff not only exhibits effective flow field reconstruction capabilities,but also provides two reconstruction strategies,both of which show effective transferability.展开更多
Obtaining unsteady hydrodynamic performance is of great significance for seaplane design.Common methods for obtaining unsteady hydrodynamic performance data include tank test and Computational Fluid Dynamics(CFD)numer...Obtaining unsteady hydrodynamic performance is of great significance for seaplane design.Common methods for obtaining unsteady hydrodynamic performance data include tank test and Computational Fluid Dynamics(CFD)numerical simulation,which are costly and time-consuming.Therefore,it is necessary to obtain unsteady hydrodynamic performance in a low-cost and high-precision manner.Due to the strong nonlinearity,complex data distribution,and temporal characteristics of unsteady hydrodynamic performance,the prediction of it is challenging.This paper proposes a Temporal Convolutional Diffusion Model(TCDM)for predicting the unsteady hydrodynamic performance of seaplanes given design parameters.Under the framework of a classifier-free guided diffusion model,TCDM learns the distribution patterns of unsteady hydrodynamic performance data with the designed denoising module based on temporal convolutional network and captures the temporal features of unsteady hydrodynamic performance data.Using CFD simulation data,the proposed method is compared with the alternative methods to demonstrate its accuracy and generalization.This paper provides a method that enables the rapid and accurate prediction of unsteady hydrodynamic performance data,expecting to shorten the design cycle of seaplanes.展开更多
Transformer models have emerged as dominant networks for various tasks in computer vision compared to Convolutional Neural Networks(CNNs).The transformers demonstrate the ability to model long-range dependencies by ut...Transformer models have emerged as dominant networks for various tasks in computer vision compared to Convolutional Neural Networks(CNNs).The transformers demonstrate the ability to model long-range dependencies by utilizing a self-attention mechanism.This study aims to provide a comprehensive survey of recent transformerbased approaches in image and video applications,as well as diffusion models.We begin by discussing existing surveys of vision transformers and comparing them to this work.Then,we review the main components of a vanilla transformer network,including the self-attention mechanism,feed-forward network,position encoding,etc.In the main part of this survey,we review recent transformer-based models in three categories:Transformer for downstream tasks,Vision Transformer for Generation,and Vision Transformer for Segmentation.We also provide a comprehensive overview of recent transformer models for video tasks and diffusion models.We compare the performance of various hierarchical transformer networks for multiple tasks on popular benchmark datasets.Finally,we explore some future research directions to further improve the field.展开更多
Multiple-input multiple-output(MIMO)systems are essential for improving capacity and reliability in semantic communications.Existing methods mainly design the channel-aware neural networks but neglect the underlying s...Multiple-input multiple-output(MIMO)systems are essential for improving capacity and reliability in semantic communications.Existing methods mainly design the channel-aware neural networks but neglect the underlying signal distribution.In this paper,we develop a denoising diffusion null-space model-based module over MIMO channels(DDNM-MIMO),which is a plug-in module deployed at the receiver.By modeling the MIMO channel,precoding,and equalization as a linear transformation with additive noise,we design corresponding linear and scaling matrices to construct a sampling process for denoising the received signal.The DDNM-MIMO integrates channel state information(CSI)embedding,supporting both closed-loop MIMO with CSI at the transmitter and open-loop MIMO with CSI at the receiver,thereby improving channel adaptability across various noise levels.As a plug-in,the DDNM-MIMO module operates independently of the joint source-channel coding(JSCC)coder structure,offering flexible integration into diverse systems.Experimental results show that DDNM-MIMO effectively reduces the mean square errors(MSE)between the encoded and equalized signals.Consequently,the proposed DDNM-MIMO semantic communication system achieves superior image reconstruction performance compared to existing JSCC-based semantic communication method.展开更多
Nuclear magnetic resonance(NMR)spectroscopy is a key method for molecular structure elucidation.However,interpreting NMR spectra to deduce molecular structures remains challenging due to the complexity of spectral dat...Nuclear magnetic resonance(NMR)spectroscopy is a key method for molecular structure elucidation.However,interpreting NMR spectra to deduce molecular structures remains challenging due to the complexity of spectral data and the vastness of the chemical space.Here we introduce DiffNMR,a novel end-to-end framework that leverages a conditional discrete diffusion model for de novo molecular structure elucidation from NMR spectra.DiffNMR refines molecular graphs iteratively through a diffusion-based generative process,ensuring global consistency and mitigating error accumulation inherent in autoregressive methods.The framework integrates a two-stage pretraining strategy that aligns spectral and molecular representations via a diffusion autoencoder and contrastive learning.It also incorporates retrieval initialization and similarity filtering during inference.Our experimental results demonstrate that DiffNMR achieves competitive performance for NMR-based structure elucidation,especially outperforming autoregressive models in domain generalization and robustness,thereby offering an efficient and robust solution for automated molecular analysis.展开更多
Recent advancements in diffusion models have significantly impacted content creation,leading to the emergence of per-sonalized content synthesis(PCS).By utilizing a small set of user-provided examples featuring the sa...Recent advancements in diffusion models have significantly impacted content creation,leading to the emergence of per-sonalized content synthesis(PCS).By utilizing a small set of user-provided examples featuring the same subject,PCS aims to tailor this subject to specific user-defined prompts.Over the past two years,more than 150 methods have been introduced in this area.However,existing surveys primarily focus on text-to-image generation,with few providing up-to-date summaries on PCS.This pa-per provides a comprehensive survey of PCS,introducing the general frameworks of PCS research,which can be categorized into test-time fine-tuning(TTF)and pre-trained adaptation(PTA)approaches.We analyze the strengths,limitations and key tech-niques of these methodologies.Additionally,we explore specialized tasks within the field,such as object,face and style personaliza-tion,while highlighting their unique challenges and innovations.Despite the promising progress,we also discuss ongoing challenges,including overfitting and the trade-off between subject fidelity and text alignment.Through this detailed overview and analysis,we propose future directions to further the development of PCS.展开更多
Denoising diffusion models have demonstrated tremendous success in modeling data distributions and synthesizing high-quality samples.In the 2D image domain,they have become the state-of-the-art and are capable of gene...Denoising diffusion models have demonstrated tremendous success in modeling data distributions and synthesizing high-quality samples.In the 2D image domain,they have become the state-of-the-art and are capable of generating photo-realistic images with high controllability.More recently,researchers have begun to explore how to utilize diffusion models to generate 3D data,as doing so has more potential in real-world applications.This requires careful design choices in two key ways:identifying a suitable 3D representation and determining how to apply the diffusion process.In this survey,we provide the first comprehensive review of diffusion models for manipulating 3D content,including 3D generation,reconstruction,and 3D-aware image synthesis.We classify existing methods into three major categories:2D space diffusion with pretrained models,2D space diffusion without pretrained models,and 3D space diffusion.We also summarize popular datasets used for 3D generation with diffusion models.Along with this survey,we maintain a repository http://gffzz188fe103f8f1460askfcbcfpb6fv9696k.ffgz.tsg.suse.edu.cn/cwchenwang/awesome-3d-diffusion to track the latest relevant papers and codebases.Finally,we pose current challenges for diffusion models for 3D generation,and suggest future research directions.展开更多
Diffusion models are a type of generative deep learning model that can process medical images more efficiently than traditional generative models.They have been applied to several medical image computing tasks.This pa...Diffusion models are a type of generative deep learning model that can process medical images more efficiently than traditional generative models.They have been applied to several medical image computing tasks.This paper aims to help researchers understand the advancements of diffusion models in medical image computing.It begins by describing the fundamental principles,sampling methods,and architecture of diffusion models.Subsequently,it discusses the application of diffusion models in five medical image computing tasks:image generation,modality conversion,image segmentation,image denoising,and anomaly detection.Additionally,this paper conducts fine-tuning of a large model for image generation tasks and comparative experiments between diffusion models and traditional generative models across these five tasks.The evaluation of the fine-tuned large model shows its potential for clinical applications.Comparative experiments demonstrate that diffusion models have a distinct advantage in tasks related to image generation,modality conversion,and image denoising.However,they require further optimization in image segmentation and anomaly detection tasks to match the efficacy of traditional models.Our codes are publicly available at:http://gffzz188fe103f8f1460askfcbcfpb6fv9696k.ffgz.tsg.suse.edu.cn/hiahub/CodeForDiffusion.展开更多
The detection of zero-day malware represents one of the most significant challenges in contemporary cybersecurity.In this paper,we introduce a novel concept called“Negative-One-Day Malware Detection”,which aims to i...The detection of zero-day malware represents one of the most significant challenges in contemporary cybersecurity.In this paper,we introduce a novel concept called“Negative-One-Day Malware Detection”,which aims to identify potentially malicious software before it is actually created by threat actors.Our approach leverages recent advancements in generative AI,specifically diffusion-based generative models,to generate and analyze potential future malware variants.By doing so,we can train detection systems to recognize these variants before they emerge in the wild,thereby closing the critical protection gap that currently exists between malware creation and detection.We demonstrate the effectiveness of our approach through extensive experimentation,showing that our framework can generate executable malware samples that combine characteristics from different families while exhibiting novel behaviors.These synthetically generated samples significantly improve the detection capabilities of security systems when incorporated into training data,providing a proactive rather than reactive approach to cybersecurity.展开更多
Crack detection accuracy in computer vision is often constrained by limited annotated datasets.Although Generative Adversarial Networks(GANs)have been applied for data augmentation,they frequently introduce blurs and ...Crack detection accuracy in computer vision is often constrained by limited annotated datasets.Although Generative Adversarial Networks(GANs)have been applied for data augmentation,they frequently introduce blurs and artifacts.To address this challenge,this study leverages Denoising Diffusion Probabilistic Models(DDPMs)to generate high-quality synthetic crack images,enriching the training set with diverse and structurally consistent samples that enhance the crack segmentation.The proposed framework involves a two-stage pipeline:first,DDPMs are used to synthesize high-fidelity crack images that capture fine structural details.Second,these generated samples are combined with real data to train segmentation networks,thereby improving accuracy and robustness in crack detection.Compared with GAN-based approaches,DDPM achieved the best fidelity,with the highest Structural Similarity Index(SSIM)(0.302)and lowest Learned Perceptual Image Patch Similarity(LPIPS)(0.461),producing artifact-free images that preserve fine crack details.To validate its effectiveness,six segmentation models were tested,among which LinkNet consistently achieved the best performance,excelling in both region-level accuracy and structural continuity.Incorporating DDPM-augmented data further enhanced segmentation outcomes,increasing F1 scores by up to 1.1%and IoU by 1.7%,while also improving boundary alignment and skeleton continuity compared with models trained on real images alone.Experiments with varying augmentation ratios showed consistent improvements,with F1 rising from 0.946(no augmentation)to 0.957 and IoU from 0.897 to 0.913 at the highest ratio.These findings demonstrate the effectiveness of diffusion-based augmentation for complex crack detection in structural health monitoring.展开更多
Accurately identifying building distribution from remote sensing images with complex background information is challenging.The emergence of diffusion models has prompted the innovative idea of employing the reverse de...Accurately identifying building distribution from remote sensing images with complex background information is challenging.The emergence of diffusion models has prompted the innovative idea of employing the reverse denoising process to distill building distribution from these complex backgrounds.Building on this concept,we propose a novel framework,building extraction diffusion model(BEDiff),which meticulously refines the extraction of building footprints from remote sensing images in a stepwise fashion.Our approach begins with the design of booster guidance,a mechanism that extracts structural and semantic features from remote sensing images to serve as priors,thereby providing targeted guidance for the diffusion process.Additionally,we introduce a cross-feature fusion module(CFM)that bridges the semantic gap between different types of features,facilitating the integration of the attributes extracted by booster guidance into the diffusion process more effectively.Our proposed BEDiff marks the first application of diffusion models to the task of building extraction.Empirical evidence from extensive experiments on the Beijing building dataset demonstrates the superior performance of BEDiff,affirming its effectiveness and potential for enhancing the accuracy of building extraction in complex urban landscapes.展开更多
AlphaPanda(AlphaFold2[1]inspired protein-specific antibody design in a diffusional manner)is an advanced algorithm for designing complementary determining regions(CDRs)of the antibody targeted the specific epitope,com...AlphaPanda(AlphaFold2[1]inspired protein-specific antibody design in a diffusional manner)is an advanced algorithm for designing complementary determining regions(CDRs)of the antibody targeted the specific epitope,combining transformer[2]models,3DCNN[3],and diffusion[4]generative models.展开更多
The application of generative artificial intelligence(AI)is bringing about notable changes in anime creation.This paper surveys recent advancements and applications of diffusion and language models in anime generation...The application of generative artificial intelligence(AI)is bringing about notable changes in anime creation.This paper surveys recent advancements and applications of diffusion and language models in anime generation,focusing on their demonstrated potential to enhance production efficiency through automation and personalization.Despite these benefits,it is crucial to acknowledge the substantial initial computational investments required for training and deploying these models.We conduct an in-depth survey of cutting-edge generative AI technologies,encompassing models such as Stable Diffusion and GPT,and appraise pivotal large-scale datasets alongside quantifiable evaluation metrics.Review of the surveyed literature indicates the achievement of considerable maturity in the capacity of AI models to synthesize high-quality,aesthetically compelling anime visual images from textual prompts,alongside discernible progress in the generation of coherent narratives.However,achieving perfect long-form consistency,mitigating artifacts like flickering in video sequences,and enabling fine-grained artistic control remain critical ongoing challenges.Building upon these advancements,research efforts have increasingly pivoted towards the synthesis of higher-dimensional content,such as video and three-dimensional assets,with recent studies demonstrating significant progress in this burgeoning field.Nevertheless,formidable challenges endure amidst these advancements.Foremost among these are the substantial computational exigencies requisite for training and deploying these sophisticated models,particularly pronounced in the realm of high-dimensional generation such as video synthesis.Additional persistent hurdles include maintaining spatial-temporal consistency across complex scenes and mitigating ethical considerations surrounding bias and the preservation of human creative autonomy.This research underscores the transformative potential and inherent complexities of AI-driven synergy within the creative industries.We posit that future research should be dedicated to the synergistic fusion of diffusion and autoregressive models,the integration of multimodal inputs,and the balanced consideration of ethical implications,particularly regarding bias and the preservation of human creative autonomy,thereby establishing a robust foundation for the advancement of anime creation and the broader landscape of AI-driven content generation.展开更多
The application of machine learning in fluid dynamics has become increasingly prevalent for accelerating computations in solving forward and inverse problems governed by partial differential equations,particularly flo...The application of machine learning in fluid dynamics has become increasingly prevalent for accelerating computations in solving forward and inverse problems governed by partial differential equations,particularly flow field reconstruction.However,existing end-to-end methods often depend heavily on specific low-fidelity patterns or sparsity rates during training,limiting their effectiveness in real-world scenarios where inputs may deviate from the training distribution or contain unanticipated noise.Diffusion models present a promising alternative by learning to transform various low-fidelity distributions into high-fidelity ones,offering greater flexibility than direct mapping approaches that are typically constrained to single problem types.In this work,we introduce the physics-informed residual diffusion model,which generalizes across diverse inputs,including evenly down-sampled sensor data,sparse sensor data with Gaussian noise,and randomly sampled sparse sensor data.By incorporating partial differential equation constraints into the objective function,our approach significantly improves the accuracy of reconstructed high-fidelity flow fields while ensuring adherence to underlying physical laws.Experimental results demonstrate that our model successfully generates high-quality outcomes for two-dimensional Navier-Stokes equation flow and Kolmogorov flow under varied low-fidelity input conditions,without the need for retraining.展开更多
基金funding from the European Commission by the Ruralities project(grant agreement no.101060876).
摘要In this paper,we propose a new privacy-aware transmission scheduling algorithm for 6G ad hoc networks.This system enables end nodes to select the optimum time and scheme to transmit private data safely.In 6G dynamic heterogeneous infrastructures,unstable links and non-uniform hardware capabilities create critical issues regarding security and privacy.Traditional protocols are often too computationally heavy to allow 6G services to achieve their expected Quality-of-Service(QoS).As the transport network is built of ad hoc nodes,there is no guarantee about their trustworthiness or behavior,and transversal functionalities are delegated to the extreme nodes.However,while security can be guaranteed in extreme-to-extreme solutions,privacy cannot,as all intermediate nodes still have to handle the data packets they are transporting.Besides,traditional schemes for private anonymous ad hoc communications are vulnerable against modern intelligent attacks based on learning models.The proposed scheme fulfills this gap.Findings show the probability of a successful intelligent attack reduces by up to 65%compared to ad hoc networks with no privacy protection strategy when used the proposed technology.While congestion probability can remain below 0.001%,as required in 6G services.
基金supported by the National Natural Science Foundation of China(32471964)。
摘要The collection and annotation of lar ge-scale bird datasets are resource-intensive and time-consuming processes that significantly limit the scalability and accuracy of biodiversity monitoring systems.While self-supervised learning(SSL)has emerged as a promising approach for leveraging unannotated data,current SSL methods face two critical challenges in bird species recognition:(1)long-tailed data distributions that result in poor performance on underrepresented species;and(2)domain shift issues caused by data augmentation strategies designed to mitigate class imbalance.Here we present SDNet,a novel SSL-based bird recognition framework that integrates diffusion models with large language models(LLMs)to overcome these limitations.SDNet employs LLMs to generate semantically rich textual descriptions for tail-class species by prompting the models with species taxonomy,morphological attributes,and habitat information,producing detailed natural language priors that capture fine-grained visual characteristics(e.g.,plumage patterns,body proportions,and distinctive markings).These textual descriptions are subsequently used by a conditional diffusion model to synthesize new bird image samples through cross-attention mechanisms that fuse textual embeddings with intermediate visual feature representations during the denoising process,ensuring generated images preserve species-specific morphological details while maintaining photorealistic quality.Additionally,we incorporate a Swin Transformer as the feature extraction backbone whose hierarchical window-based attention mechanism and shifted windowing scheme enable multi-scale local feature extraction that proves particularly effective at capturing finegrained discriminative patterns(such as beak shape and feather texture)while mitigating domain shift between synthetic and original images through consistent feature representations across both data sources.SDNet is validated on both a self-constructed dataset(Bird_BXS)an d a publicly available benchmark(Birds_25),demonstrating substantial improvements over conventional SSL approaches.Our results indicate that the synergistic integration of LLMs,diffusion models,and the Swin Transformer architecture contributes significantly to recognition accuracy,particularly for rare and morphologically similar species.These findings highlight the potential of SDNet for addressing fundamental limitations of existing SSL methods in avian recognition tasks and establishing a new paradigm for efficient self-supervised learning in large-scale ornithological vision applications.
基金supported by the Public Interest Research Grant Programs of National Research Institutes[grant number GY2024G-6]the National Natural Science Foundation of China[grant numbers 62261160653 and 62441237].
摘要Diffusion models have emerged as powerful generative tools.Recently,numerous image watermarking techniques have been developed to ensure copyright protection and content traceability.These techniques also play a vital role in multimedia forensics,where digital watermarks are used to verify authenticity,detect tampering,and support legal attribution.Conventional watermark removal techniques operate at the image level,employing various post-processing strategies to degrade embedded watermarks.However,these methods often incur substantial computational costs,limiting their practical applicability.In this work,we propose a novel watermark removal approach targeting the diffusion model itself,specifically focusing on fine-tuning the Variational Autoencoder(VAE)decoder.Utilising Low-Rank Adaptation(LoRA),we efficiently optimise a subset of parameters to suppress watermark signals while preserving image fidelity.An adversarial classifier is incorporated to guide the training process,ensuring effective watermark removal.Furthermore,perceptual and structural similarity losses are integrated to mitigate degradation in image quality.Experimental results show that our approach effectively eliminates embedded watermarks,providing a robust and lightweight solution for watermark removal in diffusion-based generative models.This work also provides a technical perspective for analysing the reliability and resilience of watermark-based mechanisms in generative AI systems,offering insight into their potential limitations in forensic identification.
摘要Image colorization has attracted considerable research interest over the past few decades.However,current methodologies frequently struggle with limited local colorization flexibility and produce unnatural color outputs,primarily due to the absence of comprehensive understanding of color perception.In this work,we propose an expressive diffusion network(EDN)that leverages a robust diffusion network to significantly enhance both colorization accuracy and diversity.The EDN consists of two main components:a pre-trained latent diffusion model and a perceptual luminance model based on VQ-Diffusion.These components work together to generate rich and vibrant colors while maintaining high fidelity to the structural features of the original grayscale image.The EDN incorporates controllable creative diffusion(CCD)to direct the color generation process toward more realistic outcomes.Extensive experiments demonstrate that the EDN outperforms existing methods in perceptual quality,offering notable improvements in visual realism and vibrancy across various scenes.The proposed EDN showcases significant improvements over ChromaGAN and InstColor,confirming its robustness in both simple and complex scenarios.
基金supported in part by the National Natural Science Foundation of China under Grants 62102450,62272478 and the Independent Research Project of a Certain Unit under Grant ZZKY20243127。
摘要Traditional steganography conceals information by modifying cover data,but steganalysis tools easily detect such alterations.While deep learning-based steganography often involves high training costs and complex deployment.Diffusion model-based methods face security vulnerabilities,particularly due to potential information leakage during generation.We propose a fixed neural network image steganography framework based on secure diffu-sion models to address these challenges.Unlike conventional approaches,our method minimizes cover modifications through neural network optimization,achieving superior steganographic performance in human visual perception and computer vision analyses.The cover images are generated in an anime style using state-of-the-art diffusion models,ensuring the transmitted images appear more natural.This study introduces fixed neural network technology that allows senders to transmit only minimal critical information alongside stego-images.Recipients can accurately reconstruct secret images using this compact data,significantly reducing transmission overhead compared to conventional deep steganography.Furthermore,our framework innovatively integrates ElGamal,a cryptographic algorithm,to protect critical information during transmission,enhancing overall system security and ensuring end-to-end information protection.This dual optimization of payload reduction and cryptographic reinforcement establishes a new paradigm for secure and efficient image steganography.
基金co-supported by the National Natural Science Foundation of China(Nos.61806219,61876189 and 61703426)the Young Talent Fund of University Association for Science and Technology in Shaanxi,China(Nos.20190108 and 20220106)the Innvation Talent Supporting Project of Shaanxi,China(No.2020KJXX-065)。
摘要Air target intent recognition holds significant importance in aiding commanders to assess battlefield situations and secure a competitive edge in decision-making.Progress in this domain has been hindered by challenges posed by imbalanced battlefield data and the limited robustness of traditional recognition models.Inspired by the success of diffusion models in addressing visual domain sample imbalances,this paper introduces a new approach that utilizes the Markov Transfer Field(MTF)method for time series data visualization.This visualization,when combined with the Denoising Diffusion Probabilistic Model(DDPM),effectively enhances sample data and mitigates noise within the original dataset.Additionally,a transformer-based model tailored for time series visualization and air target intent recognition is developed.Comprehensive experimental results,encompassing comparative,ablation,and denoising validations,reveal that the proposed method achieves a notable 98.86%accuracy in air target intent recognition while demonstrating exceptional robustness and generalization capabilities.This approach represents a promising avenue for advancing air target intent recognition.
基金supported by the National Natural Science Foundation of China(No.12472265)。
摘要High-Resolution(HR)data on flow fields are critical for accurately evaluating the aerodynamic performance of aircraft.However,acquiring such data through large-scale numerical simulations or wind tunnel experiments is highly resource intensive.This paper proposes a FlowViT-Diff framework that integrates a Vision Transformer(ViT)with an enhanced denoising diffusion probabilistic model for the Super-Resolution(SR)reconstruction of HR flow fields based on low-resolution inputs.It provides a quick initial prediction of the HR flow field by optimizing the ViT architecture,and incorporates this preliminary output as guidance within an enhanced diffusion model.The latter captures the Gaussian noise distribution during forward diffusion and progressively removes it during backward diffusion to generate the flow field.Experiments on various supercritical airfoils under different flow conditions show that FlowViT-Diff can robustly reconstruct the flow field across multiple levels of downsampling.It obtains more consistent global and local features than traditional SR methods,and yields a 3.6-fold increase in its training speed via transfer learning.Its accuracy of reconstruction of the flow field is 99.7%under ultra-low downsampling.The results demonstrate that Flow Vi T-Diff not only exhibits effective flow field reconstruction capabilities,but also provides two reconstruction strategies,both of which show effective transferability.
基金supported by the Aeronautical Science Foundation of China(Nos.2018ZA52002,2019ZA052011)the National Natural Science Foundation of China(No.12472236).
摘要Obtaining unsteady hydrodynamic performance is of great significance for seaplane design.Common methods for obtaining unsteady hydrodynamic performance data include tank test and Computational Fluid Dynamics(CFD)numerical simulation,which are costly and time-consuming.Therefore,it is necessary to obtain unsteady hydrodynamic performance in a low-cost and high-precision manner.Due to the strong nonlinearity,complex data distribution,and temporal characteristics of unsteady hydrodynamic performance,the prediction of it is challenging.This paper proposes a Temporal Convolutional Diffusion Model(TCDM)for predicting the unsteady hydrodynamic performance of seaplanes given design parameters.Under the framework of a classifier-free guided diffusion model,TCDM learns the distribution patterns of unsteady hydrodynamic performance data with the designed denoising module based on temporal convolutional network and captures the temporal features of unsteady hydrodynamic performance data.Using CFD simulation data,the proposed method is compared with the alternative methods to demonstrate its accuracy and generalization.This paper provides a method that enables the rapid and accurate prediction of unsteady hydrodynamic performance data,expecting to shorten the design cycle of seaplanes.
基金supported in part by the National Natural Science Foundation of China under Grants 61502162,61702175,and 61772184in part by the Fund of the State Key Laboratory of Geo-information Engineering under Grant SKLGIE2016-M-4-2+4 种基金in part by the Hunan Natural Science Foundation of China under Grant 2018JJ2059in part by the Key R&D Project of Hunan Province of China under Grant 2018GK2014in part by the Open Fund of the State Key Laboratory of Integrated Services Networks under Grant ISN17-14Chinese Scholarship Council(CSC)through College of Computer Science and Electronic Engineering,Changsha,410082Hunan University with Grant CSC No.2018GXZ020784.
摘要Transformer models have emerged as dominant networks for various tasks in computer vision compared to Convolutional Neural Networks(CNNs).The transformers demonstrate the ability to model long-range dependencies by utilizing a self-attention mechanism.This study aims to provide a comprehensive survey of recent transformerbased approaches in image and video applications,as well as diffusion models.We begin by discussing existing surveys of vision transformers and comparing them to this work.Then,we review the main components of a vanilla transformer network,including the self-attention mechanism,feed-forward network,position encoding,etc.In the main part of this survey,we review recent transformer-based models in three categories:Transformer for downstream tasks,Vision Transformer for Generation,and Vision Transformer for Segmentation.We also provide a comprehensive overview of recent transformer models for video tasks and diffusion models.We compare the performance of various hierarchical transformer networks for multiple tasks on popular benchmark datasets.Finally,we explore some future research directions to further improve the field.
基金supported by the National Natural Science Foundation of China(NSFC)under grant 62125108the National Science and Technology Major Project-Mobile Information Networks under Grant No.2024ZD1300700.
摘要Multiple-input multiple-output(MIMO)systems are essential for improving capacity and reliability in semantic communications.Existing methods mainly design the channel-aware neural networks but neglect the underlying signal distribution.In this paper,we develop a denoising diffusion null-space model-based module over MIMO channels(DDNM-MIMO),which is a plug-in module deployed at the receiver.By modeling the MIMO channel,precoding,and equalization as a linear transformation with additive noise,we design corresponding linear and scaling matrices to construct a sampling process for denoising the received signal.The DDNM-MIMO integrates channel state information(CSI)embedding,supporting both closed-loop MIMO with CSI at the transmitter and open-loop MIMO with CSI at the receiver,thereby improving channel adaptability across various noise levels.As a plug-in,the DDNM-MIMO module operates independently of the joint source-channel coding(JSCC)coder structure,offering flexible integration into diverse systems.Experimental results show that DDNM-MIMO effectively reduces the mean square errors(MSE)between the encoded and equalized signals.Consequently,the proposed DDNM-MIMO semantic communication system achieves superior image reconstruction performance compared to existing JSCC-based semantic communication method.
基金supported by the National Science and Technology Major Project(Grants No.2023ZD0120702)Basic Research Program of Jiangsu(BK20231215)+1 种基金National Natural Science Foundation of China(Grant No.82401075)Natural Science Foundation of Jiangsu Province Major Project(BK20232012).
摘要Nuclear magnetic resonance(NMR)spectroscopy is a key method for molecular structure elucidation.However,interpreting NMR spectra to deduce molecular structures remains challenging due to the complexity of spectral data and the vastness of the chemical space.Here we introduce DiffNMR,a novel end-to-end framework that leverages a conditional discrete diffusion model for de novo molecular structure elucidation from NMR spectra.DiffNMR refines molecular graphs iteratively through a diffusion-based generative process,ensuring global consistency and mitigating error accumulation inherent in autoregressive methods.The framework integrates a two-stage pretraining strategy that aligns spectral and molecular representations via a diffusion autoencoder and contrastive learning.It also incorporates retrieval initialization and similarity filtering during inference.Our experimental results demonstrate that DiffNMR achieves competitive performance for NMR-based structure elucidation,especially outperforming autoregressive models in domain generalization and robustness,thereby offering an efficient and robust solution for automated molecular analysis.
基金supported in part by Chinese National Natural Science Foundation Projects,China(Nos.U23B2054,62276254 and 62372314)Beijing Natural Science Foundation,China(No.L221013)+1 种基金InnoHK program,and Hong Kong Research Grants Council through Research Impact Fund,China(No.R1015-23)Open access funding provided by The Hong Kong Polytechnic University,China.
摘要Recent advancements in diffusion models have significantly impacted content creation,leading to the emergence of per-sonalized content synthesis(PCS).By utilizing a small set of user-provided examples featuring the same subject,PCS aims to tailor this subject to specific user-defined prompts.Over the past two years,more than 150 methods have been introduced in this area.However,existing surveys primarily focus on text-to-image generation,with few providing up-to-date summaries on PCS.This pa-per provides a comprehensive survey of PCS,introducing the general frameworks of PCS research,which can be categorized into test-time fine-tuning(TTF)and pre-trained adaptation(PTA)approaches.We analyze the strengths,limitations and key tech-niques of these methodologies.Additionally,we explore specialized tasks within the field,such as object,face and style personaliza-tion,while highlighting their unique challenges and innovations.Despite the promising progress,we also discuss ongoing challenges,including overfitting and the trade-off between subject fidelity and text alignment.Through this detailed overview and analysis,we propose future directions to further the development of PCS.
摘要Denoising diffusion models have demonstrated tremendous success in modeling data distributions and synthesizing high-quality samples.In the 2D image domain,they have become the state-of-the-art and are capable of generating photo-realistic images with high controllability.More recently,researchers have begun to explore how to utilize diffusion models to generate 3D data,as doing so has more potential in real-world applications.This requires careful design choices in two key ways:identifying a suitable 3D representation and determining how to apply the diffusion process.In this survey,we provide the first comprehensive review of diffusion models for manipulating 3D content,including 3D generation,reconstruction,and 3D-aware image synthesis.We classify existing methods into three major categories:2D space diffusion with pretrained models,2D space diffusion without pretrained models,and 3D space diffusion.We also summarize popular datasets used for 3D generation with diffusion models.Along with this survey,we maintain a repository http://gffzz188fe103f8f1460askfcbcfpb6fv9696k.ffgz.tsg.suse.edu.cn/cwchenwang/awesome-3d-diffusion to track the latest relevant papers and codebases.Finally,we pose current challenges for diffusion models for 3D generation,and suggest future research directions.
基金supported by the National Natural Science Foundation of China(Nos.62366050,61966033,and 61866035)。
摘要Diffusion models are a type of generative deep learning model that can process medical images more efficiently than traditional generative models.They have been applied to several medical image computing tasks.This paper aims to help researchers understand the advancements of diffusion models in medical image computing.It begins by describing the fundamental principles,sampling methods,and architecture of diffusion models.Subsequently,it discusses the application of diffusion models in five medical image computing tasks:image generation,modality conversion,image segmentation,image denoising,and anomaly detection.Additionally,this paper conducts fine-tuning of a large model for image generation tasks and comparative experiments between diffusion models and traditional generative models across these five tasks.The evaluation of the fine-tuned large model shows its potential for clinical applications.Comparative experiments demonstrate that diffusion models have a distinct advantage in tasks related to image generation,modality conversion,and image denoising.However,they require further optimization in image segmentation and anomaly detection tasks to match the efficacy of traditional models.Our codes are publicly available at:http://gffzz188fe103f8f1460askfcbcfpb6fv9696k.ffgz.tsg.suse.edu.cn/hiahub/CodeForDiffusion.
基金supported by the Ministry of Higher Education(MOHE)under the 2023 Translational Research Program for the Energy Sustainability Focus Area(Project ID:MMUE/240001)the 2024 ASEAN IVO(Project ID:2024-02),Multimedia University,and Deanship of Research,Islamic University of Madinah.
摘要The detection of zero-day malware represents one of the most significant challenges in contemporary cybersecurity.In this paper,we introduce a novel concept called“Negative-One-Day Malware Detection”,which aims to identify potentially malicious software before it is actually created by threat actors.Our approach leverages recent advancements in generative AI,specifically diffusion-based generative models,to generate and analyze potential future malware variants.By doing so,we can train detection systems to recognize these variants before they emerge in the wild,thereby closing the critical protection gap that currently exists between malware creation and detection.We demonstrate the effectiveness of our approach through extensive experimentation,showing that our framework can generate executable malware samples that combine characteristics from different families while exhibiting novel behaviors.These synthetically generated samples significantly improve the detection capabilities of security systems when incorporated into training data,providing a proactive rather than reactive approach to cybersecurity.
基金the National Natural Science Foundation of China(Grant No.:52508343)the Fundamental Research Funds for the Central Universities(Grant No.:B250201004).
摘要Crack detection accuracy in computer vision is often constrained by limited annotated datasets.Although Generative Adversarial Networks(GANs)have been applied for data augmentation,they frequently introduce blurs and artifacts.To address this challenge,this study leverages Denoising Diffusion Probabilistic Models(DDPMs)to generate high-quality synthetic crack images,enriching the training set with diverse and structurally consistent samples that enhance the crack segmentation.The proposed framework involves a two-stage pipeline:first,DDPMs are used to synthesize high-fidelity crack images that capture fine structural details.Second,these generated samples are combined with real data to train segmentation networks,thereby improving accuracy and robustness in crack detection.Compared with GAN-based approaches,DDPM achieved the best fidelity,with the highest Structural Similarity Index(SSIM)(0.302)and lowest Learned Perceptual Image Patch Similarity(LPIPS)(0.461),producing artifact-free images that preserve fine crack details.To validate its effectiveness,six segmentation models were tested,among which LinkNet consistently achieved the best performance,excelling in both region-level accuracy and structural continuity.Incorporating DDPM-augmented data further enhanced segmentation outcomes,increasing F1 scores by up to 1.1%and IoU by 1.7%,while also improving boundary alignment and skeleton continuity compared with models trained on real images alone.Experiments with varying augmentation ratios showed consistent improvements,with F1 rising from 0.946(no augmentation)to 0.957 and IoU from 0.897 to 0.913 at the highest ratio.These findings demonstrate the effectiveness of diffusion-based augmentation for complex crack detection in structural health monitoring.
基金supported by the National Natural Science Foundation of China(Nos.61906168,62202429 and 62272267)the Zhejiang Provincial Natural Science Foundation of China(No.LY23F020023)the Construction of Hubei Provincial Key Laboratory for Intelligent Visual Monitoring of Hydropower Projects(No.2022SDSJ01)。
摘要Accurately identifying building distribution from remote sensing images with complex background information is challenging.The emergence of diffusion models has prompted the innovative idea of employing the reverse denoising process to distill building distribution from these complex backgrounds.Building on this concept,we propose a novel framework,building extraction diffusion model(BEDiff),which meticulously refines the extraction of building footprints from remote sensing images in a stepwise fashion.Our approach begins with the design of booster guidance,a mechanism that extracts structural and semantic features from remote sensing images to serve as priors,thereby providing targeted guidance for the diffusion process.Additionally,we introduce a cross-feature fusion module(CFM)that bridges the semantic gap between different types of features,facilitating the integration of the attributes extracted by booster guidance into the diffusion process more effectively.Our proposed BEDiff marks the first application of diffusion models to the task of building extraction.Empirical evidence from extensive experiments on the Beijing building dataset demonstrates the superior performance of BEDiff,affirming its effectiveness and potential for enhancing the accuracy of building extraction in complex urban landscapes.
基金supported by the Key Project of International Cooperation of Qilu University of Technology(Grant No.:QLUTGJHZ2018008)Shandong Provincial Natural Science Foundation Committee,China(Grant No.:ZR2016HB54)Shandong Provincial Key Laboratory of Microbial Engineering(SME).
摘要AlphaPanda(AlphaFold2[1]inspired protein-specific antibody design in a diffusional manner)is an advanced algorithm for designing complementary determining regions(CDRs)of the antibody targeted the specific epitope,combining transformer[2]models,3DCNN[3],and diffusion[4]generative models.
基金supported by the National Natural Science Foundation of China(Grant No.62202210).
摘要The application of generative artificial intelligence(AI)is bringing about notable changes in anime creation.This paper surveys recent advancements and applications of diffusion and language models in anime generation,focusing on their demonstrated potential to enhance production efficiency through automation and personalization.Despite these benefits,it is crucial to acknowledge the substantial initial computational investments required for training and deploying these models.We conduct an in-depth survey of cutting-edge generative AI technologies,encompassing models such as Stable Diffusion and GPT,and appraise pivotal large-scale datasets alongside quantifiable evaluation metrics.Review of the surveyed literature indicates the achievement of considerable maturity in the capacity of AI models to synthesize high-quality,aesthetically compelling anime visual images from textual prompts,alongside discernible progress in the generation of coherent narratives.However,achieving perfect long-form consistency,mitigating artifacts like flickering in video sequences,and enabling fine-grained artistic control remain critical ongoing challenges.Building upon these advancements,research efforts have increasingly pivoted towards the synthesis of higher-dimensional content,such as video and three-dimensional assets,with recent studies demonstrating significant progress in this burgeoning field.Nevertheless,formidable challenges endure amidst these advancements.Foremost among these are the substantial computational exigencies requisite for training and deploying these sophisticated models,particularly pronounced in the realm of high-dimensional generation such as video synthesis.Additional persistent hurdles include maintaining spatial-temporal consistency across complex scenes and mitigating ethical considerations surrounding bias and the preservation of human creative autonomy.This research underscores the transformative potential and inherent complexities of AI-driven synergy within the creative industries.We posit that future research should be dedicated to the synergistic fusion of diffusion and autoregressive models,the integration of multimodal inputs,and the balanced consideration of ethical implications,particularly regarding bias and the preservation of human creative autonomy,thereby establishing a robust foundation for the advancement of anime creation and the broader landscape of AI-driven content generation.
摘要The application of machine learning in fluid dynamics has become increasingly prevalent for accelerating computations in solving forward and inverse problems governed by partial differential equations,particularly flow field reconstruction.However,existing end-to-end methods often depend heavily on specific low-fidelity patterns or sparsity rates during training,limiting their effectiveness in real-world scenarios where inputs may deviate from the training distribution or contain unanticipated noise.Diffusion models present a promising alternative by learning to transform various low-fidelity distributions into high-fidelity ones,offering greater flexibility than direct mapping approaches that are typically constrained to single problem types.In this work,we introduce the physics-informed residual diffusion model,which generalizes across diverse inputs,including evenly down-sampled sensor data,sparse sensor data with Gaussian noise,and randomly sampled sparse sensor data.By incorporating partial differential equation constraints into the objective function,our approach significantly improves the accuracy of reconstructed high-fidelity flow fields while ensuring adherence to underlying physical laws.Experimental results demonstrate that our model successfully generates high-quality outcomes for two-dimensional Navier-Stokes equation flow and Kolmogorov flow under varied low-fidelity input conditions,without the need for retraining.