Discriminative region localization and efficient feature encoding are crucial for fine-grained object recognition.However,existing data augmentation methods struggle to accurately locate discriminative regions in comp...Discriminative region localization and efficient feature encoding are crucial for fine-grained object recognition.However,existing data augmentation methods struggle to accurately locate discriminative regions in complex backgrounds,small target objects,and limited training data,leading to poor recognition.Fine-grained images exhibit“small inter-class differences,”and while second-order feature encoding enhances discrimination,it often requires dual Convolutional Neural Networks(CNN),increasing training time and complexity.This study proposes a model integrating discriminative region localization and efficient second-order feature encoding.By ranking feature map channels via a fully connected layer,it selects high-importance channels to generate an enhanced map,accurately locating discriminative regions.Cropping and erasing augmentations further refine recognition.To improve efficiency,a novel second-order feature encoding module generates an attention map from the fourth convolutional group of Residual Network 50 layers(ResNet-50)and multiplies it with features from the fifth group,producing second-order features while reducing dimensionality and training time.Experiments on Caltech-University of California,San Diego Birds-200-2011(CUB-200-2011),Stanford Car,and Fine-Grained Visual Classification of Aircraft(FGVC Aircraft)datasets show state-of-the-art accuracy of 88.9%,94.7%,and 93.3%,respectively.展开更多
Remote sensing object detection aims to identify and localize specific targets in satellite or aerial imagery.Spiking Neural Networks(SNNs),benefiting from their implicit feedback-based and event-driven brain-inspired...Remote sensing object detection aims to identify and localize specific targets in satellite or aerial imagery.Spiking Neural Networks(SNNs),benefiting from their implicit feedback-based and event-driven brain-inspired dynamics,offer a promising solution to alleviate the high energy consumption of conventional ANN-based detection models.However,existing SNN-based approaches for remote sensing object detection—particularly for small,arbitrarily rotated objects—are still in their infancy and suffer from a substantial performance gap compared with ANN counterparts.In this work,we draw inspiration from the hierarchical sparse perception mechanisms of biological vision and integrate dynamic receptive field modulation into the encoding stage,proposing a high-precision spiking object detection framework tailored for remote sensing image.Specifically,we design a Hierarchical Feedback-based Gaussian Encoding(HFG)scheme,in which the parameters of Gaussian kernels are dynamically adjusted through spike-triggered top-down feedback connections.This mechanism enables the encoding process to adaptively respond to complex geometric variations of remote sensing objects,including rotation and scale changes.Based on the proposed encoding strategy,we develop DGRDet(Dynamic Gaussian Receptive Field Encoding-based Spiking Neural Networks for Remote Sensing Object Detection),a directly trained deep SNN detector for remote sensing image.Extensive evaluations on the large-scale public DOTA dataset demonstrate that DGRDet achieves competitive detection accuracy,outperforming existing SNN-based object detection methods.Moreover,compared with ANN models of comparable detection performance,DGRDet reduces spike activity by 81.31%and requires only 0.12%of the inference energy consumption,achieving a favorable balance between detection accuracy,efficiency,and energy efficiency.展开更多
Transformers have become the dominant architecture for sequence modeling in natural language processing;however,their effectiveness critically depends on how positional information is encoded.Conventional positional e...Transformers have become the dominant architecture for sequence modeling in natural language processing;however,their effectiveness critically depends on how positional information is encoded.Conventional positional encodings,while effective,may have limited structural flexibility for capturing complex global sequence relationships.Recent quantum-inspired approaches have sought to address this limitation,yetmany either oversimplify quantum principles or introduce substantial computational or hardware overhead.We introduce a novel Quantum Fourier Transform(QFT)-inspired positional encoding scheme for transformers,motivated by the structured frequency representation of the QFT.Unlike prior approaches that either emulate quantum operations superficially or require complex circuit constructions,the proposed method provides a learnable hybrid encoding that preserves quantuminspired structure while remaining aligned with hardware-efficient circuit primitives and structurally compatible with future near-term quantum implementations.Experiments on WikiText-103 indicate that the proposed encoding achieves competitive perplexity,improved robustness to input scrambling,and stable training behavior relative to alternative quantum-inspired baselines under the evaluated settings.Preliminary circuit-level simulations further suggest favorable noise resilience of the associated encoding primitives.These findings support the potential utility of incorporating quantum-inspired design principles into deep learning architectures and provide a foundation for future exploration at the interface of quantum computing and transformer-based natural language processing(NLP).展开更多
Understanding how birds perceive and recognize visual objects remains a fundamental question in neuroscience.The entopallium,a key node in the avian tectofugal pathway,has long been implicated in complex visual proces...Understanding how birds perceive and recognize visual objects remains a fundamental question in neuroscience.The entopallium,a key node in the avian tectofugal pathway,has long been implicated in complex visual processing,yet its internal functional architecture remains incompletely understood.In this study,neuronal activity in the pigeon entopallium was systematically mapped using controlled visual stimuli that independently varied in color,shape,and motion.Recordings revealed marked hue selectivity that remained invariant across luminance levels,pronounced orientation tuning in response to shape stimuli,and robust direction selectivity for moving stimuli.Spatial mapping further revealed distinct functional segregation,with color-selective neurons localized anteroventrally,shape-selective neurons dorsally,and motion-selective neurons posteriorly.At the same time,partial overlap among these response classes was observed,with a subset of neurons exhibiting joint tuning across stimulus dimensions,suggesting an organizational scheme characterized by regional specialization and partial cross-feature integration.Notably,entopallium neurons exhibited a moderate level of visual feature integration and shared important functional properties with early to intermediate stages of mammalian visual processing.Together,these findings establish the entopallium as a major site for multidimensional visual analysis in birds and provide evidence for convergent principles underlying the evolution of complex visual systems across vertebrates.展开更多
This study proposes a deep learning-based method termed frequency-flexible chemical exchange saturation transfer(CEST)imaging network(FlexCENT),which enables robust CEST quantification across variable frequency offset...This study proposes a deep learning-based method termed frequency-flexible chemical exchange saturation transfer(CEST)imaging network(FlexCENT),which enables robust CEST quantification across variable frequency offset schemes without requiring retraining.FlexCENT integrates frequency offset encoding with a three-dimensional(3D)U-Net to process CEST images and frequency offsets as inputs and predict Lorentzian parameters of the 4-pool model(water,MT,APT,rNOE),including B0 inhomogeneity.By transforming frequency offsets into a continuous spectral feature representation,the frequency offset encoding allows FlexCENT to generalize to unseen frequency offset schemes.Trained on synthetic data generated from the 4-pool Lorentzian model,FlexCENT was validated through numerical simulations,tumor-bearing mouse experiments,and a human brain experiment,alongside comparisons with 4-pool Lorentzian fitting,DeepCEST,and LKAN networks.The results demonstrate that FlexCENT successfully quantified CEST parameters across all experiments,maintaining consistent performance under varying frequency offset conditions without retraining.It exhibited superior noise robustness in numerical simulations and enhanced anatomical delineation in vivo parametric mapping compared to other methods.In conclusion,by combining spectral information with spatial information,FlexCENT provides an efficient,flexible,and robust quantitative approach for CEST imaging.It significantly enhance the quantification capability and clinical potential of CEST imaging.展开更多
Numerous neuropsychiatric disorders are characterized by significant impairments in decision-making function.These include impulsive decision-making in attention-deficit hyperactivity disorder(ADHD)[1],excessive risk-...Numerous neuropsychiatric disorders are characterized by significant impairments in decision-making function.These include impulsive decision-making in attention-deficit hyperactivity disorder(ADHD)[1],excessive risk-taking during manic episodes in bipolar disorder,and the distorted prioritization observed in substance use disorders.Decisionmaking involves reflecting on the outcomes of past actions and weighing the potential consequences of future actions.In this complex balancing process,mesolimbic dopamine influences reward value assessment,the strength of motivation,and the initiation of action[2].展开更多
We demonstrate a thermally tunable mode-locked fiber laser integrating a polarization-sensitive SMF-PMFSMF modulator and a Bi2TeSe2 saturable absorber(SMF:single-mode fiber;PMF:polarization-maintaining fiber).By...We demonstrate a thermally tunable mode-locked fiber laser integrating a polarization-sensitive SMF-PMFSMF modulator and a Bi2TeSe2 saturable absorber(SMF:single-mode fiber;PMF:polarization-maintaining fiber).By controlling the PMF temperature,reversible switching among conventional,dissipative,and boundstate solitons is achieved.The wavelength tuning ranges are about 5 nm and 2.8 nm for conventional and dissipative solitons,respectively,with a tuning efficiency of 0.35 nm/℃.Numerical simulations based on temperatureinduced birefringence variation reproduce the observed dynamics.Furthermore,a wavelength-encoding scheme utilizing thermally driven soliton shifts is proposed,providing a feasible approach for soliton-state-controlled optical communication.展开更多
Non-line-of-sight(NLOS)imaging aims to reconstruct objects beyond line-of-sight view,offering potential applications in various fields.However,conventional transient NLOS methods are constrained by the detection regio...Non-line-of-sight(NLOS)imaging aims to reconstruct objects beyond line-of-sight view,offering potential applications in various fields.However,conventional transient NLOS methods are constrained by the detection region,preventing the reconstruction of targets outside its normal space and thereby limiting practical applicability.In this paper,a computational imaging method for super-field-of-view(Super-FoV)reconstruction based on spatial encoding of a translated point spread function(PSF)is proposed.展开更多
Temporal ghost imaging(TGI)enables ultrafast signal reconstruction beyond electronic bandwidth limits.Extending this concept to the mid-infrared(MIR)regime through nonlinear frequency conversion offers new opportuniti...Temporal ghost imaging(TGI)enables ultrafast signal reconstruction beyond electronic bandwidth limits.Extending this concept to the mid-infrared(MIR)regime through nonlinear frequency conversion offers new opportunities for high-fidelity temporal detection,but it remains constrained by the stringent phase-matching condition,limited spectral coverage,and intricate optical alignment.Here,we propose and demonstrate a broad-band MIR TGI system based on non-degenerate two-photon absorption.展开更多
Highly programmable shape morphing of 4D-printed microanostructures is urgently desired for applications in robotics and intelligent systems.However,due to the lack of autonomous holistic strategies throughout the tar...Highly programmable shape morphing of 4D-printed microanostructures is urgently desired for applications in robotics and intelligent systems.However,due to the lack of autonomous holistic strategies throughout the target shape input,optimal material distribution generation,and fabrication program output,4D nanoprinting that permits arbitrary shape morphing remains a challenging task for manual design.In this study,we report an autonomous inverse encoding strategy to decipher the genetic code for material property distributions that can guide the encoded modeling toward arbitrarily pre-programmed 4D shape morphing.By tuning the laser power of each voxel at the nanoscale,the genetic code can be spatially programmed and controllable shape morphing can be realized through the inverse encoding process.Using this strategy,the 4D-printed structures can be designed and accurately shift to the target morphing of arbitrarily hand-drawn lines under stimulation.Furthermore,as a proof-of-concept,a flexible fiber micromanipulator that can approach the target region through pre-programmed shape morphing is autonomously inversely encoded according to the localized spatial environment.This strategy may contribute to the modeling and arbitrary shape morphing of microanostructures fabricated via 4D nanoprinting,leading to cutting-edge applications in microfluidics,micro-robotics,minimally invasive robotic surgery,and tissue engineering.展开更多
Deep learning(DL)methods like multilayer perceptrons(MLPs)and convolutional neural networks(CNNs)have been applied to predict the complex traits in animal and plant breeding.However,improving the genomic prediction ac...Deep learning(DL)methods like multilayer perceptrons(MLPs)and convolutional neural networks(CNNs)have been applied to predict the complex traits in animal and plant breeding.However,improving the genomic prediction accuracy still presents signifcant challenges.In this study,we applied CNNs to predict swine traits using previously published data.Specifcally,we extensively evaluated the CNN model's performance by employing various sets of single nucleotide polymorphisms(SNPs)and concluded that the CNN model achieved optimal performance when utilizing SNP sets comprising 1,000 SNPs.Furthermore,we adopted a novel approach using the one-hot encoding method that transforms the 16 different genotypes into sets of eight binary variables.This innovative encoding method signifcantly enhanced the CNN's prediction accuracy for swine traits,outperforming the traditional one-hot encoding techniques.Our fndings suggest that the expanded one-hot encoding method can improve the accuracy of DL methods in the genomic prediction of swine agricultural economic traits.This discovery has significant implications for swine breeding programs,where genomic prediction is pivotal in improving breeding strategies.Furthermore,future research endeavors can explore additional enhancements to DL methods by incorporating advanced data pre-processing techniques.展开更多
Multimodal sentiment analysis aims to understand emotions from text,speech,and video data.However,current methods often overlook the dominant role of text and suffer from feature loss during integration.Given the vary...Multimodal sentiment analysis aims to understand emotions from text,speech,and video data.However,current methods often overlook the dominant role of text and suffer from feature loss during integration.Given the varying importance of each modality across different contexts,a central and pressing challenge in multimodal sentiment analysis lies in maximizing the use of rich intra-modal features while minimizing information loss during the fusion process.In response to these critical limitations,we propose a novel framework that integrates spatial position encoding and fusion embedding modules to address these issues.In our model,text is treated as the core modality,while speech and video features are selectively incorporated through a unique position-aware fusion process.The spatial position encoding strategy preserves the internal structural information of speech and visual modalities,enabling the model to capture localized intra-modal dependencies that are often overlooked.This design enhances the richness and discriminative power of the fused representation,enabling more accurate and context-aware sentiment prediction.Finally,we conduct comprehensive evaluations on two widely recognized standard datasets in the field—CMU-MOSI and CMU-MOSEI to validate the performance of the proposed model.The experimental results demonstrate that our model exhibits good performance and effectiveness for sentiment analysis tasks.展开更多
Quantum communication networks,such as quantum key distribution(QKD)networks,typically employ the measurement-resend mechanism between two users using quantum communication devices based on different quantum encoding ...Quantum communication networks,such as quantum key distribution(QKD)networks,typically employ the measurement-resend mechanism between two users using quantum communication devices based on different quantum encoding types.To achieve direct communication between the devices with different quantum encoding types,in this paper,we propose encoding conversion schemes between the polarization bases(rectilinear,diagonal and circular bases)and the time-bin phase bases(two phase bases and time-bin basis)and design the quantum encoding converters.The theoretical analysis of the encoding conversion schemes is given in detail,and the basis correspondence of encoding conversion and the property of bit flip are revealed.The conversion relationship between polarization bases and time-bin phase bases can be easily selected by controlling a phase shifter.Since no optical switches are used in our scheme,the converter can be operated with high speed.The converters can also be modularized,which may be utilized to realize miniaturization in the future.展开更多
With the rapid expansion of social media,analyzing emotions and their causes in texts has gained significant importance.Emotion-cause pair extraction enables the identification of causal relationships between emotions...With the rapid expansion of social media,analyzing emotions and their causes in texts has gained significant importance.Emotion-cause pair extraction enables the identification of causal relationships between emotions and their triggers within a text,facilitating a deeper understanding of expressed sentiments and their underlying reasons.This comprehension is crucial for making informed strategic decisions in various business and societal contexts.However,recent research approaches employing multi-task learning frameworks for modeling often face challenges such as the inability to simultaneouslymodel extracted features and their interactions,or inconsistencies in label prediction between emotion-cause pair extraction and independent assistant tasks like emotion and cause extraction.To address these issues,this study proposes an emotion-cause pair extraction methodology that incorporates joint feature encoding and task alignment mechanisms.The model consists of two primary components:First,joint feature encoding simultaneously generates features for emotion-cause pairs and clauses,enhancing feature interactions between emotion clauses,cause clauses,and emotion-cause pairs.Second,the task alignment technique is applied to reduce the labeling distance between emotion-cause pair extraction and the two assistant tasks,capturing deep semantic information interactions among tasks.The proposed method is evaluated on a Chinese benchmark corpus using 10-fold cross-validation,assessing key performance metrics such as precision,recall,and F1 score.Experimental results demonstrate that the model achieves an F1 score of 76.05%,surpassing the state-of-the-art by 1.03%.The proposed model exhibits significant improvements in emotion-cause pair extraction(ECPE)and cause extraction(CE)compared to existing methods,validating its effectiveness.This research introduces a novel approach based on joint feature encoding and task alignment mechanisms,contributing to advancements in emotion-cause pair extraction.However,the study’s limitation lies in the data sources,potentially restricting the generalizability of the findings.展开更多
The Gaussian phase distribution approximation enables analysis of restricted diffusion encoded by general gradient waveforms but fails to account for the diffraction-like features that may occur for simple pore geomet...The Gaussian phase distribution approximation enables analysis of restricted diffusion encoded by general gradient waveforms but fails to account for the diffraction-like features that may occur for simple pore geometries.We investigate the range of validity of the approximation by random walk simulations of restricted diffusion in a cylinder using isotropic diffusion encoding sequences as well as conventional single gradient pulse pairs and oscillating gradient waveforms.The results show that clear deviations from the approximation may be observed at relative signal attenuations below 0.1 for onedimensional sequences with few oscillation periods.Increasing the encoding dimensionality and/or number of oscillations while extending the total duration of the waveform diminishes the non-Gaussian effects while preserving the low apparent diffusivities characteristic of restriction.展开更多
Sensitivity encoding(SENSE)is a parallel magnetic resonance imaging(MRI)reconstruction model by utilizing the sensitivity information of receiver coils to achieve image reconstruction.The existing SENSE-based reconstr...Sensitivity encoding(SENSE)is a parallel magnetic resonance imaging(MRI)reconstruction model by utilizing the sensitivity information of receiver coils to achieve image reconstruction.The existing SENSE-based reconstruction algorithms usually used nonadaptive sparsifying transforms,resulting in a limited reconstruction accuracy.Therefore,we proposed a new model for accurate parallel MRI reconstruction by combining the L0 norm regularization term based on the efficient sum of outer products dictionary learning(SOUPDIL)with the SENSE model,called SOUPDIL-SENSE.The SOUPDIL-SENSE model is mainly solved by utilizing the variable splitting and alternating direction method of multipliers techniques.The experimental results on four human datasets show that the proposed algorithm effectively promotes the image sparsity,eliminates the noise and artifacts of the reconstructed images,and improves the reconstruction accuracy.展开更多
Blockchain,as a distributed ledger,inherently possesses tamper-resistant capabilities,creating a natural channel for covert communication.However,the immutable nature of data storage might introduce challenges to comm...Blockchain,as a distributed ledger,inherently possesses tamper-resistant capabilities,creating a natural channel for covert communication.However,the immutable nature of data storage might introduce challenges to communication security.This study introduces a blockchain-based covert communication model utilizing dynamic Base-K encoding.The proposed encoding scheme utilizes the input address sequence to determine K to encode the secret message and determines the order of transactions based on K,thus ensuring effective concealment of the message.The dynamic encoding parameters enhance flexibility and address issues related to identical transaction amounts for the same secret message.Experimental results demonstrate that the proposed method maintains smooth communication and low susceptibility to tampering,achieving commendable concealment and embedding rates.展开更多
High-Speed Trains (HSTs) have emerged as a mainstream mode of transportation in China, owing to their exceptional safety and efficiency. Ensuring the reliable operation of HSTs is of paramount economic and societal im...High-Speed Trains (HSTs) have emerged as a mainstream mode of transportation in China, owing to their exceptional safety and efficiency. Ensuring the reliable operation of HSTs is of paramount economic and societal importance. As critical rotating mechanical components of the transmission system, bearings make their fault diagnosis a topic of extensive attention. This paper provides a systematic review of image encoding-based bearing fault diagnosis methods tailored to the condition monitoring of HSTs. First, it categorizes the image encoding techniques applied in the field of bearing fault diagnosis. Then, a review of state-of-the-art studies has been presented, encompassing both monomodal image conversion and multimodal image fusion approaches. Finally, it highlights current challenges and proposes future research directions to advance intelligent fault diagnosis in HSTs, aiming to provide a valuable reference for researchers and engineers in the field of intelligent operation and maintenance.展开更多
Retinal blood vessel segmentation is crucial for diagnosing ocular and cardiovascular diseases.Although the introduction of U-Net in 2015 by Olaf Ronneberger significantly advanced this field,yet issues like limited t...Retinal blood vessel segmentation is crucial for diagnosing ocular and cardiovascular diseases.Although the introduction of U-Net in 2015 by Olaf Ronneberger significantly advanced this field,yet issues like limited training data,imbalance data distribution,and inadequate feature extraction persist,hindering both the segmentation performance and optimal model generalization.Addressing these critical issues,the DEFFA-Unet is proposed featuring an additional encoder to process domain-invariant pre-processed inputs,thereby improving both richer feature encoding and enhanced model generalization.A feature filtering fusion module is developed to ensure the precise feature filtering and robust hybrid feature fusion.In response to the task-specific need for higher precision where false positives are very costly,traditional skip connections are replaced with the attention-guided feature reconstructing fusion module.Additionally,innovative data augmentation and balancing methods are proposed to counter data scarcity and distribution imbalance,further boosting the robustness and generalization of the model.With a comprehensive suite of evaluation metrics,extensive validations on four benchmark datasets(DRIVE,CHASEDB1,STARE,and HRF)and an SLO dataset(IOSTAR),demonstrate the proposed method’s superiority over both baseline and state-of-the-art models.Particularly the proposed method significantly outperforms the compared methods in cross-validation model generalization.展开更多
基金supported,in part,by the National Nature Science Foundation of China under Grant 62272236,62376128 and 62306139the Natural Science Foundation of Jiangsu Province under Grant BK20201136,BK20191401.
摘要Discriminative region localization and efficient feature encoding are crucial for fine-grained object recognition.However,existing data augmentation methods struggle to accurately locate discriminative regions in complex backgrounds,small target objects,and limited training data,leading to poor recognition.Fine-grained images exhibit“small inter-class differences,”and while second-order feature encoding enhances discrimination,it often requires dual Convolutional Neural Networks(CNN),increasing training time and complexity.This study proposes a model integrating discriminative region localization and efficient second-order feature encoding.By ranking feature map channels via a fully connected layer,it selects high-importance channels to generate an enhanced map,accurately locating discriminative regions.Cropping and erasing augmentations further refine recognition.To improve efficiency,a novel second-order feature encoding module generates an attention map from the fourth convolutional group of Residual Network 50 layers(ResNet-50)and multiplies it with features from the fifth group,producing second-order features while reducing dimensionality and training time.Experiments on Caltech-University of California,San Diego Birds-200-2011(CUB-200-2011),Stanford Car,and Fine-Grained Visual Classification of Aircraft(FGVC Aircraft)datasets show state-of-the-art accuracy of 88.9%,94.7%,and 93.3%,respectively.
基金funded by the National Key R&D Program of China Grant No.2022YFB4500900.
摘要Remote sensing object detection aims to identify and localize specific targets in satellite or aerial imagery.Spiking Neural Networks(SNNs),benefiting from their implicit feedback-based and event-driven brain-inspired dynamics,offer a promising solution to alleviate the high energy consumption of conventional ANN-based detection models.However,existing SNN-based approaches for remote sensing object detection—particularly for small,arbitrarily rotated objects—are still in their infancy and suffer from a substantial performance gap compared with ANN counterparts.In this work,we draw inspiration from the hierarchical sparse perception mechanisms of biological vision and integrate dynamic receptive field modulation into the encoding stage,proposing a high-precision spiking object detection framework tailored for remote sensing image.Specifically,we design a Hierarchical Feedback-based Gaussian Encoding(HFG)scheme,in which the parameters of Gaussian kernels are dynamically adjusted through spike-triggered top-down feedback connections.This mechanism enables the encoding process to adaptively respond to complex geometric variations of remote sensing objects,including rotation and scale changes.Based on the proposed encoding strategy,we develop DGRDet(Dynamic Gaussian Receptive Field Encoding-based Spiking Neural Networks for Remote Sensing Object Detection),a directly trained deep SNN detector for remote sensing image.Extensive evaluations on the large-scale public DOTA dataset demonstrate that DGRDet achieves competitive detection accuracy,outperforming existing SNN-based object detection methods.Moreover,compared with ANN models of comparable detection performance,DGRDet reduces spike activity by 81.31%and requires only 0.12%of the inference energy consumption,achieving a favorable balance between detection accuracy,efficiency,and energy efficiency.
基金Prince Sattam bin Abdulaziz University for funding this research work through the project number(PSAU/2025/01/35090).
摘要Transformers have become the dominant architecture for sequence modeling in natural language processing;however,their effectiveness critically depends on how positional information is encoded.Conventional positional encodings,while effective,may have limited structural flexibility for capturing complex global sequence relationships.Recent quantum-inspired approaches have sought to address this limitation,yetmany either oversimplify quantum principles or introduce substantial computational or hardware overhead.We introduce a novel Quantum Fourier Transform(QFT)-inspired positional encoding scheme for transformers,motivated by the structured frequency representation of the QFT.Unlike prior approaches that either emulate quantum operations superficially or require complex circuit constructions,the proposed method provides a learnable hybrid encoding that preserves quantuminspired structure while remaining aligned with hardware-efficient circuit primitives and structurally compatible with future near-term quantum implementations.Experiments on WikiText-103 indicate that the proposed encoding achieves competitive perplexity,improved robustness to input scrambling,and stable training behavior relative to alternative quantum-inspired baselines under the evaluated settings.Preliminary circuit-level simulations further suggest favorable noise resilience of the associated encoding primitives.These findings support the potential utility of incorporating quantum-inspired design principles into deep learning architectures and provide a foundation for future exploration at the interface of quantum computing and transformer-based natural language processing(NLP).
基金supported by the National Natural Science Foundation of China(62206253)China Postdoctoral Science Foundation(2024M752934)。
摘要Understanding how birds perceive and recognize visual objects remains a fundamental question in neuroscience.The entopallium,a key node in the avian tectofugal pathway,has long been implicated in complex visual processing,yet its internal functional architecture remains incompletely understood.In this study,neuronal activity in the pigeon entopallium was systematically mapped using controlled visual stimuli that independently varied in color,shape,and motion.Recordings revealed marked hue selectivity that remained invariant across luminance levels,pronounced orientation tuning in response to shape stimuli,and robust direction selectivity for moving stimuli.Spatial mapping further revealed distinct functional segregation,with color-selective neurons localized anteroventrally,shape-selective neurons dorsally,and motion-selective neurons posteriorly.At the same time,partial overlap among these response classes was observed,with a subset of neurons exhibiting joint tuning across stimulus dimensions,suggesting an organizational scheme characterized by regional specialization and partial cross-feature integration.Notably,entopallium neurons exhibited a moderate level of visual feature integration and shared important functional properties with early to intermediate stages of mammalian visual processing.Together,these findings establish the entopallium as a major site for multidimensional visual analysis in birds and provide evidence for convergent principles underlying the evolution of complex visual systems across vertebrates.
基金supported by National Key R&D Program of China[grant number 2023YFA1607502]National Natural Science Foundation of China[grant numbers 12375291,82071913]Guangdong Basic and Applied Basic Research Foundation[grant number 2024A1515011262].
摘要This study proposes a deep learning-based method termed frequency-flexible chemical exchange saturation transfer(CEST)imaging network(FlexCENT),which enables robust CEST quantification across variable frequency offset schemes without requiring retraining.FlexCENT integrates frequency offset encoding with a three-dimensional(3D)U-Net to process CEST images and frequency offsets as inputs and predict Lorentzian parameters of the 4-pool model(water,MT,APT,rNOE),including B0 inhomogeneity.By transforming frequency offsets into a continuous spectral feature representation,the frequency offset encoding allows FlexCENT to generalize to unseen frequency offset schemes.Trained on synthetic data generated from the 4-pool Lorentzian model,FlexCENT was validated through numerical simulations,tumor-bearing mouse experiments,and a human brain experiment,alongside comparisons with 4-pool Lorentzian fitting,DeepCEST,and LKAN networks.The results demonstrate that FlexCENT successfully quantified CEST parameters across all experiments,maintaining consistent performance under varying frequency offset conditions without retraining.It exhibited superior noise robustness in numerical simulations and enhanced anatomical delineation in vivo parametric mapping compared to other methods.In conclusion,by combining spectral information with spatial information,FlexCENT provides an efficient,flexible,and robust quantitative approach for CEST imaging.It significantly enhance the quantification capability and clinical potential of CEST imaging.
基金supported by the grants from the National Natural Science Foundation of China(82404599)the China Postdoctoral Science Foundation-funded project(2025T180963).
摘要Numerous neuropsychiatric disorders are characterized by significant impairments in decision-making function.These include impulsive decision-making in attention-deficit hyperactivity disorder(ADHD)[1],excessive risk-taking during manic episodes in bipolar disorder,and the distorted prioritization observed in substance use disorders.Decisionmaking involves reflecting on the outcomes of past actions and weighing the potential consequences of future actions.In this complex balancing process,mesolimbic dopamine influences reward value assessment,the strength of motivation,and the initiation of action[2].
基金supported by the National Natural Science Foundation of China(Grant Nos.12275240,12261131495,and 12475008)the Natural Science Foundation of Zhejiang Province(Grant No.LY24A050002)。
摘要We demonstrate a thermally tunable mode-locked fiber laser integrating a polarization-sensitive SMF-PMFSMF modulator and a Bi2TeSe2 saturable absorber(SMF:single-mode fiber;PMF:polarization-maintaining fiber).By controlling the PMF temperature,reversible switching among conventional,dissipative,and boundstate solitons is achieved.The wavelength tuning ranges are about 5 nm and 2.8 nm for conventional and dissipative solitons,respectively,with a tuning efficiency of 0.35 nm/℃.Numerical simulations based on temperatureinduced birefringence variation reproduce the observed dynamics.Furthermore,a wavelength-encoding scheme utilizing thermally driven soliton shifts is proposed,providing a feasible approach for soliton-state-controlled optical communication.
基金National Natural Science Foundation of China(62427803,62031018,U23A20283)Jiangsu Provincial Key Research and Development Program(BE2022391)Fundamental Research Funds for the Central Universities(30924010812)。
摘要Non-line-of-sight(NLOS)imaging aims to reconstruct objects beyond line-of-sight view,offering potential applications in various fields.However,conventional transient NLOS methods are constrained by the detection region,preventing the reconstruction of targets outside its normal space and thereby limiting practical applicability.In this paper,a computational imaging method for super-field-of-view(Super-FoV)reconstruction based on spatial encoding of a translated point spread function(PSF)is proposed.
基金Shanghai Pilot Program for Basic Research(TQ20220104)National Natural Science Foundation of China(62175064,62235019)+1 种基金Postdoctoral Fellowship Program(GZC20250545)China Postdoctoral Science Foundation(2024M760918,2025T180224)。
摘要Temporal ghost imaging(TGI)enables ultrafast signal reconstruction beyond electronic bandwidth limits.Extending this concept to the mid-infrared(MIR)regime through nonlinear frequency conversion offers new opportunities for high-fidelity temporal detection,but it remains constrained by the stringent phase-matching condition,limited spectral coverage,and intricate optical alignment.Here,we propose and demonstrate a broad-band MIR TGI system based on non-degenerate two-photon absorption.
基金supported by the National Key Research and Development Project(Grant No.2023YFB4705300)the National Natural Science Foundation of China(NSFC)(Grant Nos.62205200 and 62375168)the Natural Science Foundation of Shanghai(Grant No.22ZR1431600)。
摘要Highly programmable shape morphing of 4D-printed microanostructures is urgently desired for applications in robotics and intelligent systems.However,due to the lack of autonomous holistic strategies throughout the target shape input,optimal material distribution generation,and fabrication program output,4D nanoprinting that permits arbitrary shape morphing remains a challenging task for manual design.In this study,we report an autonomous inverse encoding strategy to decipher the genetic code for material property distributions that can guide the encoded modeling toward arbitrarily pre-programmed 4D shape morphing.By tuning the laser power of each voxel at the nanoscale,the genetic code can be spatially programmed and controllable shape morphing can be realized through the inverse encoding process.Using this strategy,the 4D-printed structures can be designed and accurately shift to the target morphing of arbitrarily hand-drawn lines under stimulation.Furthermore,as a proof-of-concept,a flexible fiber micromanipulator that can approach the target region through pre-programmed shape morphing is autonomously inversely encoded according to the localized spatial environment.This strategy may contribute to the modeling and arbitrary shape morphing of microanostructures fabricated via 4D nanoprinting,leading to cutting-edge applications in microfluidics,micro-robotics,minimally invasive robotic surgery,and tissue engineering.
基金supported by the National Natural Science Foundation of China(32102513)the National Key Scientific Research Project(2023YFF1001100)+1 种基金the Shenzhen Innovation and Entrepreneurship PlanMajor Special Project of Science and Technology,China(KJZD20230923115003006)the Innovation Project of Chinese Academy of Agricultural Sciences(CAAS-ZDRW202006)。
摘要Deep learning(DL)methods like multilayer perceptrons(MLPs)and convolutional neural networks(CNNs)have been applied to predict the complex traits in animal and plant breeding.However,improving the genomic prediction accuracy still presents signifcant challenges.In this study,we applied CNNs to predict swine traits using previously published data.Specifcally,we extensively evaluated the CNN model's performance by employing various sets of single nucleotide polymorphisms(SNPs)and concluded that the CNN model achieved optimal performance when utilizing SNP sets comprising 1,000 SNPs.Furthermore,we adopted a novel approach using the one-hot encoding method that transforms the 16 different genotypes into sets of eight binary variables.This innovative encoding method signifcantly enhanced the CNN's prediction accuracy for swine traits,outperforming the traditional one-hot encoding techniques.Our fndings suggest that the expanded one-hot encoding method can improve the accuracy of DL methods in the genomic prediction of swine agricultural economic traits.This discovery has significant implications for swine breeding programs,where genomic prediction is pivotal in improving breeding strategies.Furthermore,future research endeavors can explore additional enhancements to DL methods by incorporating advanced data pre-processing techniques.
基金supported by the Collaborative Tackling Project of the Yangtze River Delta SciTech Innovation Community(Nos.2024CSJGG01503,2024CSJGG01500)Guangxi Key Research and Development Program(No.AB24010317)Jiangxi Provincial Key Laboratory of Electronic Data Control and Forensics(Jiangxi Police College)(No.2025JXJYKFJJ002).
摘要Multimodal sentiment analysis aims to understand emotions from text,speech,and video data.However,current methods often overlook the dominant role of text and suffer from feature loss during integration.Given the varying importance of each modality across different contexts,a central and pressing challenge in multimodal sentiment analysis lies in maximizing the use of rich intra-modal features while minimizing information loss during the fusion process.In response to these critical limitations,we propose a novel framework that integrates spatial position encoding and fusion embedding modules to address these issues.In our model,text is treated as the core modality,while speech and video features are selectively incorporated through a unique position-aware fusion process.The spatial position encoding strategy preserves the internal structural information of speech and visual modalities,enabling the model to capture localized intra-modal dependencies that are often overlooked.This design enhances the richness and discriminative power of the fused representation,enabling more accurate and context-aware sentiment prediction.Finally,we conduct comprehensive evaluations on two widely recognized standard datasets in the field—CMU-MOSI and CMU-MOSEI to validate the performance of the proposed model.The experimental results demonstrate that our model exhibits good performance and effectiveness for sentiment analysis tasks.
基金supported by the National Natural Science Foundation of China(Grant No.62001440).
摘要Quantum communication networks,such as quantum key distribution(QKD)networks,typically employ the measurement-resend mechanism between two users using quantum communication devices based on different quantum encoding types.To achieve direct communication between the devices with different quantum encoding types,in this paper,we propose encoding conversion schemes between the polarization bases(rectilinear,diagonal and circular bases)and the time-bin phase bases(two phase bases and time-bin basis)and design the quantum encoding converters.The theoretical analysis of the encoding conversion schemes is given in detail,and the basis correspondence of encoding conversion and the property of bit flip are revealed.The conversion relationship between polarization bases and time-bin phase bases can be easily selected by controlling a phase shifter.Since no optical switches are used in our scheme,the converter can be operated with high speed.The converters can also be modularized,which may be utilized to realize miniaturization in the future.
摘要With the rapid expansion of social media,analyzing emotions and their causes in texts has gained significant importance.Emotion-cause pair extraction enables the identification of causal relationships between emotions and their triggers within a text,facilitating a deeper understanding of expressed sentiments and their underlying reasons.This comprehension is crucial for making informed strategic decisions in various business and societal contexts.However,recent research approaches employing multi-task learning frameworks for modeling often face challenges such as the inability to simultaneouslymodel extracted features and their interactions,or inconsistencies in label prediction between emotion-cause pair extraction and independent assistant tasks like emotion and cause extraction.To address these issues,this study proposes an emotion-cause pair extraction methodology that incorporates joint feature encoding and task alignment mechanisms.The model consists of two primary components:First,joint feature encoding simultaneously generates features for emotion-cause pairs and clauses,enhancing feature interactions between emotion clauses,cause clauses,and emotion-cause pairs.Second,the task alignment technique is applied to reduce the labeling distance between emotion-cause pair extraction and the two assistant tasks,capturing deep semantic information interactions among tasks.The proposed method is evaluated on a Chinese benchmark corpus using 10-fold cross-validation,assessing key performance metrics such as precision,recall,and F1 score.Experimental results demonstrate that the model achieves an F1 score of 76.05%,surpassing the state-of-the-art by 1.03%.The proposed model exhibits significant improvements in emotion-cause pair extraction(ECPE)and cause extraction(CE)compared to existing methods,validating its effectiveness.This research introduces a novel approach based on joint feature encoding and task alignment mechanisms,contributing to advancements in emotion-cause pair extraction.However,the study’s limitation lies in the data sources,potentially restricting the generalizability of the findings.
基金financially supported by the Swedish Research Council(2022-04422_VR)。
摘要The Gaussian phase distribution approximation enables analysis of restricted diffusion encoded by general gradient waveforms but fails to account for the diffraction-like features that may occur for simple pore geometries.We investigate the range of validity of the approximation by random walk simulations of restricted diffusion in a cylinder using isotropic diffusion encoding sequences as well as conventional single gradient pulse pairs and oscillating gradient waveforms.The results show that clear deviations from the approximation may be observed at relative signal attenuations below 0.1 for onedimensional sequences with few oscillation periods.Increasing the encoding dimensionality and/or number of oscillations while extending the total duration of the waveform diminishes the non-Gaussian effects while preserving the low apparent diffusivities characteristic of restriction.
基金the National Natural Science Foundation of China(No.61861023)the Yunnan Fundamental Research Project(No.202301AT070452)。
摘要Sensitivity encoding(SENSE)is a parallel magnetic resonance imaging(MRI)reconstruction model by utilizing the sensitivity information of receiver coils to achieve image reconstruction.The existing SENSE-based reconstruction algorithms usually used nonadaptive sparsifying transforms,resulting in a limited reconstruction accuracy.Therefore,we proposed a new model for accurate parallel MRI reconstruction by combining the L0 norm regularization term based on the efficient sum of outer products dictionary learning(SOUPDIL)with the SENSE model,called SOUPDIL-SENSE.The SOUPDIL-SENSE model is mainly solved by utilizing the variable splitting and alternating direction method of multipliers techniques.The experimental results on four human datasets show that the proposed algorithm effectively promotes the image sparsity,eliminates the noise and artifacts of the reconstructed images,and improves the reconstruction accuracy.
基金sponsored by the National Natural Science Foundation of China No.U24B201114,6247070859,62302114 and No.62172353Innovation Fund Program of the Engineering Research Center for Integration and Application of Digital Learning Technology of Ministry of Education No.1331007 and No.1311022Natural Science Foundation of Guangdong Province No.2024A1515010177.
摘要Blockchain,as a distributed ledger,inherently possesses tamper-resistant capabilities,creating a natural channel for covert communication.However,the immutable nature of data storage might introduce challenges to communication security.This study introduces a blockchain-based covert communication model utilizing dynamic Base-K encoding.The proposed encoding scheme utilizes the input address sequence to determine K to encode the secret message and determines the order of transactions based on K,thus ensuring effective concealment of the message.The dynamic encoding parameters enhance flexibility and address issues related to identical transaction amounts for the same secret message.Experimental results demonstrate that the proposed method maintains smooth communication and low susceptibility to tampering,achieving commendable concealment and embedding rates.
基金supported by the Fundamental Research Funds for the Central Universities(No.2024JBZX027)the National Natural Science Foundation of China(No.52375078).
摘要High-Speed Trains (HSTs) have emerged as a mainstream mode of transportation in China, owing to their exceptional safety and efficiency. Ensuring the reliable operation of HSTs is of paramount economic and societal importance. As critical rotating mechanical components of the transmission system, bearings make their fault diagnosis a topic of extensive attention. This paper provides a systematic review of image encoding-based bearing fault diagnosis methods tailored to the condition monitoring of HSTs. First, it categorizes the image encoding techniques applied in the field of bearing fault diagnosis. Then, a review of state-of-the-art studies has been presented, encompassing both monomodal image conversion and multimodal image fusion approaches. Finally, it highlights current challenges and proposes future research directions to advance intelligent fault diagnosis in HSTs, aiming to provide a valuable reference for researchers and engineers in the field of intelligent operation and maintenance.
摘要Retinal blood vessel segmentation is crucial for diagnosing ocular and cardiovascular diseases.Although the introduction of U-Net in 2015 by Olaf Ronneberger significantly advanced this field,yet issues like limited training data,imbalance data distribution,and inadequate feature extraction persist,hindering both the segmentation performance and optimal model generalization.Addressing these critical issues,the DEFFA-Unet is proposed featuring an additional encoder to process domain-invariant pre-processed inputs,thereby improving both richer feature encoding and enhanced model generalization.A feature filtering fusion module is developed to ensure the precise feature filtering and robust hybrid feature fusion.In response to the task-specific need for higher precision where false positives are very costly,traditional skip connections are replaced with the attention-guided feature reconstructing fusion module.Additionally,innovative data augmentation and balancing methods are proposed to counter data scarcity and distribution imbalance,further boosting the robustness and generalization of the model.With a comprehensive suite of evaluation metrics,extensive validations on four benchmark datasets(DRIVE,CHASEDB1,STARE,and HRF)and an SLO dataset(IOSTAR),demonstrate the proposed method’s superiority over both baseline and state-of-the-art models.Particularly the proposed method significantly outperforms the compared methods in cross-validation model generalization.