Accurate quantification of crop residue cover(CRC)is crucial for monitoring and evaluating conservation tillage practices,yet it poses a significant image segmentation challenge.The subtle visual distinctions between ...Accurate quantification of crop residue cover(CRC)is crucial for monitoring and evaluating conservation tillage practices,yet it poses a significant image segmentation challenge.The subtle visual distinctions between fragmented residue and soil,compounded by variable illumination and shadows in field imagery,often lead to poor segmentation performance.To overcome these limitations,we introduce RCTUnet,a novel deep learning architecture designed for robust crop-residue-soil segmentation and precise CRC estimation.RCTUnet’s architecture synergistically integrates three key components:(1)a ResNet50 backbone for deep,multi-scale feature extraction;(2)a convolutional block attention module(CBAM)to adaptively focus on salient residue features across both channel and spatial dimensions;and(3)a transformer-based global context fusion module(GCFM)to model long-range spatial dependencies,which is critical for interpreting heterogeneous residue patterns.We evaluated RCTUnet on a dataset of 1220 field-acquired images spanning four typical crop rotations.Experimental results show that,compared to traditional models:(1)RCTUnet achieves significantly higher crop-residue-soil segmentation accuracy than classic models including Unet,Unet++,DeepLabV3,segmentation network(SegNet),and fully convolutional network(FCN),with improvements of 3.24%,3.42%,4.88%,8.28%,and 6.05%in overall accuracy,respectively;(2)RCTUnet yields superior residue-soil segmentation performance,with increases in residue recall of 7.67%,7.37%,14.09%,27.05%,and 16.91%,respectively;(3)RCTUnet shows enhanced CRC estimation accuracy,achieving a root mean square error(RMSE)of 4.875,representing a 45.5%improvement over Unet(RMSE=8.941).These results demonstrate the efficacy of our hybrid approach,which combines deep hierarchical features,dual-domain attention,and global context modeling.RCTUnet provides a robust and reliable tool for automated CRC assessment,advancing the capabilities of in-field agricultural monitoring.展开更多
Models based on U-shaped networks have achieved widespread success in the field of medical image segmentation,but their performance is generally limited by structural bottlenecks in the network.At this stage,feature m...Models based on U-shaped networks have achieved widespread success in the field of medical image segmentation,but their performance is generally limited by structural bottlenecks in the network.At this stage,feature maps experience a sharp decline in spatial resolution due to continuous downsampling,resulting in significant loss of critical boundaries and structural details.Additionally,the local receptive fields of convolutions limit the effective modelling of global context.To address this core issue,we propose a novel enhanced segmentation network called FMTNet.FMTNet fundamentally enhances the expressive power of deep features by integrating an innovative composite enhancement module at the bottleneck of the U-Net.This module consists of three synergistically working submodules:the Fourier spatial fusion module,which introduces a frequency-domain perspective to compensate for and reconstruct high-frequency structural information lost in the spatial domain;the hybrid mamba–transformer module,which efficiently captures cross-regional long-range dependencies to establish global context and the multi-scale context Aggregation module,which fuses features of different scales to adapt to objects of varying sizes.We conducted extensive experiments on multiple public multi-modal datasets,including colonoscopy polyps,dermatoscopy lesions,breast ultrasound and dental X-rays.The results demonstrate that FMTNet comprehensively outperforms SOTA methods across all key metrics,showcasing exceptional segmentation accuracy and generalisation capabilities.Our research study demonstrates that by synergistically enhancing deep features across three dimensions—frequency,global,and multi-scale—FMTNet provides a general and efficient solution to address the bottleneck issues of U-Net,significantly enhancing the accuracy and robustness of medical image segmentation.The source code and pre-trained weights are available at http://gffzz188fe103f8f1460asqkpuxpv06uqv6xxc.ffgz.tsg.suse.edu.cn/shiguiling0-has/FMTNet.展开更多
Satellite image segmentation plays a crucial role in remote sensing,supporting applications such as environmental monitoring,land use analysis,and disaster management.However,traditional segmentation methods often rel...Satellite image segmentation plays a crucial role in remote sensing,supporting applications such as environmental monitoring,land use analysis,and disaster management.However,traditional segmentation methods often rely on large amounts of labeled data,which are costly and time-consuming to obtain,especially in largescale or dynamic environments.To address this challenge,we propose the Semi-Supervised Multi-View Picture Fuzzy Clustering(SS-MPFC)algorithm,which improves segmentation accuracy and robustness,particularly in complex and uncertain remote sensing scenarios.SS-MPFC unifies three paradigms:semi-supervised learning,multi-view clustering,and picture fuzzy set theory.This integration allows the model to effectively utilize a small number of labeled samples,fuse complementary information from multiple data views,and handle the ambiguity and uncertainty inherent in satellite imagery.We design a novel objective function that jointly incorporates picture fuzzy membership functions across multiple views of the data,and embeds pairwise semi-supervised constraints(must-link and cannot-link)directly into the clustering process to enhance segmentation accuracy.Experiments conducted on several benchmark satellite datasets demonstrate that SS-MPFC significantly outperforms existing state-of-the-art methods in segmentation accuracy,noise robustness,and semantic interpretability.On the Augsburg dataset,SS-MPFC achieves a Purity of 0.8158 and an Accuracy of 0.6860,highlighting its outstanding robustness and efficiency.These results demonstrate that SSMPFC offers a scalable and effective solution for real-world satellite-based monitoring systems,particularly in scenarios where rapid annotation is infeasible,such as wildfire tracking,agricultural monitoring,and dynamic urban mapping.展开更多
Multilevel image segmentation is a critical task in image analysis,which imposes high requirements on the global search capability and convergence efficiency of segmentation algorithms.In this paper,an improved Artifi...Multilevel image segmentation is a critical task in image analysis,which imposes high requirements on the global search capability and convergence efficiency of segmentation algorithms.In this paper,an improved Artificial Protozoa Optimization algorithm,termed the two-stage Taguchi-assisted Gaussian–Levy Artificial Protozoa Optimization(TGAPO)algorithm,is proposed and applied tomultilevel image segmentation.The proposed algorithm adopts a two-stage evolutionary mechanism.In the first stage,Gaussian perturbation is introduced to enhance local search capability;in the second stage,Levy flight is incorporated to expand the global search range;and finally,the Taguchi strategy is employed to further refine the optimal solution.Consequently,the global optimization performance and robustness of the algorithm are significantly improved.To evaluate the effectiveness of the proposed TGAPO algorithm,comparative experiments are conducted with representative optimization algorithms,including the Grey Wolf Optimizer(GWO)and Particle Swarm Optimization(PSO),in the context ofmultilevel image segmentation.The segmentation quality is assessed using the minimum cross-entropy function as the performance metric.Experimental results demonstrate that the TGAPO algorithm outperforms the comparison algorithms in terms of segmentation accuracy and convergence speed,and exhibits superior stability in high-threshold segmentation tasks.Furthermore,the proposedmethod achieves excellentmulti-threshold segmentation performance for color images and shows strong potential for practical applications.展开更多
Microscopy imaging is fundamental in analyzing bacterial morphology and dynamics,offering critical insights into bacterial physiology and pathogenicity.Image segmentation techniques enable quantitative analysis of bac...Microscopy imaging is fundamental in analyzing bacterial morphology and dynamics,offering critical insights into bacterial physiology and pathogenicity.Image segmentation techniques enable quantitative analysis of bacterial structures,facilitating precise measurement of morphological variations and population behaviors at single-cell resolution.This paper reviews advancements in bacterial image segmentation,emphasizing the shift from traditional thresholding and watershed methods to deep learning-driven approaches.Convolutional neural networks(CNNs),U-Net architectures,and three-dimensional(3D)frameworks excel at segmenting dense biofilms and resolving antibiotic-induced morphological changes.These methods combine automated feature extraction with physics-informed postprocessing.Despite progress,challenges persist in computational efficiency,cross-species generalizability,and integration with multimodal experimental workflows.Future progress will depend on improving model robustness across species and imaging modalities,integrating multimodal data for phenotype-function mapping,and developing standard pipelines that link computational tools with clinical diagnostics.These innovations will expand microbial phenotyping beyond structural analysis,enabling deeper insights into bacterial physiology and ecological interactions.展开更多
This research introduces an innovative lightweight image segmentation framework where models of hybrid architectures work together to predict the output and also have self-adapting ability,along with maintaining data ...This research introduces an innovative lightweight image segmentation framework where models of hybrid architectures work together to predict the output and also have self-adapting ability,along with maintaining data privacy.In this framework,data is distributed and trained in a decentralized way using different deep learning architectures.That is how the advantages of all these models will be integrated into the system.Each trained model makes its own prediction,and the final output is determined through cooperation among these models.Here,the confidence-level and pixel-wise voting majority algorithms will be utilized for the co-operation-based output prediction system.Due to the efficient setup of the operations of these two algorithms,each input will get its accurate output.Additionally,the federated learning-based self-adapting feature facilitated the proposed framework for advancing its performance consistently by interacting with the inputs.Here,UNet,SegNet and FCNN models have been trained and integrated into the prediction framework.Here,the Oxford-IIT pet dataset was used.And all the data of this dataset is distributed among these three models.The framework’s effectiveness was measured using metrics like average pixel accuracy,IoU,F1 score,precision,and recall,which resulted in scores of 89.26%,71.48%,81.29%,83.96%and 81.16%,respectively.Another notable feature of this proposed framework is allocating comparatively fewer computational resources and taking less time.To validate these claims,the proposed system is compared with three other state-of-the-art models,and the proposed system delivered superior performance among all.展开更多
In this study,we present a novel approach to multi-threshold image segmentation using an adaptive method that combines the Ebola Optimization Search Algorithm(EOSA)with the Aquila Optimizer,termed the Integrated Enhan...In this study,we present a novel approach to multi-threshold image segmentation using an adaptive method that combines the Ebola Optimization Search Algorithm(EOSA)with the Aquila Optimizer,termed the Integrated Enhanced Ebola Optimization Search Algorithm(IEOSA).Our approach leverages this integration to produce high-quality segmented images.The IEOSA method introduces two distinct optimization mechanisms to identify optimal solutions.By blending the randomness of the Aquila Optimizer with the capabilities of EOSA,we enhance the exploration potential of the algorithm.Additionally,we incorporate a self-transition learning system within the IEOSA to further boost its performance.To tackle multi-level threshold image segmentation,we apply Kapur’s entropy between-class variance within the IEOSA framework.Our findings show that the IEOSA-based techniques outperform other comparable methods,offering faster convergence and more stable segmentation results.Through comparative analysis using standard test images,we demonstrate that IEOSA achieves higher solution accuracy than other methods.Ultimately,the proposed IEOSA methodologies effectively address multi-level threshold image segmentation challenges,accurately segmenting even the minor errors that are often overlooked in high-resolution images.展开更多
Medical image segmentation is of critical importance in the domain of contemporary medical imaging.However,U-Net and its variants exhibit limitations in capturing complex nonlinear patterns and global contextual infor...Medical image segmentation is of critical importance in the domain of contemporary medical imaging.However,U-Net and its variants exhibit limitations in capturing complex nonlinear patterns and global contextual information.Although the subsequent U-KAN model enhances nonlinear representation capabilities,it still faces challenges such as gradient vanishing during deep network training and spatial detail loss during feature downsampling,resulting in insufficient segmentation accuracy for edge structures and minute lesions.To address these challenges,this paper proposes the RE-UKAN model,which innovatively improves upon U-KAN.Firstly,a residual network is introduced into the encoder to effectively mitigate gradient vanishing through cross-layer identity mappings,thus enhancing modelling capabilities for complex pathological structures.Secondly,Efficient Local Attention(ELA)is integrated to suppress spatial detail loss during downsampling,thereby improving the perception of edge structures and minute lesions.Experimental results on four public datasets demonstrate that RE-UKAN outperforms existing medical image segmentation methods across multiple evaluation metrics,with particularly outstanding performance on the TN-SCUI 2020 dataset,achieving IoU of 88.18%and Dice of 93.57%.Compared to the baseline model,it achieves improvements of 3.05%and 1.72%,respectively.These results fully demonstrate RE-UKAN’s superior detail retention capability and boundary recognition accuracy in complex medical image segmentation tasks,providing a reliable solution for clinical precision segmentation.展开更多
This article studies the problem of image segmentation-based semantic communication in autonomous driving.In real traffic scenes,the detecting of objects(e.g.,vehicles and pedestrians)is more important to guarantee dr...This article studies the problem of image segmentation-based semantic communication in autonomous driving.In real traffic scenes,the detecting of objects(e.g.,vehicles and pedestrians)is more important to guarantee driving safety,which is always ignored in existing works.Therefore,we propose a vehicular image segmentation-oriented semantic communication system,termed VIS-SemCom,focusing on transmitting and recovering image semantic features of high-important objects to reduce transmission redundancy.First,we develop a semantic codec based on Swin Transformer architecture,which expands the perceptual field thus improving the segmentation accuracy.To highlight the important objects'accuracy,we propose a multi-scale semantic extraction method by assigning the number of Swin Transformer blocks for diverse resolution semantic features.Also,an importance-aware loss incorporating important levels is devised,and an online hard example mining(OHEM)strategy is proposed to handle small sample issues in the dataset.Finally,experimental results demonstrate that the proposed VIS-SemCom can achieve a significant mean intersection over union(mIoU)performance in the SNR regions,a reduction of transmitted data volume by about 60%at 60%mIoU,and improve the segmentation accuracy of important objects,compared to baseline image communication.展开更多
U-Net,a fully convolutional neural network(FCNN)with U-shaped features,has demonstrated significant success in biomedical image segmentation.However,the locality of convolution operations in the U-Net limits its abili...U-Net,a fully convolutional neural network(FCNN)with U-shaped features,has demonstrated significant success in biomedical image segmentation.However,the locality of convolution operations in the U-Net limits its ability to learn long-range dependencies.Transformers,originally developed for natural language processing,have recently been adapted for image segmentation because of their global self-attention mechanisms.Inspired by the long-range feature learning capability of transformers,we propose Dense-Transformer(DenT),an architecture designed for volumetric microscopy image segmentation.DenT incorporates transformers as encoders within each convolutional layer to capture global contextual information.Additionally,dense skip connections at multiple resolutions enhance feature propagation,enabling precise localization.We evaluated DenT on mitochondrial segmentation using our confocal microscopy dataset and a public fluorescence microscope dataset from the Allen Institute for Cell Science.The experimental results demonstrate that DenT incrementally improves the segmentation of mitochondria and mitochondrial DNA substructures from transmitted light microscopy images.DenT offers a tool for visualization,measurement,and analysis of mitochondrial morphology and mitochondrial DNA in label-free microscopy.展开更多
Automatic and accurate medical image segmentation remains a fundamental task in computer-aided diagnosis and treatment planning.Recent advances in foundation models,such as the medical-focused Segment AnythingModel(Me...Automatic and accurate medical image segmentation remains a fundamental task in computer-aided diagnosis and treatment planning.Recent advances in foundation models,such as the medical-focused Segment AnythingModel(MedSAM),have demonstrated strong performance but face challenges inmanymedical applications due to anatomical complexity and a limited domain-specific prompt.Thiswork introduces amethodology that enhances segmentation robustness and precision by automatically generating multiple informative point prompts,rather than relying on single inputs.The proposed approach randomly samples sets of spatially distributed point prompts based on image features,enabling MedSAM to better capture fine-grained anatomical structures and boundaries.During inference,probability maps are aggregated to reduce local misclassifications without additional model training.Extensive experiments on various computed tomography(CT)and magnetic resonance imaging(MRI)datasets demonstrate improvements in Dice Similarity Coefficient(DSC)and Normalized Surface Dice(NSD)metrics compared to baseline SAM and Scribble Prompt models.A semi-automatic point sampling version based on the ground truth segmentations yielded enhanced results,achieving up to 92.1%DSC and 86.6%NSD,with significant gains in delineating complex organs such as the pancreas,colon,kidney,and brain tumours.The main novelty of our method consists of effectively combining the results of multiple point prompts into the medical segmentation pipeline so that single-point prompt methods are outperformed.Overall,the proposed model offers a straightforward yet effective approach to improve medical image segmentation performance while maintaining computational efficiency.展开更多
Traditional Mamba-UNet integrations employ four-stage architectures,replacing conventional five-stage UNets with VMamba blocks for global dependency modeling.Unlike Transformers,which suffer from quadratic complexity ...Traditional Mamba-UNet integrations employ four-stage architectures,replacing conventional five-stage UNets with VMamba blocks for global dependency modeling.Unlike Transformers,which suffer from quadratic complexity and high memory consumption in self-attention,Mamba-UNet achieves efficient global modeling through linear-complexity state space modeling.This paper proposes TriLVM-UNet,a lightweight three-stage architecture that integrates parameter-efficient VMamba blocks and enhances cross-stage feature interaction via an improved skip-attention bridge(SAB)module inspired by UltraLight VM-UNet.The model incorporates a Lightweight Vision Mamba(LVM)layer for high-resolution feature extraction,alongside multi-scale dilated convolution(MSDC)and convolutional block attention module(CBAM)for enhanced feature fusion.Evaluated on the 3D ACDC dataset against six baseline models,TriLVM-UNet achieves 98.57%accuracy.The GitHub repository is available at:http://gffzz188fe103f8f1460asqkpuxpv06uqv6xxc.ffgz.tsg.suse.edu.cn/730432ch/TriLVM-UNet.展开更多
This study presents a novel hybrid optimization model that combines the complementary aspects of Sand Cat Swarm Optimization(SCSO)and Whale Optimization Algorithm(WOA)to solve the multi-level image thresholding proble...This study presents a novel hybrid optimization model that combines the complementary aspects of Sand Cat Swarm Optimization(SCSO)and Whale Optimization Algorithm(WOA)to solve the multi-level image thresholding problem.The proposed approach utilizes an adaptive two-stage mechanism that balances the high exploration capacity of SCSO with the concentrated local search capability of WOA,aiming to maximize inter-class variance in the histogram-based thresholding process.Various experiments are conducted on lung cancer,prostate,and mixed medical image datasets.Results demonstrate that modified SCSOWOA achieves superior performance across all datasets.For LC25000,it attains PSNR 27.9453 dB,SSIM 0.9340,FSIM 0.9542,Dice coefficient 0.8901,and Jaccard index 0.8034 at T=12.For prostate,PSNR reaches 28.3965 dB,SSIM 0.7532,FSIM 0.8170,Dice 0.9215,and Jaccard 0.8593.In the MSD dataset,SCSOWOA achieves PSNR 29.3244 dB,SSIM 0.7118,FSIM 0.7562,Dice 0.8901,and Jaccard 0.8034,indicating consistent performance across diverse organs and modalities.The method also demonstrates high computational efficiency,with an average execution time of 1.3221 s,offering up to 40%speed improvement over conventional metaheuristics such as PSO and GWO.Overall,proposed method provides high-accuracy,low-variance,and computationally efficient segmentation,preserving both structural and perceptual fidelity.These results confirm the method’s robustness,generalizability,and practical applicability for AI-assisted diagnostic systems across histopathological and medical imaging contexts,balancing precision,structural preservation,and speed for real-world clinical deployment.展开更多
The Transformer has achieved great success in the field of medical image segmentation,but its quadratic computational complexity limits its application in dense medical image prediction.Recently,the receptance weighte...The Transformer has achieved great success in the field of medical image segmentation,but its quadratic computational complexity limits its application in dense medical image prediction.Recently,the receptance weighted key value(RWKV)architecture has garnered widespread attention due to its linear computational complexity and its capability of parallel computation during training.Despite the RWKV model's proficiency in addressing long-range modeling tasks with linear computational complexity,most current RWKV-based approaches employ static scanning patterns.These patterns may inadvertently incorporate biased prior knowledge into the model's predictions.To address this challenge,we propose a multi-head scan strategy combined with padding methods to effectively simulate spatial continuity in 2D images.Within the Feature Aggregation Attention(FAA)module,asymmetric convolutions are designed to aggregate 1D sequence features along a single dimension,thereby expanding effective receptive fields while preserving structural sparsity.Additionally,panoramic token shift(P-Shift)effectively models local dependency relationships by moving tokens from a wide receptive field.Extensive experiments conducted on the ISIC17/18 and ACDC datasets demonstrate that our method exhibits superior performance in dense medical image prediction tasks.展开更多
The development of oil and gas is constrained by difficulties in dynamically characterizing pore structures.Traditional methods inadequately represent the complex interactions between mineral dissolution,precipitation...The development of oil and gas is constrained by difficulties in dynamically characterizing pore structures.Traditional methods inadequately represent the complex interactions between mineral dissolution,precipitation,and fluid flow.This study addresses these gaps by introducing a Transformer U-Neural Network(TransUNet)for computed tomography(CT)image segmentation.The integrated workflow combines conventional CT(Resolution of 5.4μm)and synchrotron radiation CT(Resolution of0.8μm)for dynamic flooding,imaging,segmentation,and precise 3D pore network extraction,overcoming resolution limits.TransUNet's strong global attention and feature extraction reduce overfitting and deliver high-accuracy segmentation of minerals,pores,and argillaceous microporous networks(AMN),achieving 74.92%intersection over union(IoU)for AMN.A porosity correction method improves conventional CT porosity accuracy to 94%of gas-measured values.Alkaline flooding experiments reveal:(1)initial clay swelling reduces small pore size by~50%as alkaline ions destabilize clay;(2)mineral dissolution,such as dolomite,creates secondary pores,increasing 80μm pores by 1.8 times;(3)silicate dissolution increases porosity and leads to a 93.7%rise in permeability.Clay reorganization enhances the AMN by 46.1%.The pore size distribution shifts to log-no rmal at steady state,and throat connectivity improves flow capacity.This work pioneers Transformer-based CT image segmentation,introduces cross-resolution prediction,and clarifies pore regulation by mineral phase changes,establishing a new paradigm for chemical flooding in sandstone reservoirs.展开更多
The rising need for precision farming and sustainable land management has catalyzed the requirement for sophisticated means of deriving practical data from remote sensing images.Image segmentation,or the process of di...The rising need for precision farming and sustainable land management has catalyzed the requirement for sophisticated means of deriving practical data from remote sensing images.Image segmentation,or the process of dividing the image into semantically relevant parts,has become a groundbreaking technology that allows resolving the problem of transitioning the pixel-level data to a parcel-level analysis.This review is a synthesis of the segmentation methods and their use in crop research and geospatial science.The architectures of pixel-based,object-based,and deep learning(convolutional neural networks,U-Net,Mask R-CNN,and Transformer models)are considered in terms of principles,capabilities,and limitations.Multi-spectral,hyperspectral,LiDAR,and SAR data are integrated to improve the efficiency of segmentation,allowing the possible delineation of fields,the classification of crops,health monitoring,monitoring of yields,and stress identification.In addition to agriculture,segmentation helps in land use and land cover mapping,identification of temporal change,monitoring of the environment,and is used in combination with GIS-based spatial modeling.Nevertheless,issues related to data heterogeneity,mixed pixels,computational requirements,and inadequate availability of labelled data still exist despite the major progress.The future directions involve multi-source data fusion,pixel-to-parcel pipeline automation,and predictive models based on AI,which are used to enhance its scalability,robustness,and the ability to monitor in real-time.This review makes it clear that the use of image segmentation as a tool in generating precision agriculture,sustainable land use,and informed geospatial.展开更多
In the field of medical image processing,combining global and local relationship modeling constitutes an effective strategy for precise segmentation.Prior research has established the validity of Convolutional Neural ...In the field of medical image processing,combining global and local relationship modeling constitutes an effective strategy for precise segmentation.Prior research has established the validity of Convolutional Neural Networks(CNN)in modeling local relationships.Conversely,Transformers have demonstrated their capability to effectively capture global contextual information.However,when utilized to address CNNs’limitations in modeling global relationships,Transformers are hindered by substantial computational complexity.To address this issue,we introduce Mamba,a State-Space Model(SSM)that exhibits exceptional proficiency in modeling long-range dependencies in sequential data.Given Mamba’s demonstrated potential in 2D medical image segmentation in previous studies,we have designed a Dual-encoder Global-local Feature Extraction Network based on Mamba,termed DGFE-Mamba,to accurately capture and fuse long-range dependencies and local dependencies within multi-scale features.Compared to Transformer-based methods,the DGFE-Mamba model excels in comprehensive feature modeling and demonstrates significantly improved segmentation accuracy.To validate the effectiveness and practicality of DGFE-Mamba,we conducted tests on the Automatic Cardiac Diagnosis Challenge(ACDC)dataset,the Synapse multi-organ CT abdominal segmentation dataset,and the Colorectal Cancer Clinic(CVC-ClinicDB)dataset.The results showed that DGFE-Mamba achieved Dice coefficients of 92.20,83.67,and 94.13,respectively.These findings comprehensively validate the effectiveness and practicality of the proposed DGFE-Mamba architecture.展开更多
Convolutional neural networks(CNNs)-based medical image segmentation technologies have been widely used in medical image segmentation because of their strong representation and generalization abilities.However,due to ...Convolutional neural networks(CNNs)-based medical image segmentation technologies have been widely used in medical image segmentation because of their strong representation and generalization abilities.However,due to the inability to effectively capture global information from images,CNNs can easily lead to loss of contours and textures in segmentation results.Notice that the transformer model can effectively capture the properties of long-range dependencies in the image,and furthermore,combining the CNN and the transformer can effectively extract local details and global contextual features of the image.Motivated by this,we propose a multi-branch and multi-scale attention network(M2ANet)for medical image segmentation,whose architecture consists of three components.Specifically,in the first component,we construct an adaptive multi-branch patch module for parallel extraction of image features to reduce information loss caused by downsampling.In the second component,we apply residual block to the well-known convolutional block attention module to enhance the network’s ability to recognize important features of images and alleviate the phenomenon of gradient vanishing.In the third component,we design a multi-scale feature fusion module,in which we adopt adaptive average pooling and position encoding to enhance contextual features,and then multi-head attention is introduced to further enrich feature representation.Finally,we validate the effectiveness and feasibility of the proposed M2ANet method through comparative experiments on four benchmark medical image segmentation datasets,particularly in the context of preserving contours and textures.展开更多
Medical image segmentation is a crucial task in clinical applications.However,obtaining labeled data for medical images is often challenging.This has led to the appeal of semi-supervised learning(SSL),a technique adep...Medical image segmentation is a crucial task in clinical applications.However,obtaining labeled data for medical images is often challenging.This has led to the appeal of semi-supervised learning(SSL),a technique adept at leveraging a modest amount of labeled data.Nonetheless,most prevailing SSL segmentation methods for medical images either rely on the single consistency training method or directly fine-tune SSL methods designed for natural images.In this paper,we propose an innovative semi-supervised method called multi-consistency training(MCT)for medical image segmentation.Our approach transcends the constraints of prior methodologies by considering consistency from a dual perspective:output consistency across different up-sampling methods and output consistency of the same data within the same network under various perturbations to the intermediate features.We design distinct semi-supervised loss regression methods for these two types of consistencies.To enhance the application of our MCT model,we also develop a dedicated decoder as the core of our neural network.Thorough experiments were conducted on the polyp dataset and the dental dataset,rigorously compared against other SSL methods.Experimental results demonstrate the superiority of our approach,achieving higher segmentation accuracy.Moreover,comprehensive ablation studies and insightful discussion substantiate the efficacy of our approach in navigating the intricacies of medical image segmentation.展开更多
Existing semi-supervisedmedical image segmentation algorithms use copy-paste data augmentation to correct the labeled-unlabeled data distribution mismatch.However,current copy-paste methods have three limitations:(1)t...Existing semi-supervisedmedical image segmentation algorithms use copy-paste data augmentation to correct the labeled-unlabeled data distribution mismatch.However,current copy-paste methods have three limitations:(1)training the model solely with copy-paste mixed pictures from labeled and unlabeled input loses a lot of labeled information;(2)low-quality pseudo-labels can cause confirmation bias in pseudo-supervised learning on unlabeled data;(3)the segmentation performance in low-contrast and local regions is less than optimal.We design a Stochastic Augmentation-Based Dual-Teaching Auxiliary Training Strategy(SADT),which enhances feature diversity and learns high-quality features to overcome these problems.To be more precise,SADT trains the Student Network by using pseudo-label-based training from Teacher Network 1 and supervised learning with labeled data,which prevents the loss of rare labeled data.We introduce a bi-directional copy-pastemask with progressive high-entropy filtering to reduce data distribution disparities and mitigate confirmation bias in pseudo-supervision.For the mixed images,Deep-Shallow Spatial Contrastive Learning(DSSCL)is proposed in the feature spaces of Teacher Network 2 and the Student Network to improve the segmentation capabilities in low-contrast and local areas.In this procedure,the features retrieved by the Student Network are subjected to a random feature perturbation technique.On two openly available datasets,extensive trials show that our proposed SADT performs much better than the state-ofthe-art semi-supervised medical segmentation techniques.Using only 10%of the labeled data for training,SADT was able to acquire a Dice score of 90.10%on the ACDC(Automatic Cardiac Diagnosis Challenge)dataset.展开更多
基金supported by the National Natural Science Foundation of China(No.42101362)the Natural Science Foundation of Henan Province(No.252300421158)+1 种基金the Shenzhen Science and Technology Program(No.JCYJ20220530162001003)the Science and Technology Development Program of Henan Province(No.242300421639),China。
摘要Accurate quantification of crop residue cover(CRC)is crucial for monitoring and evaluating conservation tillage practices,yet it poses a significant image segmentation challenge.The subtle visual distinctions between fragmented residue and soil,compounded by variable illumination and shadows in field imagery,often lead to poor segmentation performance.To overcome these limitations,we introduce RCTUnet,a novel deep learning architecture designed for robust crop-residue-soil segmentation and precise CRC estimation.RCTUnet’s architecture synergistically integrates three key components:(1)a ResNet50 backbone for deep,multi-scale feature extraction;(2)a convolutional block attention module(CBAM)to adaptively focus on salient residue features across both channel and spatial dimensions;and(3)a transformer-based global context fusion module(GCFM)to model long-range spatial dependencies,which is critical for interpreting heterogeneous residue patterns.We evaluated RCTUnet on a dataset of 1220 field-acquired images spanning four typical crop rotations.Experimental results show that,compared to traditional models:(1)RCTUnet achieves significantly higher crop-residue-soil segmentation accuracy than classic models including Unet,Unet++,DeepLabV3,segmentation network(SegNet),and fully convolutional network(FCN),with improvements of 3.24%,3.42%,4.88%,8.28%,and 6.05%in overall accuracy,respectively;(2)RCTUnet yields superior residue-soil segmentation performance,with increases in residue recall of 7.67%,7.37%,14.09%,27.05%,and 16.91%,respectively;(3)RCTUnet shows enhanced CRC estimation accuracy,achieving a root mean square error(RMSE)of 4.875,representing a 45.5%improvement over Unet(RMSE=8.941).These results demonstrate the efficacy of our hybrid approach,which combines deep hierarchical features,dual-domain attention,and global context modeling.RCTUnet provides a robust and reliable tool for automated CRC assessment,advancing the capabilities of in-field agricultural monitoring.
基金funded by UKRI(Grants EP/W020408/1 and RS718)through Doctoral Training Centre at Swansea UniversityQilu Medical Talent Cultivation Project of Shandong Health Commission(Grant[2023]78)National Natural Science Foundation of China(Grant 82405459)。
摘要Models based on U-shaped networks have achieved widespread success in the field of medical image segmentation,but their performance is generally limited by structural bottlenecks in the network.At this stage,feature maps experience a sharp decline in spatial resolution due to continuous downsampling,resulting in significant loss of critical boundaries and structural details.Additionally,the local receptive fields of convolutions limit the effective modelling of global context.To address this core issue,we propose a novel enhanced segmentation network called FMTNet.FMTNet fundamentally enhances the expressive power of deep features by integrating an innovative composite enhancement module at the bottleneck of the U-Net.This module consists of three synergistically working submodules:the Fourier spatial fusion module,which introduces a frequency-domain perspective to compensate for and reconstruct high-frequency structural information lost in the spatial domain;the hybrid mamba–transformer module,which efficiently captures cross-regional long-range dependencies to establish global context and the multi-scale context Aggregation module,which fuses features of different scales to adapt to objects of varying sizes.We conducted extensive experiments on multiple public multi-modal datasets,including colonoscopy polyps,dermatoscopy lesions,breast ultrasound and dental X-rays.The results demonstrate that FMTNet comprehensively outperforms SOTA methods across all key metrics,showcasing exceptional segmentation accuracy and generalisation capabilities.Our research study demonstrates that by synergistically enhancing deep features across three dimensions—frequency,global,and multi-scale—FMTNet provides a general and efficient solution to address the bottleneck issues of U-Net,significantly enhancing the accuracy and robustness of medical image segmentation.The source code and pre-trained weights are available at http://gffzz188fe103f8f1460asqkpuxpv06uqv6xxc.ffgz.tsg.suse.edu.cn/shiguiling0-has/FMTNet.
基金funded by the Research Project:THTETN.05/24-25,Vietnam Academy of Science and Technology.
摘要Satellite image segmentation plays a crucial role in remote sensing,supporting applications such as environmental monitoring,land use analysis,and disaster management.However,traditional segmentation methods often rely on large amounts of labeled data,which are costly and time-consuming to obtain,especially in largescale or dynamic environments.To address this challenge,we propose the Semi-Supervised Multi-View Picture Fuzzy Clustering(SS-MPFC)algorithm,which improves segmentation accuracy and robustness,particularly in complex and uncertain remote sensing scenarios.SS-MPFC unifies three paradigms:semi-supervised learning,multi-view clustering,and picture fuzzy set theory.This integration allows the model to effectively utilize a small number of labeled samples,fuse complementary information from multiple data views,and handle the ambiguity and uncertainty inherent in satellite imagery.We design a novel objective function that jointly incorporates picture fuzzy membership functions across multiple views of the data,and embeds pairwise semi-supervised constraints(must-link and cannot-link)directly into the clustering process to enhance segmentation accuracy.Experiments conducted on several benchmark satellite datasets demonstrate that SS-MPFC significantly outperforms existing state-of-the-art methods in segmentation accuracy,noise robustness,and semantic interpretability.On the Augsburg dataset,SS-MPFC achieves a Purity of 0.8158 and an Accuracy of 0.6860,highlighting its outstanding robustness and efficiency.These results demonstrate that SSMPFC offers a scalable and effective solution for real-world satellite-based monitoring systems,particularly in scenarios where rapid annotation is infeasible,such as wildfire tracking,agricultural monitoring,and dynamic urban mapping.
摘要Multilevel image segmentation is a critical task in image analysis,which imposes high requirements on the global search capability and convergence efficiency of segmentation algorithms.In this paper,an improved Artificial Protozoa Optimization algorithm,termed the two-stage Taguchi-assisted Gaussian–Levy Artificial Protozoa Optimization(TGAPO)algorithm,is proposed and applied tomultilevel image segmentation.The proposed algorithm adopts a two-stage evolutionary mechanism.In the first stage,Gaussian perturbation is introduced to enhance local search capability;in the second stage,Levy flight is incorporated to expand the global search range;and finally,the Taguchi strategy is employed to further refine the optimal solution.Consequently,the global optimization performance and robustness of the algorithm are significantly improved.To evaluate the effectiveness of the proposed TGAPO algorithm,comparative experiments are conducted with representative optimization algorithms,including the Grey Wolf Optimizer(GWO)and Particle Swarm Optimization(PSO),in the context ofmultilevel image segmentation.The segmentation quality is assessed using the minimum cross-entropy function as the performance metric.Experimental results demonstrate that the TGAPO algorithm outperforms the comparison algorithms in terms of segmentation accuracy and convergence speed,and exhibits superior stability in high-threshold segmentation tasks.Furthermore,the proposedmethod achieves excellentmulti-threshold segmentation performance for color images and shows strong potential for practical applications.
基金financially supported by the Open Project Program of Wuhan National Laboratory for Optoelectronics(No.2022WNLOKF009)the National Natural Science Foundation of China(No.62475216)+2 种基金the Key Research and Development Program of Shaanxi(No.2024GH-ZDXM-37)the Fujian Provincial Natural Science Foundation of China(No.2024J01060)the Startup Program of XMU,and the Fundamental Research Funds for the Central Universities.
摘要Microscopy imaging is fundamental in analyzing bacterial morphology and dynamics,offering critical insights into bacterial physiology and pathogenicity.Image segmentation techniques enable quantitative analysis of bacterial structures,facilitating precise measurement of morphological variations and population behaviors at single-cell resolution.This paper reviews advancements in bacterial image segmentation,emphasizing the shift from traditional thresholding and watershed methods to deep learning-driven approaches.Convolutional neural networks(CNNs),U-Net architectures,and three-dimensional(3D)frameworks excel at segmenting dense biofilms and resolving antibiotic-induced morphological changes.These methods combine automated feature extraction with physics-informed postprocessing.Despite progress,challenges persist in computational efficiency,cross-species generalizability,and integration with multimodal experimental workflows.Future progress will depend on improving model robustness across species and imaging modalities,integrating multimodal data for phenotype-function mapping,and developing standard pipelines that link computational tools with clinical diagnostics.These innovations will expand microbial phenotyping beyond structural analysis,enabling deeper insights into bacterial physiology and ecological interactions.
基金funded by Multimedia University,Cyberjaya,Selangor,Malaysia(Grant Number:PostDoc(MMUI/240029)).
摘要This research introduces an innovative lightweight image segmentation framework where models of hybrid architectures work together to predict the output and also have self-adapting ability,along with maintaining data privacy.In this framework,data is distributed and trained in a decentralized way using different deep learning architectures.That is how the advantages of all these models will be integrated into the system.Each trained model makes its own prediction,and the final output is determined through cooperation among these models.Here,the confidence-level and pixel-wise voting majority algorithms will be utilized for the co-operation-based output prediction system.Due to the efficient setup of the operations of these two algorithms,each input will get its accurate output.Additionally,the federated learning-based self-adapting feature facilitated the proposed framework for advancing its performance consistently by interacting with the inputs.Here,UNet,SegNet and FCNN models have been trained and integrated into the prediction framework.Here,the Oxford-IIT pet dataset was used.And all the data of this dataset is distributed among these three models.The framework’s effectiveness was measured using metrics like average pixel accuracy,IoU,F1 score,precision,and recall,which resulted in scores of 89.26%,71.48%,81.29%,83.96%and 81.16%,respectively.Another notable feature of this proposed framework is allocating comparatively fewer computational resources and taking less time.To validate these claims,the proposed system is compared with three other state-of-the-art models,and the proposed system delivered superior performance among all.
基金King Saud University,Saudi Arabia for funding this work through Ongoing Research Funding Program,(ORF-2026-704)supported by the National Natural Science Foundation of China under Grant 62471493+1 种基金partially supported by the Natural Science Foundation of Shandong Province under Grant ZR2023LZH017,ZR2024MF066partially supported by the National Vocational Education Teacher Teaching Innovation Team Characteristic Project under Grant CXTD003.
摘要In this study,we present a novel approach to multi-threshold image segmentation using an adaptive method that combines the Ebola Optimization Search Algorithm(EOSA)with the Aquila Optimizer,termed the Integrated Enhanced Ebola Optimization Search Algorithm(IEOSA).Our approach leverages this integration to produce high-quality segmented images.The IEOSA method introduces two distinct optimization mechanisms to identify optimal solutions.By blending the randomness of the Aquila Optimizer with the capabilities of EOSA,we enhance the exploration potential of the algorithm.Additionally,we incorporate a self-transition learning system within the IEOSA to further boost its performance.To tackle multi-level threshold image segmentation,we apply Kapur’s entropy between-class variance within the IEOSA framework.Our findings show that the IEOSA-based techniques outperform other comparable methods,offering faster convergence and more stable segmentation results.Through comparative analysis using standard test images,we demonstrate that IEOSA achieves higher solution accuracy than other methods.Ultimately,the proposed IEOSA methodologies effectively address multi-level threshold image segmentation challenges,accurately segmenting even the minor errors that are often overlooked in high-resolution images.
摘要Medical image segmentation is of critical importance in the domain of contemporary medical imaging.However,U-Net and its variants exhibit limitations in capturing complex nonlinear patterns and global contextual information.Although the subsequent U-KAN model enhances nonlinear representation capabilities,it still faces challenges such as gradient vanishing during deep network training and spatial detail loss during feature downsampling,resulting in insufficient segmentation accuracy for edge structures and minute lesions.To address these challenges,this paper proposes the RE-UKAN model,which innovatively improves upon U-KAN.Firstly,a residual network is introduced into the encoder to effectively mitigate gradient vanishing through cross-layer identity mappings,thus enhancing modelling capabilities for complex pathological structures.Secondly,Efficient Local Attention(ELA)is integrated to suppress spatial detail loss during downsampling,thereby improving the perception of edge structures and minute lesions.Experimental results on four public datasets demonstrate that RE-UKAN outperforms existing medical image segmentation methods across multiple evaluation metrics,with particularly outstanding performance on the TN-SCUI 2020 dataset,achieving IoU of 88.18%and Dice of 93.57%.Compared to the baseline model,it achieves improvements of 3.05%and 1.72%,respectively.These results fully demonstrate RE-UKAN’s superior detail retention capability and boundary recognition accuracy in complex medical image segmentation tasks,providing a reliable solution for clinical precision segmentation.
基金National Natural Science Foundation of China under Grants No.62171047,U22B2001,62271065,62001051Beijing Natural Science Foundation under Grant L223027BUPT Excellent Ph.D Students Foundation under Grants CX2021114。
摘要This article studies the problem of image segmentation-based semantic communication in autonomous driving.In real traffic scenes,the detecting of objects(e.g.,vehicles and pedestrians)is more important to guarantee driving safety,which is always ignored in existing works.Therefore,we propose a vehicular image segmentation-oriented semantic communication system,termed VIS-SemCom,focusing on transmitting and recovering image semantic features of high-important objects to reduce transmission redundancy.First,we develop a semantic codec based on Swin Transformer architecture,which expands the perceptual field thus improving the segmentation accuracy.To highlight the important objects'accuracy,we propose a multi-scale semantic extraction method by assigning the number of Swin Transformer blocks for diverse resolution semantic features.Also,an importance-aware loss incorporating important levels is devised,and an online hard example mining(OHEM)strategy is proposed to handle small sample issues in the dataset.Finally,experimental results demonstrate that the proposed VIS-SemCom can achieve a significant mean intersection over union(mIoU)performance in the SNR regions,a reduction of transmitted data volume by about 60%at 60%mIoU,and improve the segmentation accuracy of important objects,compared to baseline image communication.
基金Taiwan University Center for Advanced Computing and Imaging in Biomedicine(NTU-114L900701).
摘要U-Net,a fully convolutional neural network(FCNN)with U-shaped features,has demonstrated significant success in biomedical image segmentation.However,the locality of convolution operations in the U-Net limits its ability to learn long-range dependencies.Transformers,originally developed for natural language processing,have recently been adapted for image segmentation because of their global self-attention mechanisms.Inspired by the long-range feature learning capability of transformers,we propose Dense-Transformer(DenT),an architecture designed for volumetric microscopy image segmentation.DenT incorporates transformers as encoders within each convolutional layer to capture global contextual information.Additionally,dense skip connections at multiple resolutions enhance feature propagation,enabling precise localization.We evaluated DenT on mitochondrial segmentation using our confocal microscopy dataset and a public fluorescence microscope dataset from the Allen Institute for Cell Science.The experimental results demonstrate that DenT incrementally improves the segmentation of mitochondria and mitochondrial DNA substructures from transmitted light microscopy images.DenT offers a tool for visualization,measurement,and analysis of mitochondrial morphology and mitochondrial DNA in label-free microscopy.
基金supported by the Autonomous Government of Andalusia(Spain)under project UMA20-FEDERJA-108also by the Ministry of Science and Innovation of Spain,grant number PID2022-136764OA-I00+1 种基金It includes funds fromthe European Regional Development Fund(ERDF),It is also partially supported by the Fundación Unicaja(PUNI-003_2023)the Instituto de Investigación Biomédica de Málaga y Plataforma en Nanomedicina-IBIMA Plataforma BIONAND(ATECH-25-02).
摘要Automatic and accurate medical image segmentation remains a fundamental task in computer-aided diagnosis and treatment planning.Recent advances in foundation models,such as the medical-focused Segment AnythingModel(MedSAM),have demonstrated strong performance but face challenges inmanymedical applications due to anatomical complexity and a limited domain-specific prompt.Thiswork introduces amethodology that enhances segmentation robustness and precision by automatically generating multiple informative point prompts,rather than relying on single inputs.The proposed approach randomly samples sets of spatially distributed point prompts based on image features,enabling MedSAM to better capture fine-grained anatomical structures and boundaries.During inference,probability maps are aggregated to reduce local misclassifications without additional model training.Extensive experiments on various computed tomography(CT)and magnetic resonance imaging(MRI)datasets demonstrate improvements in Dice Similarity Coefficient(DSC)and Normalized Surface Dice(NSD)metrics compared to baseline SAM and Scribble Prompt models.A semi-automatic point sampling version based on the ground truth segmentations yielded enhanced results,achieving up to 92.1%DSC and 86.6%NSD,with significant gains in delineating complex organs such as the pancreas,colon,kidney,and brain tumours.The main novelty of our method consists of effectively combining the results of multiple point prompts into the medical segmentation pipeline so that single-point prompt methods are outperformed.Overall,the proposed model offers a straightforward yet effective approach to improve medical image segmentation performance while maintaining computational efficiency.
基金The Key Project of Ningxia Natural Science Foundation under grant 2025AAC020006the Regional Program of National Natural Science Foundation of China under grant 62561002The Shaanxi University of Technology Foundation Project under grant SLGRCQD2137.
摘要Traditional Mamba-UNet integrations employ four-stage architectures,replacing conventional five-stage UNets with VMamba blocks for global dependency modeling.Unlike Transformers,which suffer from quadratic complexity and high memory consumption in self-attention,Mamba-UNet achieves efficient global modeling through linear-complexity state space modeling.This paper proposes TriLVM-UNet,a lightweight three-stage architecture that integrates parameter-efficient VMamba blocks and enhances cross-stage feature interaction via an improved skip-attention bridge(SAB)module inspired by UltraLight VM-UNet.The model incorporates a Lightweight Vision Mamba(LVM)layer for high-resolution feature extraction,alongside multi-scale dilated convolution(MSDC)and convolutional block attention module(CBAM)for enhanced feature fusion.Evaluated on the 3D ACDC dataset against six baseline models,TriLVM-UNet achieves 98.57%accuracy.The GitHub repository is available at:http://gffzz188fe103f8f1460asqkpuxpv06uqv6xxc.ffgz.tsg.suse.edu.cn/730432ch/TriLVM-UNet.
摘要This study presents a novel hybrid optimization model that combines the complementary aspects of Sand Cat Swarm Optimization(SCSO)and Whale Optimization Algorithm(WOA)to solve the multi-level image thresholding problem.The proposed approach utilizes an adaptive two-stage mechanism that balances the high exploration capacity of SCSO with the concentrated local search capability of WOA,aiming to maximize inter-class variance in the histogram-based thresholding process.Various experiments are conducted on lung cancer,prostate,and mixed medical image datasets.Results demonstrate that modified SCSOWOA achieves superior performance across all datasets.For LC25000,it attains PSNR 27.9453 dB,SSIM 0.9340,FSIM 0.9542,Dice coefficient 0.8901,and Jaccard index 0.8034 at T=12.For prostate,PSNR reaches 28.3965 dB,SSIM 0.7532,FSIM 0.8170,Dice 0.9215,and Jaccard 0.8593.In the MSD dataset,SCSOWOA achieves PSNR 29.3244 dB,SSIM 0.7118,FSIM 0.7562,Dice 0.8901,and Jaccard 0.8034,indicating consistent performance across diverse organs and modalities.The method also demonstrates high computational efficiency,with an average execution time of 1.3221 s,offering up to 40%speed improvement over conventional metaheuristics such as PSO and GWO.Overall,proposed method provides high-accuracy,low-variance,and computationally efficient segmentation,preserving both structural and perceptual fidelity.These results confirm the method’s robustness,generalizability,and practical applicability for AI-assisted diagnostic systems across histopathological and medical imaging contexts,balancing precision,structural preservation,and speed for real-world clinical deployment.
基金Supported by Zhejiang Provincial Natural Science Foundation of China(LY22F020025)the National Natural Science Foundation of China(62072126)。
摘要The Transformer has achieved great success in the field of medical image segmentation,but its quadratic computational complexity limits its application in dense medical image prediction.Recently,the receptance weighted key value(RWKV)architecture has garnered widespread attention due to its linear computational complexity and its capability of parallel computation during training.Despite the RWKV model's proficiency in addressing long-range modeling tasks with linear computational complexity,most current RWKV-based approaches employ static scanning patterns.These patterns may inadvertently incorporate biased prior knowledge into the model's predictions.To address this challenge,we propose a multi-head scan strategy combined with padding methods to effectively simulate spatial continuity in 2D images.Within the Feature Aggregation Attention(FAA)module,asymmetric convolutions are designed to aggregate 1D sequence features along a single dimension,thereby expanding effective receptive fields while preserving structural sparsity.Additionally,panoramic token shift(P-Shift)effectively models local dependency relationships by moving tokens from a wide receptive field.Extensive experiments conducted on the ISIC17/18 and ACDC datasets demonstrate that our method exhibits superior performance in dense medical image prediction tasks.
基金supported by Program for Young Talents of Basic Research in Universities of Heilongjiang Province(YQJH2024036)Collaborative Innovation Projects of“Double First-class”Disciplines in Heilongjiang Province(LJGXCG2024-P20)Outstanding Talent Cultivation Foundation of Northeast Petroleum University(SJQHB202004)。
摘要The development of oil and gas is constrained by difficulties in dynamically characterizing pore structures.Traditional methods inadequately represent the complex interactions between mineral dissolution,precipitation,and fluid flow.This study addresses these gaps by introducing a Transformer U-Neural Network(TransUNet)for computed tomography(CT)image segmentation.The integrated workflow combines conventional CT(Resolution of 5.4μm)and synchrotron radiation CT(Resolution of0.8μm)for dynamic flooding,imaging,segmentation,and precise 3D pore network extraction,overcoming resolution limits.TransUNet's strong global attention and feature extraction reduce overfitting and deliver high-accuracy segmentation of minerals,pores,and argillaceous microporous networks(AMN),achieving 74.92%intersection over union(IoU)for AMN.A porosity correction method improves conventional CT porosity accuracy to 94%of gas-measured values.Alkaline flooding experiments reveal:(1)initial clay swelling reduces small pore size by~50%as alkaline ions destabilize clay;(2)mineral dissolution,such as dolomite,creates secondary pores,increasing 80μm pores by 1.8 times;(3)silicate dissolution increases porosity and leads to a 93.7%rise in permeability.Clay reorganization enhances the AMN by 46.1%.The pore size distribution shifts to log-no rmal at steady state,and throat connectivity improves flow capacity.This work pioneers Transformer-based CT image segmentation,introduces cross-resolution prediction,and clarifies pore regulation by mineral phase changes,establishing a new paradigm for chemical flooding in sandstone reservoirs.
基金supported under the 2024 Foshan City Self-Funded Science and Technology Innovation Project“Research on Image Segmentation Technology Based on Convolutional Neural Networks in Crop Images”(Project Number:2420001004686).
摘要The rising need for precision farming and sustainable land management has catalyzed the requirement for sophisticated means of deriving practical data from remote sensing images.Image segmentation,or the process of dividing the image into semantically relevant parts,has become a groundbreaking technology that allows resolving the problem of transitioning the pixel-level data to a parcel-level analysis.This review is a synthesis of the segmentation methods and their use in crop research and geospatial science.The architectures of pixel-based,object-based,and deep learning(convolutional neural networks,U-Net,Mask R-CNN,and Transformer models)are considered in terms of principles,capabilities,and limitations.Multi-spectral,hyperspectral,LiDAR,and SAR data are integrated to improve the efficiency of segmentation,allowing the possible delineation of fields,the classification of crops,health monitoring,monitoring of yields,and stress identification.In addition to agriculture,segmentation helps in land use and land cover mapping,identification of temporal change,monitoring of the environment,and is used in combination with GIS-based spatial modeling.Nevertheless,issues related to data heterogeneity,mixed pixels,computational requirements,and inadequate availability of labelled data still exist despite the major progress.The future directions involve multi-source data fusion,pixel-to-parcel pipeline automation,and predictive models based on AI,which are used to enhance its scalability,robustness,and the ability to monitor in real-time.This review makes it clear that the use of image segmentation as a tool in generating precision agriculture,sustainable land use,and informed geospatial.
基金supported by the National Natural Science Foundation of China(62276092,62303167)the Postdoctoral Fellowship Program(Grade C)of China Postdoctoral Science Foundation(GZC20230707)+3 种基金the Key Science and Technology Program of Henan Province,China(242102211051)MRC(MC_PC_17171)Royal Society(RP202G0230)BHF(AA/18/3/34220)。
摘要In the field of medical image processing,combining global and local relationship modeling constitutes an effective strategy for precise segmentation.Prior research has established the validity of Convolutional Neural Networks(CNN)in modeling local relationships.Conversely,Transformers have demonstrated their capability to effectively capture global contextual information.However,when utilized to address CNNs’limitations in modeling global relationships,Transformers are hindered by substantial computational complexity.To address this issue,we introduce Mamba,a State-Space Model(SSM)that exhibits exceptional proficiency in modeling long-range dependencies in sequential data.Given Mamba’s demonstrated potential in 2D medical image segmentation in previous studies,we have designed a Dual-encoder Global-local Feature Extraction Network based on Mamba,termed DGFE-Mamba,to accurately capture and fuse long-range dependencies and local dependencies within multi-scale features.Compared to Transformer-based methods,the DGFE-Mamba model excels in comprehensive feature modeling and demonstrates significantly improved segmentation accuracy.To validate the effectiveness and practicality of DGFE-Mamba,we conducted tests on the Automatic Cardiac Diagnosis Challenge(ACDC)dataset,the Synapse multi-organ CT abdominal segmentation dataset,and the Colorectal Cancer Clinic(CVC-ClinicDB)dataset.The results showed that DGFE-Mamba achieved Dice coefficients of 92.20,83.67,and 94.13,respectively.These findings comprehensively validate the effectiveness and practicality of the proposed DGFE-Mamba architecture.
基金supported by the Natural Science Foundation of the Anhui Higher Education Institutions of China(Grant Nos.2023AH040149 and 2024AH051915)the Anhui Provincial Natural Science Foundation(Grant No.2208085MF168)+1 种基金the Science and Technology Innovation Tackle Plan Project of Maanshan(Grant No.2024RGZN001)the Scientific Research Fund Project of Anhui Medical University(Grant No.2023xkj122).
摘要Convolutional neural networks(CNNs)-based medical image segmentation technologies have been widely used in medical image segmentation because of their strong representation and generalization abilities.However,due to the inability to effectively capture global information from images,CNNs can easily lead to loss of contours and textures in segmentation results.Notice that the transformer model can effectively capture the properties of long-range dependencies in the image,and furthermore,combining the CNN and the transformer can effectively extract local details and global contextual features of the image.Motivated by this,we propose a multi-branch and multi-scale attention network(M2ANet)for medical image segmentation,whose architecture consists of three components.Specifically,in the first component,we construct an adaptive multi-branch patch module for parallel extraction of image features to reduce information loss caused by downsampling.In the second component,we apply residual block to the well-known convolutional block attention module to enhance the network’s ability to recognize important features of images and alleviate the phenomenon of gradient vanishing.In the third component,we design a multi-scale feature fusion module,in which we adopt adaptive average pooling and position encoding to enhance contextual features,and then multi-head attention is introduced to further enrich feature representation.Finally,we validate the effectiveness and feasibility of the proposed M2ANet method through comparative experiments on four benchmark medical image segmentation datasets,particularly in the context of preserving contours and textures.
基金the Innovation Program of Shanghai Industrial Synergy(No.XTCX-KJ-2023-2-12)。
摘要Medical image segmentation is a crucial task in clinical applications.However,obtaining labeled data for medical images is often challenging.This has led to the appeal of semi-supervised learning(SSL),a technique adept at leveraging a modest amount of labeled data.Nonetheless,most prevailing SSL segmentation methods for medical images either rely on the single consistency training method or directly fine-tune SSL methods designed for natural images.In this paper,we propose an innovative semi-supervised method called multi-consistency training(MCT)for medical image segmentation.Our approach transcends the constraints of prior methodologies by considering consistency from a dual perspective:output consistency across different up-sampling methods and output consistency of the same data within the same network under various perturbations to the intermediate features.We design distinct semi-supervised loss regression methods for these two types of consistencies.To enhance the application of our MCT model,we also develop a dedicated decoder as the core of our neural network.Thorough experiments were conducted on the polyp dataset and the dental dataset,rigorously compared against other SSL methods.Experimental results demonstrate the superiority of our approach,achieving higher segmentation accuracy.Moreover,comprehensive ablation studies and insightful discussion substantiate the efficacy of our approach in navigating the intricacies of medical image segmentation.
基金supported by the Natural Science Foundation of China(No.41804112,author:Chengyun Song).
摘要Existing semi-supervisedmedical image segmentation algorithms use copy-paste data augmentation to correct the labeled-unlabeled data distribution mismatch.However,current copy-paste methods have three limitations:(1)training the model solely with copy-paste mixed pictures from labeled and unlabeled input loses a lot of labeled information;(2)low-quality pseudo-labels can cause confirmation bias in pseudo-supervised learning on unlabeled data;(3)the segmentation performance in low-contrast and local regions is less than optimal.We design a Stochastic Augmentation-Based Dual-Teaching Auxiliary Training Strategy(SADT),which enhances feature diversity and learns high-quality features to overcome these problems.To be more precise,SADT trains the Student Network by using pseudo-label-based training from Teacher Network 1 and supervised learning with labeled data,which prevents the loss of rare labeled data.We introduce a bi-directional copy-pastemask with progressive high-entropy filtering to reduce data distribution disparities and mitigate confirmation bias in pseudo-supervision.For the mixed images,Deep-Shallow Spatial Contrastive Learning(DSSCL)is proposed in the feature spaces of Teacher Network 2 and the Student Network to improve the segmentation capabilities in low-contrast and local areas.In this procedure,the features retrieved by the Student Network are subjected to a random feature perturbation technique.On two openly available datasets,extensive trials show that our proposed SADT performs much better than the state-ofthe-art semi-supervised medical segmentation techniques.Using only 10%of the labeled data for training,SADT was able to acquire a Dice score of 90.10%on the ACDC(Automatic Cardiac Diagnosis Challenge)dataset.