Although Transformer-based image restoration methods have demonstrated impressive performance,existing Transformers still insufficiently exploit multiscale information.Previous non-Transformer-based studies have shown...Although Transformer-based image restoration methods have demonstrated impressive performance,existing Transformers still insufficiently exploit multiscale information.Previous non-Transformer-based studies have shown that incorporating multiscale features is crucial for improving restoration results.In this paper,we propose a multiscale Transformer(MST)that captures cross-scale attention among tokens,thereby effectively leveraging the multiscale patch recurrence prior of natural images.Furthermore,we introduce a channel-gate feed-forward network(CGFN)to enhance inter-channel information aggregation and reduce channel redundancy.To simultaneously utilise global,local and multiscale features,we design a multitype feature integration block(MFIB).Extensive experiments on both image super-resolution and HEVC compressed video artefact reduction demonstrate that the proposed MST achieves state-of-the-art performance.Ablation studies further verify the effectiveness of each proposed module.展开更多
After more than forty years of development,the accuracy of digital image correlation(DIC)methods has reached an extremely high level.However,the interpolation bias of DIC has not been resolved.With the flourishing of ...After more than forty years of development,the accuracy of digital image correlation(DIC)methods has reached an extremely high level.However,the interpolation bias of DIC has not been resolved.With the flourishing of deep learning in the field of image superresolution,it has become possible to use deep learning-based image superresolution methods to reduce DIC interpolation bias.To achieve this goal,this paper improves the local implicit image function(LIIF)method based on the characteristics of speckle images to obtain local implicit image function suitable for speckle images(LIIF-S),achieving continuous image representation and arbitrary resolution interpolation.Subsequently,LIIF-S is used as the interpolation algorithm of the inverse compositional-Gaussian Newton method to reduce the interpolation bias.The simulation experiment results show that LIIF-S not only improves the accuracy by more than one order of magnitude compared to traditional interpolation algorithms but also that the interpolation bias does not have sinusoidal characteristics.In addition,the effectiveness and generalization of the LIIF-S method in unseen real-world scenarios have also been demonstrated through physical experiments.The code and dataset are publicly available at http://gffzz188fe103f8f1460asnufu9uo0xc0k69bk.ffgz.tsg.suse.edu.cn/LianpoWang/SLIIF.展开更多
In order to thoroughly investigate the impact of lasers on the image quality of charge-coupled device(CCD)image sensors,this paper employs objective image quality assessment methods to study the damage effects of 1064...In order to thoroughly investigate the impact of lasers on the image quality of charge-coupled device(CCD)image sensors,this paper employs objective image quality assessment methods to study the damage effects of 1064 nm continuous lasers on CCD image sensors at different power densities.Through a comprehensive analysis of grayscale distribution histograms,peak signal-to-noise ratio(PSNR)and structural similarity index measure(SSIM),the variations in image quality with increasing power density are revealed.The research results indicate that at low power density levels,the image quality decreases slowly;however,once the power density reaches a threshold,the image quality rapidly declines.When the image quality is severely damaged,the frequency of the main grayscale value drops to 0,and the degree of image damage saturates.Additionally,with the increase in power density,the dynamic range of the image first expands and then shrinks,eventually shifting towards the low grayscale value range.展开更多
For low-light image enhancement tasks,RAW images surpass RGB images due to their high information content,however,their noise and single-channel nature challenge feature extraction.Existing methods using multi-stage c...For low-light image enhancement tasks,RAW images surpass RGB images due to their high information content,however,their noise and single-channel nature challenge feature extraction.Existing methods using multi-stage convolutional neural network(CNN)frameworks struggle with global feature extraction,while single-stage CNN-transformer fusions often result in residual noise.To overcome these limitations,this paper introduces a multi-stage RAW image enhancement network combining CNN and transformer.Considering the characteristics inherent to the task,we devised a CNN-based denoising block for the denoising stage and incorporated wavelet information to enhance frequency features.A transformer-based correction block has been designed for the color and white balance recovery stage,with the white balance being adjusted dynamically using a signal-to-noise ratio(SNR)map.With this design,our method outperforms other state-of-the-art models in all metrics on the Sony and Fuji datasets of see-in-the-dark(SID),and achieves optimal structural similarity index measurement(SSIM)on the mono-colored raw(MCR)dataset.展开更多
To tackle the issue where current techniques for combining infrared and visible imagery often fail to equilibrate thermal saliency and background textures in complicated heterogeneous scenarios(especially under few-sh...To tackle the issue where current techniques for combining infrared and visible imagery often fail to equilibrate thermal saliency and background textures in complicated heterogeneous scenarios(especially under few-shot scenarios),we present a subregional adaptive framework inspired by the rattlesnake visual mechanism.The method first employs an enhancement model inspired by the rattlesnake visual mechanism to improve the contrast and details of infrared images.Subsequently,a lightweight object detection network is integrated to precisely partition the image into target and non-target regions.Differentiated fusion strategies are designed for these regions:focusing on preserving infrared thermal radiation information in target regions,while maintaining visible texture structures in background regions,with a Gaussian blur mechanism ensuring smooth transitions.Finally,a colorization step is employed to enhance visual perception by restoring chromatic information to the fused image.Extensive evaluations on widely used benchmark datasets demonstrate that our framework consistently outperforms state-of-the-art bio-inspired and deep learning methods.Qualitatively,it significantly excels in enhancing the thermal contrast of targets and preserving high-fidelity background textures.Quantitatively,it delivers notable improvements in key metrics such as Edge Preservation(QAB/F),Average Gradient,and Cross Entropy,reflecting superior edge information transfer and structural fidelity.Most importantly,the proposed framework achieves these superior results specifically under few-shot settings,demonstrating exceptional data efficiency and robust generalization.展开更多
In hyperspectral remote sensing imagery,pixel interactions within defined spatial extents result in the mixing of adjacent pixels.Additionally,the high similarity of adjacent spectra leads to information redundancy,wh...In hyperspectral remote sensing imagery,pixel interactions within defined spatial extents result in the mixing of adjacent pixels.Additionally,the high similarity of adjacent spectra leads to information redundancy,which hinders the extraction of global spatial and spectral correlations.In order to solve the problems of mixed adjacent pixels and redundant adjacent spectra,this work offers a hyperspectral image classification approach that uses a global space-spectral attention mechanism.First,the proposed method's global spatial attention module uses multi-scale dilated convolution to produce a bigger receptive field to be capable of capturing global spatial correlation and obtain unmixed pixel information.Then,the global spectral attention module designs a spectral domain partition algorithm,using the combination of regional density as well as information entropy as the threshold to divide spectrum into dispersed subsets and eliminate redundant information.The global context information for entire spectral band is fully exploited,and correlation of the global spectral information is extracted.Finally,the two modules combine to provide a global correlation of space and spectrum.Experiments demonstrate that the suggested method obtains overall accuracies of 97.28%,94.73%,and 95.76%on the three WHU-Hi hyperspectral datasets,surpassing comparison methods.展开更多
This study presents a novel hybrid optimization model that combines the complementary aspects of Sand Cat Swarm Optimization(SCSO)and Whale Optimization Algorithm(WOA)to solve the multi-level image thresholding proble...This study presents a novel hybrid optimization model that combines the complementary aspects of Sand Cat Swarm Optimization(SCSO)and Whale Optimization Algorithm(WOA)to solve the multi-level image thresholding problem.The proposed approach utilizes an adaptive two-stage mechanism that balances the high exploration capacity of SCSO with the concentrated local search capability of WOA,aiming to maximize inter-class variance in the histogram-based thresholding process.Various experiments are conducted on lung cancer,prostate,and mixed medical image datasets.Results demonstrate that modified SCSOWOA achieves superior performance across all datasets.For LC25000,it attains PSNR 27.9453 dB,SSIM 0.9340,FSIM 0.9542,Dice coefficient 0.8901,and Jaccard index 0.8034 at T=12.For prostate,PSNR reaches 28.3965 dB,SSIM 0.7532,FSIM 0.8170,Dice 0.9215,and Jaccard 0.8593.In the MSD dataset,SCSOWOA achieves PSNR 29.3244 dB,SSIM 0.7118,FSIM 0.7562,Dice 0.8901,and Jaccard 0.8034,indicating consistent performance across diverse organs and modalities.The method also demonstrates high computational efficiency,with an average execution time of 1.3221 s,offering up to 40%speed improvement over conventional metaheuristics such as PSO and GWO.Overall,proposed method provides high-accuracy,low-variance,and computationally efficient segmentation,preserving both structural and perceptual fidelity.These results confirm the method’s robustness,generalizability,and practical applicability for AI-assisted diagnostic systems across histopathological and medical imaging contexts,balancing precision,structural preservation,and speed for real-world clinical deployment.展开更多
Natural fractures serve as the primary storage spaces and flow pathways in deep to ultra-deep tight sandstone reservoirs,directly influencing hydrocarbon accumulation,preservation,and production.Borehole images offer ...Natural fractures serve as the primary storage spaces and flow pathways in deep to ultra-deep tight sandstone reservoirs,directly influencing hydrocarbon accumulation,preservation,and production.Borehole images offer intuitive,continuous,and high-resolution identification of natural fractures along the entire borehole.However,relying solely on complete sinusoidal curves from borehole images for fracture identification may lead to omissions,as it overlooks cases where these curves are incomplete or truncated.To address the problems and deficiencies in fracture identification,this study systematically classifies borehole image feature patterns based on core-to-log spatial position restoring.A bidirectional comparison is conducted between natu ral fractures in cores and the fracture image features in borehole images.A quantitative relationship between fracture dip angle,thin layer thickness and borehole radius was established,accompanied by a mathematical expression describing the fracture curve morphology was proposed.These findings enabled the development of an imaging response pattern for natural fractures in deep and ultra-deep tight sandstone reservoirs,incorporating key parameters such as dip angle,through-layer connectivity,and spatial position within the borehole.In the Bashijiqike-Baxigai tight-sandstone reservoirs of the Bozi-Dabei area,we estimate that approximately 24%of coreobserved fractures display distinct linear-pattern features on borehole images,whereas approximately 91%of borehole images features can be correlated with fractures observed in core.Fracture identification rates for natural fractures increased by 17%in water-based mud and by 3%in oil-based mud through the application of the natural fracture image response pattern.Moreover,this study analyzes the deviations in the matching between core fractures and image features.Finally,we further discuss the common sources of error in natural fracture identification using borehole images from multiple perspectives,including missing core responses,inconsistencies between core and borehole image features,distortion of fracture chord curve,inaccurate fracture count,misclassification of fractures,and variations in interpretation under different mud systems.The research addresses the blind spots of traditional methods in fracture identification within thin layers,not only enhancing the detection rate of natural fractu res but also further improving the accuracy of fractu re recognitio n.At the same time,it will contribute to the optimization of fracture characterization,reservoir evaluation,and production forecasting,providing a more reliable data foundation for exploration and development under complex geological conditions.展开更多
The stability of concrete-rock interfaces is a critical issue in underground engineering.This study investigated the strain localization mechanism and energy evolution of concrete-sandstone specimens containing single...The stability of concrete-rock interfaces is a critical issue in underground engineering.This study investigated the strain localization mechanism and energy evolution of concrete-sandstone specimens containing single and double interfacial cracks at various inclination angles.Acoustic emission(AE)technology and energy theory were used to analyze energy evolution,whereas digital image correlation was employed to examine strain development and fracture mechanisms.A new approach combining digital image processing and custom binarization was introduced to characterize the fractal properties of crack patterns using the box-counting method.Experimental results showed that the AE cumulative energy and stress-strain curves divided the loading process into three stages:crack closure(Ⅰ),stable crack growth(Ⅱ),and rapid crack propagation(Ⅲ).Fractal dimensions were computed for both singleand double-crack specimens using the Otsu method and the proposed binarization technique.The Otsu method yielded values of 1.457,1.482,1.131,1.512,1.489,1.536,1.171,and 1.491,whereas the new method produced higher values—1.6038,1.6643,1.2713,1.6806,1.5594,1.6282,1.2239,and 1.6565—indicating enhanced fractal characteristics.Furthermore,the proposed method detected a three-phase evolution in fractal dimension before failure,which follows an initial increase,a stable period,and a finalrapid rise.These findingsprovide theoretical support for the application of the proposed method in underground engineering.展开更多
This paper firstly analyzes the characteristics of medical images,and then proposes a specialized compressive sensing algorithm called sparsity-precise iterative hard thresholding(SIHT),which is specifically designed ...This paper firstly analyzes the characteristics of medical images,and then proposes a specialized compressive sensing algorithm called sparsity-precise iterative hard thresholding(SIHT),which is specifically designed to address their specific features such as low sparsity and low frequency.SIHT adaptively measures sparsity and step length which becomes more precise during the iteration process to achieve a certain quality improvement in medical image reconstruction.Experimental results demonstrate that as compared to other image compressive sensing(ICS)reconstruction algorithms across three different types of medical image datasets,SIHT can achieve the best subjective recovery quality particularly in terms of mitigating blocky artifacts and noise,where a notable improvement is obtained in terms of peak signal-to-noise ratio(PSNR)and structural similarity index measurement(SSIM)of medical ICS reconstruction.展开更多
In the image fusion field,fusing infrared images(IRIs)and visible images(VIs)excelled is a key area.The differences between IRIs and VIs make it challenging to fuse both types into a high-quality image.Accordingly,eff...In the image fusion field,fusing infrared images(IRIs)and visible images(VIs)excelled is a key area.The differences between IRIs and VIs make it challenging to fuse both types into a high-quality image.Accordingly,efficiently combining the advantages of both images while overcoming their shortcomings is necessary.To handle this challenge,we developed an end-to-end IRI andVI fusionmethod based on frequency decomposition and enhancement.By applying concepts from frequency domain analysis,we used the layering mechanism to better capture the salient thermal targets from the IRIs and the rich textural information from the VIs,respectively,significantly boosting the image fusion quality and effectiveness.In addition,the backbone network combined Restormer Blocks and Dense Blocks;Restormer blocks utilize global attention to extract shallow features.Meanwhile,Dense Blocks ensure the integration between shallow and deep features,thereby avoiding the loss of shallow attributes.Extensive experiments on TNO and MSRS datasets demonstrated that the suggested method achieved state-of-the-art(SOTA)performance in various metrics:Entropy(EN),Mutual Information(MI),Standard Deviation(SD),The Structural Similarity Index Measure(SSIM),Fusion quality(Qabf),MI of the pixel(FMIpixel),and modified Visual Information Fidelity(VIFm).展开更多
Industrial anomaly detection is dedicated to identifying and locating regions that deviate from the standard appearance.The prevailing approach achieves unsupervised anomaly detection through the reconstruction of ima...Industrial anomaly detection is dedicated to identifying and locating regions that deviate from the standard appearance.The prevailing approach achieves unsupervised anomaly detection through the reconstruction of images using autoencoders.Due to the simplistic structure of some abnormal regions,the autoencoder can effectively reconstruct these areas,consequently diminishing the model’s anomaly detection capabilities.To address this issue,this paper transforms the reconstruction task into the inpainting-filling-reconstruction task to increase the reconstruction error between abnormal samples and normal samples.The masked regions inpainted by the filling network are used to fill in the input image,thereby achieving an effect similar to masking.Unlike typical masking processes,this approach retains partial authentic information in the input image,rendering it partially visible.This is beneficial for the reconstruction network to repair the masked areas.Due to the consistent structure between the masked region inpainted by the filling network and the normal region,the filled abnormal regions display a complex structure that has not been learned,making it difficult for the reconstruction network to reconstruct the abnormal regions.Experimental results indicate that our method performs better than other methods on both the MVTec AD dataset and the MVTec LOCO AD dataset.展开更多
Dear Editor,Due to the scarcity of high-quality infrared data,translating visible images to infrared has become a practical solution to meet the growing demand for infrared images in low-light and adverse conditions.D...Dear Editor,Due to the scarcity of high-quality infrared data,translating visible images to infrared has become a practical solution to meet the growing demand for infrared images in low-light and adverse conditions.Due to the large modality gap and limited prior information,existing visible-to-infrared(VIS-to-IR)image translation methods often struggle with poor structural preservation,unclear cross-modal correspondence,and loss of thermal details.展开更多
To address the issue of inconsistent image quality and data scarcity in bolt defect detection for transmission lines,this paper proposes an improved sparse region-based convolutional neural network(RCNN) based detecti...To address the issue of inconsistent image quality and data scarcity in bolt defect detection for transmission lines,this paper proposes an improved sparse region-based convolutional neural network(RCNN) based detection framework integrating image quality evaluation and text-to-image data augmentation.First,a HyperNetwork-based image quality assessment module is introduced to filter low-quality inspection images in terms of clarity and structural integrity,resulting in a high-quality training dataset.Second,a text-to-image diffusion model is utilized for sample augmentation.By designing text prompts that describe various bolt defect types under diverse lighting and viewing conditions,the model automatically generates realistic synthetic samples.The generated images are further filtered using a combination of quality and perceptual similarity metrics to ensure consistency with the real data distribution.Building upon the sparse RCNN baseline,a dynamic label assignment mechanism and a random decision path detection head are incorporated to enhance bounding box matching and prediction accuracy.Experimental results demonstrate that the proposed method significantly improves detection accuracy(mAP@0.5) over the original sparse RCNN while maintaining low computational cost,enabling more efficient and intelligent inspection of transmission line components.展开更多
The current infrared image pedestrian detectors have problems with high rates of false positives and false negatives. To solve these problems, we proposed an improved anchor-free fully convolutional one-stage object d...The current infrared image pedestrian detectors have problems with high rates of false positives and false negatives. To solve these problems, we proposed an improved anchor-free fully convolutional one-stage object detection(FCOS) algorithm. Firstly, we introduced the channel attention module squeeze excitation(SE)-Block in the FCOS backbone network, which was used to learn how to model the relative importance between different feature channels, and to achieve the weight recalibration of the features extracted from the convolution neural network, and improve the weight values that are more important for pedestrian target detection. Secondly, soft non-maximum suppression(Soft-NMS) replaced the conventional NMS within the algorithm's post-processing phase, which was used to reduce the probability of missed detection for occluded pedestrians. The experimental results show that our improved FCOS algorithm improves the average precision(AP) by 6.71% on the original dataset and 7.97% on the augmented KAIST pedestrian dataset compared with the original FCOS algorithm. Our improvements effectively meet the real-time requirements and there is no significant decrease in speed compared with the original FCOS algorithm, and decreased the false positives and false negatives for infrared image pedestrian detection.展开更多
Distributive Fluvial Systems(DFS)are critical sedimentary systems governing fluvial dynamics,sediment transport,and ecosystem sustainability in modern and ancient basins.Accurate quantification of DFS channel morpholo...Distributive Fluvial Systems(DFS)are critical sedimentary systems governing fluvial dynamics,sediment transport,and ecosystem sustainability in modern and ancient basins.Accurate quantification of DFS channel morphology is essential for advancing sedimentary modeling,optimizing water resource management,and mitigating fluvial hazards.Here,the authors present a novel automated framework that extracts DFS channel networks from remote sensing imagery by integrating multiscale image segmentation,fractal network evolution,and region-merging algorithms.Through hierarchically multiresolution feature processing,this method overcomes limitations of traditional single-scale analysis,enabling adaptive extraction while reducing segmentation heterogeneity.Specifically,the workflow consists of three stages:Image segmentation,feature extraction,and image classification.When applied to the Golmud fluvial fan(Qinghai,China),this approach achieves 90.2%overall channel extraction accuracy using 0.5 m resolution imagery,significantly outperforming traditional DEM-based(81.7%)and water spectral methods(85.4%)in resolving fine-scale channel networks.Crucially,the framework demonstrates robust adaptability to complex sedimentary environments with variable vegetation cover(<30%density)and spectral noise,providing a time-efficient,data-agnostic solution for DFS characterization.展开更多
Lossy image coding is the art of computing that is principally bounded by the image’s rate-distortion function.This bound,though never accurately characterized,has been approached practically via deep learning techno...Lossy image coding is the art of computing that is principally bounded by the image’s rate-distortion function.This bound,though never accurately characterized,has been approached practically via deep learning technologies in recent years.Indeed,learned image coding schemes allow direct optimization of the joint rate-distortion cost,thereby outperforming the handcrafted image coding schemes by a large margin.Still,it is observed that there is room for further improvement in the rate-distortion performance of learned image coding.In this article,we identify the gap between the ideal rate-distortion function forecasted by Shannon’s information theory and the empirical rate-distortion function achieved by the state-of-the-art learned image coding schemes,revealing that the gap is incurred by five different effects:modeling effect,approximation effect,amortization effect,digitization effect,and asymptotic effect.We design simulations and experiments to quantitatively evaluate the last three effects,which demonstrates the high potential of future lossy image coding technologies.展开更多
Organoids possess immense potential for unraveling the intricate functions of human tissues and facilitating preclinical disease treatment.Their applications span from high-throughput drug screening to the modeling of...Organoids possess immense potential for unraveling the intricate functions of human tissues and facilitating preclinical disease treatment.Their applications span from high-throughput drug screening to the modeling of complex diseases,with some even achieving clinical translation.Changes in the overall size,shape,boundary,and other morphological features of organoids provide a noninvasive method for assessing organoid drug sensitivity.However,the precise segmentation of organoids in bright-field microscopy images is made difficult by the complexity of the organoid morphology and interference,including overlapping organoids,bubbles,dust particles,and cell fragments.This paper introduces the precision organoid segmentation technique(POST),which is a deep-learning algorithm for segmenting challenging organoids under simple bright-field imaging conditions.Unlike existing methods,POST accurately segments each organoid and eliminates various artifacts encountered during organoid culturing and imaging.Furthermore,it is sensitive to and aligns with measurements of organoid activity in drug sensitivity experiments.POST is expected to be a valuable tool for drug screening using organoids owing to its capability of automatically and rapidly eliminating interfering substances and thereby streamlining the organoid analysis and drug screening process.展开更多
Motion artifacts and noise are key factors that determine image quality in optical coherence tomography angiography(OCTA).Although deep learning has emerged as an effective method for artifacts removal and denoising,i...Motion artifacts and noise are key factors that determine image quality in optical coherence tomography angiography(OCTA).Although deep learning has emerged as an effective method for artifacts removal and denoising,its generalization capability remains limited,and it is difficult to handle images with both motion artifacts and noise.To address this issue,we designed a Swin Transformer-based multi-scale motion artifacts and noise parallel removal network(ST-MANPR)to learn the nonlinear mapping between images with motion artifacts and noise and images without them,thereby achieving simultaneous suppression of motion artifacts and noise.The proposed network integrates the Swin window attention module,channel attention module(CAM),and adaptive multi-scale convolution denoising module(AMS-CNN)to enhance its capability in processing complex image features.At the same time,a hybrid loss function combining wavelet transform(WT)and mean square error(MSE)was introduced to facilitate high-frequency detail restoration.In addition,we constructed a dataset of OCTA images with and without motion artifacts and noise at multiple intensity levels.The created dataset was applied to train the network and the test results were evaluated both visually and numerically.The experimental results show that the proposed network can effectively remove motion artifacts and noise in OCTA images simultaneously.展开更多
Accurate quantification of crop residue cover(CRC)is crucial for monitoring and evaluating conservation tillage practices,yet it poses a significant image segmentation challenge.The subtle visual distinctions between ...Accurate quantification of crop residue cover(CRC)is crucial for monitoring and evaluating conservation tillage practices,yet it poses a significant image segmentation challenge.The subtle visual distinctions between fragmented residue and soil,compounded by variable illumination and shadows in field imagery,often lead to poor segmentation performance.To overcome these limitations,we introduce RCTUnet,a novel deep learning architecture designed for robust crop-residue-soil segmentation and precise CRC estimation.RCTUnet’s architecture synergistically integrates three key components:(1)a ResNet50 backbone for deep,multi-scale feature extraction;(2)a convolutional block attention module(CBAM)to adaptively focus on salient residue features across both channel and spatial dimensions;and(3)a transformer-based global context fusion module(GCFM)to model long-range spatial dependencies,which is critical for interpreting heterogeneous residue patterns.We evaluated RCTUnet on a dataset of 1220 field-acquired images spanning four typical crop rotations.Experimental results show that,compared to traditional models:(1)RCTUnet achieves significantly higher crop-residue-soil segmentation accuracy than classic models including Unet,Unet++,DeepLabV3,segmentation network(SegNet),and fully convolutional network(FCN),with improvements of 3.24%,3.42%,4.88%,8.28%,and 6.05%in overall accuracy,respectively;(2)RCTUnet yields superior residue-soil segmentation performance,with increases in residue recall of 7.67%,7.37%,14.09%,27.05%,and 16.91%,respectively;(3)RCTUnet shows enhanced CRC estimation accuracy,achieving a root mean square error(RMSE)of 4.875,representing a 45.5%improvement over Unet(RMSE=8.941).These results demonstrate the efficacy of our hybrid approach,which combines deep hierarchical features,dual-domain attention,and global context modeling.RCTUnet provides a robust and reliable tool for automated CRC assessment,advancing the capabilities of in-field agricultural monitoring.展开更多
基金supported in part by the National Natural Science Foundation of China under Grants 62101346 and 62301330the Guangdong Basic and Applied Basic Research Foundation under Grants 2021A1515011702 and 2022A1515110101+1 种基金the Shenzhen Science and Technology Programme under Grants JCYJ20240813141358076 and 20231121103807001the Guangdong Provincial Key Laboratory under Grant 2023B1212060076.
摘要Although Transformer-based image restoration methods have demonstrated impressive performance,existing Transformers still insufficiently exploit multiscale information.Previous non-Transformer-based studies have shown that incorporating multiscale features is crucial for improving restoration results.In this paper,we propose a multiscale Transformer(MST)that captures cross-scale attention among tokens,thereby effectively leveraging the multiscale patch recurrence prior of natural images.Furthermore,we introduce a channel-gate feed-forward network(CGFN)to enhance inter-channel information aggregation and reduce channel redundancy.To simultaneously utilise global,local and multiscale features,we design a multitype feature integration block(MFIB).Extensive experiments on both image super-resolution and HEVC compressed video artefact reduction demonstrate that the proposed MST achieves state-of-the-art performance.Ablation studies further verify the effectiveness of each proposed module.
基金supported by the National Natural Science Foundation of China(Grant No.12302245)the Basic Research Programs of Taicang 2024(Grant No.TC2024JC37)Young Talent Fund of Xi’an Association for Science and Technology(Grant No.0959202513091).
摘要After more than forty years of development,the accuracy of digital image correlation(DIC)methods has reached an extremely high level.However,the interpolation bias of DIC has not been resolved.With the flourishing of deep learning in the field of image superresolution,it has become possible to use deep learning-based image superresolution methods to reduce DIC interpolation bias.To achieve this goal,this paper improves the local implicit image function(LIIF)method based on the characteristics of speckle images to obtain local implicit image function suitable for speckle images(LIIF-S),achieving continuous image representation and arbitrary resolution interpolation.Subsequently,LIIF-S is used as the interpolation algorithm of the inverse compositional-Gaussian Newton method to reduce the interpolation bias.The simulation experiment results show that LIIF-S not only improves the accuracy by more than one order of magnitude compared to traditional interpolation algorithms but also that the interpolation bias does not have sinusoidal characteristics.In addition,the effectiveness and generalization of the LIIF-S method in unseen real-world scenarios have also been demonstrated through physical experiments.The code and dataset are publicly available at http://gffzz188fe103f8f1460asnufu9uo0xc0k69bk.ffgz.tsg.suse.edu.cn/LianpoWang/SLIIF.
摘要In order to thoroughly investigate the impact of lasers on the image quality of charge-coupled device(CCD)image sensors,this paper employs objective image quality assessment methods to study the damage effects of 1064 nm continuous lasers on CCD image sensors at different power densities.Through a comprehensive analysis of grayscale distribution histograms,peak signal-to-noise ratio(PSNR)and structural similarity index measure(SSIM),the variations in image quality with increasing power density are revealed.The research results indicate that at low power density levels,the image quality decreases slowly;however,once the power density reaches a threshold,the image quality rapidly declines.When the image quality is severely damaged,the frequency of the main grayscale value drops to 0,and the degree of image damage saturates.Additionally,with the increase in power density,the dynamic range of the image first expands and then shrinks,eventually shifting towards the low grayscale value range.
摘要For low-light image enhancement tasks,RAW images surpass RGB images due to their high information content,however,their noise and single-channel nature challenge feature extraction.Existing methods using multi-stage convolutional neural network(CNN)frameworks struggle with global feature extraction,while single-stage CNN-transformer fusions often result in residual noise.To overcome these limitations,this paper introduces a multi-stage RAW image enhancement network combining CNN and transformer.Considering the characteristics inherent to the task,we devised a CNN-based denoising block for the denoising stage and incorporated wavelet information to enhance frequency features.A transformer-based correction block has been designed for the color and white balance recovery stage,with the white balance being adjusted dynamically using a signal-to-noise ratio(SNR)map.With this design,our method outperforms other state-of-the-art models in all metrics on the Sony and Fuji datasets of see-in-the-dark(SID),and achieves optimal structural similarity index measurement(SSIM)on the mono-colored raw(MCR)dataset.
基金supported by the Key Research and Development Program of Jilin Province(Grant No.20230201043GX)the Youth Program of the National Natural Science Foundation of China(Grant No.61201368).
摘要To tackle the issue where current techniques for combining infrared and visible imagery often fail to equilibrate thermal saliency and background textures in complicated heterogeneous scenarios(especially under few-shot scenarios),we present a subregional adaptive framework inspired by the rattlesnake visual mechanism.The method first employs an enhancement model inspired by the rattlesnake visual mechanism to improve the contrast and details of infrared images.Subsequently,a lightweight object detection network is integrated to precisely partition the image into target and non-target regions.Differentiated fusion strategies are designed for these regions:focusing on preserving infrared thermal radiation information in target regions,while maintaining visible texture structures in background regions,with a Gaussian blur mechanism ensuring smooth transitions.Finally,a colorization step is employed to enhance visual perception by restoring chromatic information to the fused image.Extensive evaluations on widely used benchmark datasets demonstrate that our framework consistently outperforms state-of-the-art bio-inspired and deep learning methods.Qualitatively,it significantly excels in enhancing the thermal contrast of targets and preserving high-fidelity background textures.Quantitatively,it delivers notable improvements in key metrics such as Edge Preservation(QAB/F),Average Gradient,and Cross Entropy,reflecting superior edge information transfer and structural fidelity.Most importantly,the proposed framework achieves these superior results specifically under few-shot settings,demonstrating exceptional data efficiency and robust generalization.
基金the National Key Research and Devel-opment Program of China(No.2021YFB3900601)。
摘要In hyperspectral remote sensing imagery,pixel interactions within defined spatial extents result in the mixing of adjacent pixels.Additionally,the high similarity of adjacent spectra leads to information redundancy,which hinders the extraction of global spatial and spectral correlations.In order to solve the problems of mixed adjacent pixels and redundant adjacent spectra,this work offers a hyperspectral image classification approach that uses a global space-spectral attention mechanism.First,the proposed method's global spatial attention module uses multi-scale dilated convolution to produce a bigger receptive field to be capable of capturing global spatial correlation and obtain unmixed pixel information.Then,the global spectral attention module designs a spectral domain partition algorithm,using the combination of regional density as well as information entropy as the threshold to divide spectrum into dispersed subsets and eliminate redundant information.The global context information for entire spectral band is fully exploited,and correlation of the global spectral information is extracted.Finally,the two modules combine to provide a global correlation of space and spectrum.Experiments demonstrate that the suggested method obtains overall accuracies of 97.28%,94.73%,and 95.76%on the three WHU-Hi hyperspectral datasets,surpassing comparison methods.
摘要This study presents a novel hybrid optimization model that combines the complementary aspects of Sand Cat Swarm Optimization(SCSO)and Whale Optimization Algorithm(WOA)to solve the multi-level image thresholding problem.The proposed approach utilizes an adaptive two-stage mechanism that balances the high exploration capacity of SCSO with the concentrated local search capability of WOA,aiming to maximize inter-class variance in the histogram-based thresholding process.Various experiments are conducted on lung cancer,prostate,and mixed medical image datasets.Results demonstrate that modified SCSOWOA achieves superior performance across all datasets.For LC25000,it attains PSNR 27.9453 dB,SSIM 0.9340,FSIM 0.9542,Dice coefficient 0.8901,and Jaccard index 0.8034 at T=12.For prostate,PSNR reaches 28.3965 dB,SSIM 0.7532,FSIM 0.8170,Dice 0.9215,and Jaccard 0.8593.In the MSD dataset,SCSOWOA achieves PSNR 29.3244 dB,SSIM 0.7118,FSIM 0.7562,Dice 0.8901,and Jaccard 0.8034,indicating consistent performance across diverse organs and modalities.The method also demonstrates high computational efficiency,with an average execution time of 1.3221 s,offering up to 40%speed improvement over conventional metaheuristics such as PSO and GWO.Overall,proposed method provides high-accuracy,low-variance,and computationally efficient segmentation,preserving both structural and perceptual fidelity.These results confirm the method’s robustness,generalizability,and practical applicability for AI-assisted diagnostic systems across histopathological and medical imaging contexts,balancing precision,structural preservation,and speed for real-world clinical deployment.
基金supported by the National Natural Science Foundation of China(No.42072182)the Science and Technology Department of Sichuan Province(No.2024NSFSC0815)supported by the Natural Gas Development Research Department,Exploration and Development Research Institute,Petro China Tarim Oilfield Company。
摘要Natural fractures serve as the primary storage spaces and flow pathways in deep to ultra-deep tight sandstone reservoirs,directly influencing hydrocarbon accumulation,preservation,and production.Borehole images offer intuitive,continuous,and high-resolution identification of natural fractures along the entire borehole.However,relying solely on complete sinusoidal curves from borehole images for fracture identification may lead to omissions,as it overlooks cases where these curves are incomplete or truncated.To address the problems and deficiencies in fracture identification,this study systematically classifies borehole image feature patterns based on core-to-log spatial position restoring.A bidirectional comparison is conducted between natu ral fractures in cores and the fracture image features in borehole images.A quantitative relationship between fracture dip angle,thin layer thickness and borehole radius was established,accompanied by a mathematical expression describing the fracture curve morphology was proposed.These findings enabled the development of an imaging response pattern for natural fractures in deep and ultra-deep tight sandstone reservoirs,incorporating key parameters such as dip angle,through-layer connectivity,and spatial position within the borehole.In the Bashijiqike-Baxigai tight-sandstone reservoirs of the Bozi-Dabei area,we estimate that approximately 24%of coreobserved fractures display distinct linear-pattern features on borehole images,whereas approximately 91%of borehole images features can be correlated with fractures observed in core.Fracture identification rates for natural fractures increased by 17%in water-based mud and by 3%in oil-based mud through the application of the natural fracture image response pattern.Moreover,this study analyzes the deviations in the matching between core fractures and image features.Finally,we further discuss the common sources of error in natural fracture identification using borehole images from multiple perspectives,including missing core responses,inconsistencies between core and borehole image features,distortion of fracture chord curve,inaccurate fracture count,misclassification of fractures,and variations in interpretation under different mud systems.The research addresses the blind spots of traditional methods in fracture identification within thin layers,not only enhancing the detection rate of natural fractu res but also further improving the accuracy of fractu re recognitio n.At the same time,it will contribute to the optimization of fracture characterization,reservoir evaluation,and production forecasting,providing a more reliable data foundation for exploration and development under complex geological conditions.
基金supported by the National Natural Science Foundation of China(Grant Nos.52264006 and 52364004)Guizhou Provincial Basic Research Program(Natural Science)(Grant No.QianKeHe Basic-ZK[2025]general program 630).
摘要The stability of concrete-rock interfaces is a critical issue in underground engineering.This study investigated the strain localization mechanism and energy evolution of concrete-sandstone specimens containing single and double interfacial cracks at various inclination angles.Acoustic emission(AE)technology and energy theory were used to analyze energy evolution,whereas digital image correlation was employed to examine strain development and fracture mechanisms.A new approach combining digital image processing and custom binarization was introduced to characterize the fractal properties of crack patterns using the box-counting method.Experimental results showed that the AE cumulative energy and stress-strain curves divided the loading process into three stages:crack closure(Ⅰ),stable crack growth(Ⅱ),and rapid crack propagation(Ⅲ).Fractal dimensions were computed for both singleand double-crack specimens using the Otsu method and the proposed binarization technique.The Otsu method yielded values of 1.457,1.482,1.131,1.512,1.489,1.536,1.171,and 1.491,whereas the new method produced higher values—1.6038,1.6643,1.2713,1.6806,1.5594,1.6282,1.2239,and 1.6565—indicating enhanced fractal characteristics.Furthermore,the proposed method detected a three-phase evolution in fractal dimension before failure,which follows an initial increase,a stable period,and a finalrapid rise.These findingsprovide theoretical support for the application of the proposed method in underground engineering.
基金funded by the National Natural Science Foundation of China(Nos.62372100 and 62371118).
摘要This paper firstly analyzes the characteristics of medical images,and then proposes a specialized compressive sensing algorithm called sparsity-precise iterative hard thresholding(SIHT),which is specifically designed to address their specific features such as low sparsity and low frequency.SIHT adaptively measures sparsity and step length which becomes more precise during the iteration process to achieve a certain quality improvement in medical image reconstruction.Experimental results demonstrate that as compared to other image compressive sensing(ICS)reconstruction algorithms across three different types of medical image datasets,SIHT can achieve the best subjective recovery quality particularly in terms of mitigating blocky artifacts and noise,where a notable improvement is obtained in terms of peak signal-to-noise ratio(PSNR)and structural similarity index measurement(SSIM)of medical ICS reconstruction.
基金funded by Anhui Province University Key Science and Technology Project(2024AH053415)Anhui Province University Major Science and Technology Project(2024AH040229)+3 种基金Talent Research Initiation Fund Project of Tongling University(2024tlxyrc019)Tongling University School-Level Scientific Research Project(2024tlxyptZD07)TheUniversity Synergy Innovation Programof Anhui Province(GXXT-2023-050)Tongling City Science and Technology Major Special Project(Unveiling and Commanding Model)(200401JB004).
摘要In the image fusion field,fusing infrared images(IRIs)and visible images(VIs)excelled is a key area.The differences between IRIs and VIs make it challenging to fuse both types into a high-quality image.Accordingly,efficiently combining the advantages of both images while overcoming their shortcomings is necessary.To handle this challenge,we developed an end-to-end IRI andVI fusionmethod based on frequency decomposition and enhancement.By applying concepts from frequency domain analysis,we used the layering mechanism to better capture the salient thermal targets from the IRIs and the rich textural information from the VIs,respectively,significantly boosting the image fusion quality and effectiveness.In addition,the backbone network combined Restormer Blocks and Dense Blocks;Restormer blocks utilize global attention to extract shallow features.Meanwhile,Dense Blocks ensure the integration between shallow and deep features,thereby avoiding the loss of shallow attributes.Extensive experiments on TNO and MSRS datasets demonstrated that the suggested method achieved state-of-the-art(SOTA)performance in various metrics:Entropy(EN),Mutual Information(MI),Standard Deviation(SD),The Structural Similarity Index Measure(SSIM),Fusion quality(Qabf),MI of the pixel(FMIpixel),and modified Visual Information Fidelity(VIFm).
基金supported by the National Natural Science Foundation of China(Nos.12372020,12202106 and 12102299).
摘要Industrial anomaly detection is dedicated to identifying and locating regions that deviate from the standard appearance.The prevailing approach achieves unsupervised anomaly detection through the reconstruction of images using autoencoders.Due to the simplistic structure of some abnormal regions,the autoencoder can effectively reconstruct these areas,consequently diminishing the model’s anomaly detection capabilities.To address this issue,this paper transforms the reconstruction task into the inpainting-filling-reconstruction task to increase the reconstruction error between abnormal samples and normal samples.The masked regions inpainted by the filling network are used to fill in the input image,thereby achieving an effect similar to masking.Unlike typical masking processes,this approach retains partial authentic information in the input image,rendering it partially visible.This is beneficial for the reconstruction network to repair the masked areas.Due to the consistent structure between the masked region inpainted by the filling network and the normal region,the filled abnormal regions display a complex structure that has not been learned,making it difficult for the reconstruction network to reconstruct the abnormal regions.Experimental results indicate that our method performs better than other methods on both the MVTec AD dataset and the MVTec LOCO AD dataset.
基金supported by the National Natural Science Foundation of China(62506268,62276192)。
摘要Dear Editor,Due to the scarcity of high-quality infrared data,translating visible images to infrared has become a practical solution to meet the growing demand for infrared images in low-light and adverse conditions.Due to the large modality gap and limited prior information,existing visible-to-infrared(VIS-to-IR)image translation methods often struggle with poor structural preservation,unclear cross-modal correspondence,and loss of thermal details.
基金Supported by the Science and Technology Project from State Grid Corporation of China (No.5700-202490330A-2-1-ZX)。
摘要To address the issue of inconsistent image quality and data scarcity in bolt defect detection for transmission lines,this paper proposes an improved sparse region-based convolutional neural network(RCNN) based detection framework integrating image quality evaluation and text-to-image data augmentation.First,a HyperNetwork-based image quality assessment module is introduced to filter low-quality inspection images in terms of clarity and structural integrity,resulting in a high-quality training dataset.Second,a text-to-image diffusion model is utilized for sample augmentation.By designing text prompts that describe various bolt defect types under diverse lighting and viewing conditions,the model automatically generates realistic synthetic samples.The generated images are further filtered using a combination of quality and perceptual similarity metrics to ensure consistency with the real data distribution.Building upon the sparse RCNN baseline,a dynamic label assignment mechanism and a random decision path detection head are incorporated to enhance bounding box matching and prediction accuracy.Experimental results demonstrate that the proposed method significantly improves detection accuracy(mAP@0.5) over the original sparse RCNN while maintaining low computational cost,enabling more efficient and intelligent inspection of transmission line components.
基金supported by the Natural Science Fund of Heilongjiang Province(No.PL2024F027)the National Natural Science Foundation of China(No.61601174)。
摘要The current infrared image pedestrian detectors have problems with high rates of false positives and false negatives. To solve these problems, we proposed an improved anchor-free fully convolutional one-stage object detection(FCOS) algorithm. Firstly, we introduced the channel attention module squeeze excitation(SE)-Block in the FCOS backbone network, which was used to learn how to model the relative importance between different feature channels, and to achieve the weight recalibration of the features extracted from the convolution neural network, and improve the weight values that are more important for pedestrian target detection. Secondly, soft non-maximum suppression(Soft-NMS) replaced the conventional NMS within the algorithm's post-processing phase, which was used to reduce the probability of missed detection for occluded pedestrians. The experimental results show that our improved FCOS algorithm improves the average precision(AP) by 6.71% on the original dataset and 7.97% on the augmented KAIST pedestrian dataset compared with the original FCOS algorithm. Our improvements effectively meet the real-time requirements and there is no significant decrease in speed compared with the original FCOS algorithm, and decreased the false positives and false negatives for infrared image pedestrian detection.
基金supported by the National Natural Science Foundation of China(42130813).
摘要Distributive Fluvial Systems(DFS)are critical sedimentary systems governing fluvial dynamics,sediment transport,and ecosystem sustainability in modern and ancient basins.Accurate quantification of DFS channel morphology is essential for advancing sedimentary modeling,optimizing water resource management,and mitigating fluvial hazards.Here,the authors present a novel automated framework that extracts DFS channel networks from remote sensing imagery by integrating multiscale image segmentation,fractal network evolution,and region-merging algorithms.Through hierarchically multiresolution feature processing,this method overcomes limitations of traditional single-scale analysis,enabling adaptive extraction while reducing segmentation heterogeneity.Specifically,the workflow consists of three stages:Image segmentation,feature extraction,and image classification.When applied to the Golmud fluvial fan(Qinghai,China),this approach achieves 90.2%overall channel extraction accuracy using 0.5 m resolution imagery,significantly outperforming traditional DEM-based(81.7%)and water spectral methods(85.4%)in resolving fine-scale channel networks.Crucially,the framework demonstrates robust adaptability to complex sedimentary environments with variable vegetation cover(<30%density)and spectral noise,providing a time-efficient,data-agnostic solution for DFS characterization.
基金supported by the Fundamental Research Funds for the Central Universities(WK3490000006).
摘要Lossy image coding is the art of computing that is principally bounded by the image’s rate-distortion function.This bound,though never accurately characterized,has been approached practically via deep learning technologies in recent years.Indeed,learned image coding schemes allow direct optimization of the joint rate-distortion cost,thereby outperforming the handcrafted image coding schemes by a large margin.Still,it is observed that there is room for further improvement in the rate-distortion performance of learned image coding.In this article,we identify the gap between the ideal rate-distortion function forecasted by Shannon’s information theory and the empirical rate-distortion function achieved by the state-of-the-art learned image coding schemes,revealing that the gap is incurred by five different effects:modeling effect,approximation effect,amortization effect,digitization effect,and asymptotic effect.We design simulations and experiments to quantitatively evaluate the last three effects,which demonstrates the high potential of future lossy image coding technologies.
基金supported by the National Key R&D Program of China(No.2022YFC2504403)the National Natural Science Foundation of China(No.62172202)+1 种基金the Experiment Project of China Manned Space Program(No.HYZHXM01019)the Fundamental Research Funds for the Central Universities from Southeast University(No.3207032101C3)。
摘要Organoids possess immense potential for unraveling the intricate functions of human tissues and facilitating preclinical disease treatment.Their applications span from high-throughput drug screening to the modeling of complex diseases,with some even achieving clinical translation.Changes in the overall size,shape,boundary,and other morphological features of organoids provide a noninvasive method for assessing organoid drug sensitivity.However,the precise segmentation of organoids in bright-field microscopy images is made difficult by the complexity of the organoid morphology and interference,including overlapping organoids,bubbles,dust particles,and cell fragments.This paper introduces the precision organoid segmentation technique(POST),which is a deep-learning algorithm for segmenting challenging organoids under simple bright-field imaging conditions.Unlike existing methods,POST accurately segments each organoid and eliminates various artifacts encountered during organoid culturing and imaging.Furthermore,it is sensitive to and aligns with measurements of organoid activity in drug sensitivity experiments.POST is expected to be a valuable tool for drug screening using organoids owing to its capability of automatically and rapidly eliminating interfering substances and thereby streamlining the organoid analysis and drug screening process.
基金supported by the National Natural Science Foundation of China(Nos.62375144 and 12404345)Key Research and Development Program of Liaoning Province(No.2025JH2/102800050)+1 种基金the Funding from National Key Laboratory of Particle Transport and Separation Technology(No.KGKF-2024-3)the Fundamental Research Funds for the Central Universities",Nankai University(No.63241331).
摘要Motion artifacts and noise are key factors that determine image quality in optical coherence tomography angiography(OCTA).Although deep learning has emerged as an effective method for artifacts removal and denoising,its generalization capability remains limited,and it is difficult to handle images with both motion artifacts and noise.To address this issue,we designed a Swin Transformer-based multi-scale motion artifacts and noise parallel removal network(ST-MANPR)to learn the nonlinear mapping between images with motion artifacts and noise and images without them,thereby achieving simultaneous suppression of motion artifacts and noise.The proposed network integrates the Swin window attention module,channel attention module(CAM),and adaptive multi-scale convolution denoising module(AMS-CNN)to enhance its capability in processing complex image features.At the same time,a hybrid loss function combining wavelet transform(WT)and mean square error(MSE)was introduced to facilitate high-frequency detail restoration.In addition,we constructed a dataset of OCTA images with and without motion artifacts and noise at multiple intensity levels.The created dataset was applied to train the network and the test results were evaluated both visually and numerically.The experimental results show that the proposed network can effectively remove motion artifacts and noise in OCTA images simultaneously.
基金supported by the National Natural Science Foundation of China(No.42101362)the Natural Science Foundation of Henan Province(No.252300421158)+1 种基金the Shenzhen Science and Technology Program(No.JCYJ20220530162001003)the Science and Technology Development Program of Henan Province(No.242300421639),China。
摘要Accurate quantification of crop residue cover(CRC)is crucial for monitoring and evaluating conservation tillage practices,yet it poses a significant image segmentation challenge.The subtle visual distinctions between fragmented residue and soil,compounded by variable illumination and shadows in field imagery,often lead to poor segmentation performance.To overcome these limitations,we introduce RCTUnet,a novel deep learning architecture designed for robust crop-residue-soil segmentation and precise CRC estimation.RCTUnet’s architecture synergistically integrates three key components:(1)a ResNet50 backbone for deep,multi-scale feature extraction;(2)a convolutional block attention module(CBAM)to adaptively focus on salient residue features across both channel and spatial dimensions;and(3)a transformer-based global context fusion module(GCFM)to model long-range spatial dependencies,which is critical for interpreting heterogeneous residue patterns.We evaluated RCTUnet on a dataset of 1220 field-acquired images spanning four typical crop rotations.Experimental results show that,compared to traditional models:(1)RCTUnet achieves significantly higher crop-residue-soil segmentation accuracy than classic models including Unet,Unet++,DeepLabV3,segmentation network(SegNet),and fully convolutional network(FCN),with improvements of 3.24%,3.42%,4.88%,8.28%,and 6.05%in overall accuracy,respectively;(2)RCTUnet yields superior residue-soil segmentation performance,with increases in residue recall of 7.67%,7.37%,14.09%,27.05%,and 16.91%,respectively;(3)RCTUnet shows enhanced CRC estimation accuracy,achieving a root mean square error(RMSE)of 4.875,representing a 45.5%improvement over Unet(RMSE=8.941).These results demonstrate the efficacy of our hybrid approach,which combines deep hierarchical features,dual-domain attention,and global context modeling.RCTUnet provides a robust and reliable tool for automated CRC assessment,advancing the capabilities of in-field agricultural monitoring.