Owing to the chaotic and non-integrable nature of three-body dynamics,the conventional Keplerian elements are rendered inadequate for cataloging cislunar space objects.Currently,there has been a conspicuous absence of...Owing to the chaotic and non-integrable nature of three-body dynamics,the conventional Keplerian elements are rendered inadequate for cataloging cislunar space objects.Currently,there has been a conspicuous absence of universally recognized parameters for the characterization and cataloging of such objects,thereby posing an urgent challenge to cislunar space situational awareness.This paper proposes a novel approach to parameterize the orbits of Earth-Moon collinear libration points by leveraging the theoretical frameworks of canonical transformations.First,under the Hamiltonian-form dynamical equations of the libration point,symplectic transformations are employed to extract 3 modes of motion from locally linearized part.A subsequent canonical transformation then decouples the hyperbolic invariant manifold from the center manifold within the nonlinear remainder.Finally,6 characteristic parameters obtained via action-angle variables are established in a bijective correspondence with the state variables,where two parameters characterize the motion of the invariant manifold and four parameters characterize the motion of the central manifold.Furthermore,a distribution map of the Earth-Moon libration point orbits is drawn utilizing Poincare sections,which can be used to describe the distribution of libration point object.Simulation results demonstrate that the proposed parameters are not only applicable to orbit identification and object cataloging but also exhibit remarkable consistency and robustness against variations in observation arc length and observational errors.展开更多
Road Abandoned Objects(RAOs)pose significant threats to traffic safety,particularly due to their small size,irregular shapes,and unpredictable distribution in complex road environments.The primary objective of this st...Road Abandoned Objects(RAOs)pose significant threats to traffic safety,particularly due to their small size,irregular shapes,and unpredictable distribution in complex road environments.The primary objective of this study is to develop an accurate and real-time detection framework for RAOs while maintaining low computational cost for practical deployment.To achieve this,we propose RAO-YOLO,a lightweight vision-based detection framework built upon an enhanced YOLO architecture.Specifically,a Mixed Aggregation Network(MANet)is introduced to improve multi-scale feature representation,and a Lightweight Shared Detail-Enhanced Detection(LSDD)head is designed to enhance localization accuracy for small and irregular objects.Furthermore,a Focal-MPDIoU loss function is proposed to address sample imbalance and geometric irregularity during training.Extensive experiments conducted on the RAOD dataset demonstrate that the proposed method achieves superior performance compared to state-of-the-art detectors,achieving a mAP@0.5:0.95 of 56.1%while maintaining real-time inference speed.These results validate the effectiveness of the proposed framework for practical intelligent transportation applications.展开更多
To support the process of grasping objects on a tabletop for the blind or robotic arm,it is necessary to address fundamental computer vision tasks,such as detecting,recognizing,and locating objects in space,and determ...To support the process of grasping objects on a tabletop for the blind or robotic arm,it is necessary to address fundamental computer vision tasks,such as detecting,recognizing,and locating objects in space,and determining the position of the grasping information.These results can then be used to guide the visually impaired or to execute grasping tasks with a robotic arm.In this paper,we collected,annotated,and published the benchmark TQUGraspingObject dataset for testing,validation,and evaluation of deep learning(DL)models for detecting,recognizing,and localizing grasping objects in 2D and 3D space,especially 3D point cloud data.Our dataset is collected in a shared room,with common everyday objects placed on the tabletop in jumbled positions by Intel RealSense D435(IR-D435).This dataset includes more than 63k RGB-D pairs and related data such as normalized 3D object point cloud,3D object point cloud segmented,coordinate system normalizationmatrix,3D object point cloud normalized,and hand pose for grasping each object.At the same time,we also conducted experiments on fourDL networks with the best performance:SSD-MobileNetV3,ResNet50-Transformer,ResNet101-Transformer,and YOLOv12.The results present that YOLOv12 has the most suitable results in detecting and recognizing objects in images.All data,annotations,toolkit,source code,point cloud data,and results are publicly available on our project website:http://gffzz188fe103f8f1460asqu99ouuwwk5x6opf.ffgz.tsg.suse.edu.cn/HuaTThanhIT2327Tqu/datasetv2.展开更多
To improve the accuracy of small object feature detection in complex backgrounds for Unmanned Aerial Vehicle(UAV)aerial photography and reduce computational complexity,we propose the lightweight UAV aerial photography...To improve the accuracy of small object feature detection in complex backgrounds for Unmanned Aerial Vehicle(UAV)aerial photography and reduce computational complexity,we propose the lightweight UAV aerial photography small object detection method based on multi-scale feature fusion and contextual information.Firstly,by introducing the grouped content-aware reassembly(GCA)operator and designing lightweight pinwheel context convolution(LPConv),we extend the feature fusion path to the P2 layer,constructing a lightweight multi-scale feature fusion network(SG-PANet).Through the decoupling of fine-grained small object features and background interference features by the GCA operator,combined with the anisotropic receptive field constructed by LPConv,our proposed method can effectively preserve the geometric details of small objects.Furthermore,we introduce the cross-stage dense feature refinement(CSPStage)module as the pre-refining unit of the detection head,and use the full history state awareness mechanism to strengthen feature reuse and gradient propagation to solve the problem of feature degradation across layers.We utilize the Wise-IoU v3 loss function to dynamically optimize the gradient gains of high-quality and low-quality samples,thereby enhancing the detection accuracy and convergence speed of the proposed method in complex scenarios.Finally,we verified the superiority and generalization of the proposed method on the VisDrone2019 dataset and DOTAv1.5 dataset.The results show that compared with YOLOv11n,MFCI-YOLO’s detection mAP50-95 increased by 11.1%,small object mAP50 increased by 16.1%,and mAP50 reached 80.3%.It provides a practical solution for detecting small objects in dense scenes.展开更多
In modern industrial production,foreign object detection in complex environments is crucial to ensure product quality and production safety.Detection systems based on deep-learning image processing algorithms often fa...In modern industrial production,foreign object detection in complex environments is crucial to ensure product quality and production safety.Detection systems based on deep-learning image processing algorithms often face challenges with handling high-resolution images and achieving accurate detection against complex backgrounds.To address these issues,this study employs the PatchCore unsupervised anomaly detection algorithm combined with data augmentation techniques to enhance the system’s generalization capability across varying lighting conditions,viewing angles,and object scales.The proposed method is evaluated in a complex industrial detection scenario involving the bogie of an electric multiple unit(EMU).A dataset consisting of complex backgrounds,diverse lighting conditions,and multiple viewing angles is constructed to validate the performance of the detection system in real industrial environments.Experimental results show that the proposed model achieves an average area under the receiver operating characteristic curve(AUROC)of 0.92 and an average F1 score of 0.85.Combined with data augmentation,the proposed model exhibits improvements in AUROC by 0.06 and F1 score by 0.03,demonstrating enhanced accuracy and robustness for foreign object detection in complex industrial settings.In addition,the effects of key factors on detection performance are systematically analyzed,providing practical guidance for parameter selection in real industrial applications.展开更多
Structured study of spatial objects and their relationships leads to a better cognition of the geospatial information and creates the concept of context at a higher level of abstraction.This study is aimed at providin...Structured study of spatial objects and their relationships leads to a better cognition of the geospatial information and creates the concept of context at a higher level of abstraction.This study is aimed at providing a comprehensive definition of the context for geospatial objects.A combination of binary qualitative spatial relationships(i.e.direction,distance,and topological relations)among the members of a set of spatial objects will be used accordingly.In addition,by incorporating the general concept of context,obtained from either static data(attributes in a database)or dynamic data(sensors),the compact context of spatial objects will be introduced.Our framework for presentation of the involved knowledge and conception about the objects in context is also explored using ontology and description logic because of powerful conceptualization of relationships,either spatial or non-spatial,integrally.For this purpose,the hierarchies of main structure and object properties are formed at first.The constraint and characteristics of classes,such as subclasses,equivalent classes,cardinality etc.,and object properties,such as being functional,transitive,symmetric,asymmetric,inverse functional,disjoint etc.,are discovered and presented in more detail using web ontology language in description logic mode.The implementation is then performed in the framework of semantic web and extensible markup language syntaxes.The method ultimately facilitates,spatial reasoning by effective querying in a semantic framework taking pellet reasoner and SPARQL(a recursive acronym for SPARQL Protocol and RDF Query Language).展开更多
Accurate segmentation of camouflage objects in aerial imagery is vital for improving the efficiency of UAV-based reconnaissance and rescue missions.However,camouflage object segmentation is increasingly challenging du...Accurate segmentation of camouflage objects in aerial imagery is vital for improving the efficiency of UAV-based reconnaissance and rescue missions.However,camouflage object segmentation is increasingly challenging due to advances in both camouflage materials and biological mimicry.Although multispectral-RGB based technology shows promise,conventional dual-aperture multispectral-RGB imaging systems are constrained by imprecise and time-consuming registration and fusion across different modalities,limiting their performance.Here,we propose the Reconstructed Multispectral-RGB Fusion Network(RMRF-Net),which reconstructs RGB images into multispectral ones,enabling efficient multimodal segmentation using only an RGB camera.Specifically,RMRF-Net employs a divergentsimilarity feature correction strategy to minimize reconstruction errors and includes an efficient boundary-aware decoder to enhance object contours.Notably,we establish the first real-world aerial multispectral-RGB semantic segmentation of camouflage objects dataset,including 11 object categories.Experimental results demonstrate that RMRF-Net outperforms existing methods,achieving 17.38 FPS on the NVIDIA Jetson AGX Orin,with only a 0.96%drop in mIoU compared to the RTX 3090,showing its practical applicability in multimodal remote sensing.展开更多
In contrast to the nearly fixed flying altitude of satellite remote sensing platforms,aerial remote sensing(e.g.,unmanned aerial vehicles)often employs oblique photography at varying flying altitudes to observe object...In contrast to the nearly fixed flying altitude of satellite remote sensing platforms,aerial remote sensing(e.g.,unmanned aerial vehicles)often employs oblique photography at varying flying altitudes to observe objects from multiple angles and distances in real time.While the existing oriented object detection methods have already demonstrated reliable results in most satellite remote sensing scenarios and achieved high detection precision on large public datasets,such as DOTA-v1.0 and DIOR-R,these methods tend to perform suboptimally on aerial remote sensing images.This performance gap is primarily due to the following two challenges:(A)significant shape variation of objects under multi-view imaging scenarios and(B)substantial object scale variation under multi-distance imaging conditions.To address these issues,we propose the SAA-O2DINO(oriented object detection transformer with improved denoising anchor boxes and shape-adaptive assigner)method for aerial remote sensing in this paper.The proposed method is based on the recently developed AO2DINO framework.It introduces an enhanced Shape-Adaptive Assigner(SAA)that incorporates object shape information into the threshold estimation,allowing for more accurate separation of positive and negative samples,thereby improving the model's adaptability to significant shape changes across different imaging angles.Additionally,a Gradient Calibration Loss(GCL)is introduced to mitigate the problem of object scale variation.The GCL employs a gradient scaling strategy to reduce scale sensitivity during the optimisation process.We comprehensively compare the proposed method against typical oriented object detection approaches on the DOTA-v1.0 and VSAI datasets.The results show that the proposed method has substantial improvement in detection performance across all datasets,particularly for aerial remote sensing images,validating the generalisation capabilities of our model.展开更多
Visible and infrared(RGB-IR)fusion object detection plays an important role in security,disaster relief,etc.In recent years,deep-learning-based RGB-IR fusion detection methods have been developing rapidly,but still st...Visible and infrared(RGB-IR)fusion object detection plays an important role in security,disaster relief,etc.In recent years,deep-learning-based RGB-IR fusion detection methods have been developing rapidly,but still struggle to deal with the complex and changing scenarios captured by drones,mainly due to two reasons:(A)RGB-IR fusion detectors are susceptible to inferior inputs that degrade performance and stability.(B)RGB-IR fusion detectors are susceptible to redundant features that reduce accuracy and efficiency.In this paper,an innovative RGB-IR fusion detection framework based on global-local feature optimization,named GLFDet,is proposed to improve the detection performance and efficiency of drone-captured objects.The key components of GLFDet include a Global Feature Optimization(GFO)module,a Local Feature Optimization(LFO)module and a Channel Separation Fusion(CSF)module.Specifically,GFO calculates the information content of the input image from the frequency domain and optimizes the features holistically.Then,LFO dynamically selects high-value features and filters out low-value features before fusion,which significantly improves the efficiency of fusion.Finally,CSF fuses the RGB and IR features across the corresponding channels,which avoids the rearrangement of the channel relationships and enhances the model stability.Extensive experimental results show that the proposed method achieves the best performance on three popular RGB-IR datasets Drone Vehicle,VEDAI,and LLVIP.In addition,GLFDet is more lightweight than other comparable models,making it more appealing to edge devices such as drones.The code is available at http://gffzz188fe103f8f1460asqu99ouuwwk5x6opf.ffgz.tsg.suse.edu.cn/lao chen330/GLFDet.展开更多
Most image-based object detection methods employ horizontal bounding boxes(HBBs)to capture objects in tunnel images.However,these bounding boxes often fail to effectively enclose objects oriented in arbitrary directio...Most image-based object detection methods employ horizontal bounding boxes(HBBs)to capture objects in tunnel images.However,these bounding boxes often fail to effectively enclose objects oriented in arbitrary directions,resulting in reduced accuracy and suboptimal detection performance.Moreover,HBBs cannot provide directional information for rotated objects.This study proposes a rotated detection method for identifying apparent defects in shield tunnels.Specifically,the oriented region-convolutional neural network(oriented R-CNN)is utilized to detect rotated objects in tunnel images.To enhance feature extraction,a novel hybrid backbone combining CNN-based networks with Swin Transformers is proposed.A feature fusion strategy is employed to integrate features extracted from both networks.Additionally,a neck network based on the bidirectional-feature pyramid network(Bi-FPN)is designed to combine multi-scale object features.The bolt hole dataset is curated to evaluate the efficacyof the proposed method.In addition,a dedicated pre-processing approach is developed for large-sized images to accommodate the rotated,dense,and small-scale characteristics of objects in tunnel images.Experimental results demonstrate that the proposed method achieves a more than 4%improvement in mAP50-95compared to other rotated detectors and a 6.6%-12.7%improvement over mainstream horizontal detectors.Furthermore,the proposed method outperforms mainstream methods by 6.5%-14.7%in detecting leakage bolt holes,underscoring its significant engineering applicability.展开更多
The integrity of perception data transmitted over in-vehicle networks is important for the safety of autonomous driving.However,legacy protocols like the Controller Area Network(CAN)bus which lacks essential security ...The integrity of perception data transmitted over in-vehicle networks is important for the safety of autonomous driving.However,legacy protocols like the Controller Area Network(CAN)bus which lacks essential security features make In-Vehicle Networks(IVNs)vulnerable to data tampering attacks.Current research typically focuses on detecting the attack itself but ignores the information recovery from the missing data,leading to an unsafe autonomous driving system.To address the issue,we propose a 3D object recovery framework to recover the missing data caused by the tampering attack that occurred in in-vehicle networks.The proposed framework exploits both temporal and spatial context for the 3D object recovery,where a temporal branch is designed to learn the coordinate offsets of 3D objects based on historical data from previous frames,while a spatial branch employs information from the adjacent views of the attacked objects to locate the recovered objects from the overlapped regions in the current frame.By integrating the temporal and spatial clues,the framework effectively recovers the missing objects from the resting ones,thereby enhancing the immunity of in-vehicle networks for the tampering attack.Extensive experiments on the nuScenes dataset demonstrate that the proposed framework significantly improves 3D object detection performance under the attack when compared to the method without recovery.Additionally,the recovery performance becomes better as the attack intensity increases,highlighting the framework’s robustness in high-risk scenarios.The source will be available upon publication.展开更多
[Objective]Detecting dense and small aquaculture net cages in complex backgrounds is difficult,the purpose of this study is to build a specialized dataset and design a targeted detection model that enhances recognitio...[Objective]Detecting dense and small aquaculture net cages in complex backgrounds is difficult,the purpose of this study is to build a specialized dataset and design a targeted detection model that enhances recognition accuracy and robustness for practical aquaculture management.[Methods]A dataset of aquaculture net cages was constructed using highresolution remote sensing imagery collected from seven representative farming regions(Australia,Canada,Chile,Croatia,Greece,China,and the Faroe Islands),and Cage-YOLO,a deep learning model based on YOLOv5,was proposed for detecting dense and small aquaculture net cages.First,an adaptive dense perception algorithm was introduced,which automatically selects and generates feature maps that reflect the high-density distribution of small aquaculture net cages.Second,an enhanced module based on spatial pyramid pooling fast was integrated to effectively reduce background noise interference and improve global feature extraction capabilities.Finally,a mixed attention block was incorporated to further enhance the model's perception of dense and small objects.[Results and Discussions]Experimental results showed that the proposed Cage-YOLO achieved improvements over the original YOLOv5 in terms of precision,recall,and mean average precision by 5.6,21.8,and 17.4 percentage points,respectively.The model size was maintained at 16.9 MB,demonstrating both strong performance and deployment advantages.[Conclusions]This study provides a new approach for dense and small object detection and offers technical support for the intelligent management of marine cage aquaculture.展开更多
High-resolution remote sensing images(HRSIs)are now an essential data source for gathering surface information due to advancements in remote sensing data capture technologies.However,their significant scale changes an...High-resolution remote sensing images(HRSIs)are now an essential data source for gathering surface information due to advancements in remote sensing data capture technologies.However,their significant scale changes and wealth of spatial details pose challenges for semantic segmentation.While convolutional neural networks(CNNs)excel at capturing local features,they are limited in modeling long-range dependencies.Conversely,transformers utilize multihead self-attention to integrate global context effectively,but this approach often incurs a high computational cost.This paper proposes a global-local multiscale context network(GLMCNet)to extract both global and local multiscale contextual information from HRSIs.A detail-enhanced filtering module(DEFM)is proposed at the end of the encoder to refine the encoder outputs further,thereby enhancing the key details extracted by the encoder and effectively suppressing redundant information.In addition,a global-local multiscale transformer block(GLMTB)is proposed in the decoding stage to enable the modeling of rich multiscale global and local information.We also design a stair fusion mechanism to transmit deep semantic information from deep to shallow layers progressively.Finally,we propose the semantic awareness enhancement module(SAEM),which further enhances the representation of multiscale semantic features through spatial attention and covariance channel attention.Extensive ablation analyses and comparative experiments were conducted to evaluate the performance of the proposed method.Specifically,our method achieved a mean Intersection over Union(mIoU)of 86.89%on the ISPRS Potsdam dataset and 84.34%on the ISPRS Vaihingen dataset,outperforming existing models such as ABCNet and BANet.展开更多
Urban green space may impact human health through complex pathways and the effect can vary across different travel contexts.Revealing these disparities in health pathways between different travel contexts may provide ...Urban green space may impact human health through complex pathways and the effect can vary across different travel contexts.Revealing these disparities in health pathways between different travel contexts may provide essential and practical suggestions for sustainable developments in urban environments.In this study,we investi gated the impacts of travel contexts on people’s perceptions and evaluations of green space using a cross-sectional dataset collected in Hong Kong,China.Eight hundred participants in 4 representative communities were recruited through stratified sampling,and we identified 2,913 travel events from their two-day activity-travel diaries after rigorous cross-validation with GPS-derived trajectories.We also derived two green space exposure representa tions using fine-grained remote sensing imagery and 8 representative green space exposure indicators.Eighty logistical regression models and mixed-effects models were developed to investigate the associations with con trol of a range of potential uncertainties.Our results indicate solid and consistently positive associations between participants’measured green space exposure and perceived green space,and significant but variable effects of travel purposes,travel modes,and travel time on participants’perceptions and evaluations of green space.Walk ing significantly promotes participants’perceptions and positive evaluations of urban green space,buses are not significantly associated,and metro trains may depress the perception and evaluation.Our results provide solid evidence on how travel contexts may influence people’s perceptions and evaluations of urban green space and,thus,provide essential insights into environmental health studies and sustainable urban planning that consider green space as an important urban environmental setting.展开更多
Recognising human-object interactions(HOI)is a challenging task for traditional machine learning models,including convolutional neural networks(CNNs).Existing models show limited transferability across complex dataset...Recognising human-object interactions(HOI)is a challenging task for traditional machine learning models,including convolutional neural networks(CNNs).Existing models show limited transferability across complex datasets such as D3D-HOI and SYSU 3D HOI.The conventional architecture of CNNs restricts their ability to handle HOI scenarios with high complexity.HOI recognition requires improved feature extraction methods to overcome the current limitations in accuracy and scalability.This work proposes a Novel quantum gate-enabled hybrid CNN(QEH-CNN)for effectiveHOI recognition.Themodel enhancesCNNperformance by integrating quantumcomputing components.The framework begins with bilateral image filtering,followed bymulti-object tracking(MOT)and Felzenszwalb superpixel segmentation.A watershed algorithm refines object boundaries by cleaning merged superpixels.Feature extraction combines a histogram of oriented gradients(HOG),Global Image Statistics for Texture(GIST)descriptors,and a novel 23-joint keypoint extractionmethod using relative joint angles and joint proximitymeasures.A fuzzy optimization process refines the extracted features before feeding them into the QEH-CNNmodel.The proposed model achieves 95.06%accuracy on the 3D-D3D-HOI dataset and 97.29%on the SYSU3DHOI dataset.Theintegration of quantum computing enhances feature optimization,leading to improved accuracy and overall model efficiency.展开更多
In recent years,with the rapid advancement of artificial intelligence,object detection algorithms have made significant strides in accuracy and computational efficiency.Notably,research and applications of Anchor-Free...In recent years,with the rapid advancement of artificial intelligence,object detection algorithms have made significant strides in accuracy and computational efficiency.Notably,research and applications of Anchor-Free models have opened new avenues for real-time target detection in optical remote sensing images(ORSIs).However,in the realmof adversarial attacks,developing adversarial techniques tailored to Anchor-Freemodels remains challenging.Adversarial examples generated based on Anchor-Based models often exhibit poor transferability to these new model architectures.Furthermore,the growing diversity of Anchor-Free models poses additional hurdles to achieving robust transferability of adversarial attacks.This study presents an improved cross-conv-block feature fusion You Only Look Once(YOLO)architecture,meticulously engineered to facilitate the extraction ofmore comprehensive semantic features during the backpropagation process.To address the asymmetry between densely distributed objects in ORSIs and the corresponding detector outputs,a novel dense bounding box attack strategy is proposed.This approach leverages dense target bounding boxes loss in the calculation of adversarial loss functions.Furthermore,by integrating translation-invariant(TI)and momentum-iteration(MI)adversarial methodologies,the proposed framework significantly improves the transferability of adversarial attacks.Experimental results demonstrate that our method achieves superior adversarial attack performance,with adversarial transferability rates(ATR)of 67.53%on the NWPU VHR-10 dataset and 90.71%on the HRSC2016 dataset.Compared to ensemble adversarial attack and cascaded adversarial attack approaches,our method generates adversarial examples in an average of 0.64 s,representing an approximately 14.5%improvement in efficiency under equivalent conditions.展开更多
Object tracking in 3D space is a classical problem in computer vision.In this paper,an efficient and robust X-Triplet detection method is proposed based on the support vector machine(SVM) and an adjacent matrix for lo...Object tracking in 3D space is a classical problem in computer vision.In this paper,an efficient and robust X-Triplet detection method is proposed based on the support vector machine(SVM) and an adjacent matrix for locating and tracking objects through stereo vision with minimal feature points.The X-Triplet,denoted as Tri-X,is a composite marker consisting of three sequential X-corners.The definition and types of Tri-X markers are introduced at first.Then a fast and robust X-corner detector based on the block search strategy and SVM is proposed to extract X-corner candidates with sub-pixel locations and orientations.Thereafter the X-corner adjacent matrix(XAM) is constructed using the orientation angle error to describe the possibility that any X-corner pair form a valid edge vector.The Tri-X candidates are then extracted efficiently from the XAM.Finally once the Tri-X markers are detected in binocular images,their 6D pose information can be recovered through stereo matching and triangulation technique.When multiple targets are involved simultaneously,different Tri-X markers can be utilized to identify different objects.Experimental results show that the proposed method outperformed the state-of-the-art in terms of both accuracy and efficiency for Tri-X marker detection.In localization precision test,it achieved 0.1 mm error for the position and 1° error for the orientation.Our method exhibits great potential for utilization in user-defined specific tracking tasks,offering flexibility and adaptability to various tracking requirements,especially multi-tool tracking in medical robotics.展开更多
Intelligent transportation and autonomous driving systems have made urgent demands on the techniques with high performance on object detection in traffic scenes.This paper proposes an improved object detection model Y...Intelligent transportation and autonomous driving systems have made urgent demands on the techniques with high performance on object detection in traffic scenes.This paper proposes an improved object detection model YOLO-VSF over the YOLOv4 model,which is a representative work with excellent performance among YOLO series of object detection models.The main improvement measures include:The backbone feature extraction network CSPDarknet53 of YOLOv4 is replaced with VGG16 to improve the feature extraction capability;SENet attention mechanism is incorporated to improve the salient and correlation feature representation capability;Focal Loss is integrated into the loss function to overcome the sample imbalance problem.In addition,the detection performance of small targets is improved by increasing the resolution of input images.Experimental results show that on the VanJee traffic image dataset provided by Beijing VanJee Technology Co.,Ltd.,the proposed YOLO-VSF model achieves an average mean accuracy(mAP)of 92.21 percentage points,which improves the mAP by 3.04 percentage points compared with the YOLOv4 model while maintaining the detection speed of the original model.On the UA-DETRAC dataset,the average accuracy of YOLO-VSF is close to that of the latest YOLOv7 model with the number of parameters reduced by 1.329×107.The proposed method can provide a support for object detection in traffic scenes.展开更多
The fasteners employed in the railway tracks are susceptible to defects arising from their intricate composition.Foreign objects are frequently observed on the track bed in an open environment.These two types of defec...The fasteners employed in the railway tracks are susceptible to defects arising from their intricate composition.Foreign objects are frequently observed on the track bed in an open environment.These two types of defects pose potential threats to high-speed trains,thus necessitating timely and accurate track inspection.The majority of extant automatic inspection methods are predicated on the utilization of single visible light data,and the efficacy of the algorithmic processes is influenced by complex environments.Furthermore,due to the single information dimension,the detection accuracy of defects in similar,occluded,and small object categories is low.To address the aforementioned issues,this paper proposes a track defect detectionmethod based on dynamicmulti-modal fusion and challenging object enhanced perception.First,in light of the variances in the representation dimensions ofmultimodal information,this paper proposes a dynamic weighted multi-modal feature fusion module.The fused multi-modal features are assigned weights,and thenmultiplied with the extracted single-modal features atmultiple levels,achieving adaptive adjustment of the response degree of fusion features.Second,a novel stepwise multi-scale convolution feature aggregation module is proposed for challenging objects.The proposed method employs depth separable convolution and cross-scale aggregation operations of different receptive fields to enhance feature extraction and reuse,thereby reducing the degree of progressive loss of effective information.The experimental results demonstrate the efficacy of the proposed method in comparison to eight established methods,encompassing both single-modal and multi-modal methods,as evidenced by the extensive findings within the constructed RGBD dataset.展开更多
BACKGROUND Recognition of pelvic autonomic nerves(PAN)during total mesorectal excision(TME)largely depends on the surgeon’s expertise,making it susceptible to misrecognition and unintentional damage.There is an urgen...BACKGROUND Recognition of pelvic autonomic nerves(PAN)during total mesorectal excision(TME)largely depends on the surgeon’s expertise,making it susceptible to misrecognition and unintentional damage.There is an urgent need for objective and real-time support methods.AIM To develop a deep learning(DL)model for precise recognition and visual annotation of 5 categories of PAN during TME.METHODS This single-center retrospective study enrolled 120 TME videos from January 2021 to January 2023.A total of 3246 high-quality images were obtained and split 9:1 into training and internal test sets.Difficult-to-recognize characteristics were summarized.An additional 20 independent TME videos from June 2023 to January 2024 were used for external validation.The DL model performance was compared with that of surgeons and verified pathologically.χ2,Fisher’s exact test and t-tests were applied(P<0.05).RESULTS The DL model achieved a precision of 0.839,a recall of 0.769,and a mean average precision at intersection over union 50 of 0.873.The overall recognition rate in external validation was 76.0%,similar to senior surgeons(73.9%,P=0.156)but superior to junior surgeons(64.9%,P=0.001).The miss rate of 5 PAN categories ranged from 12.9%to 29.6%.Initial recognition time(2.08-2.21 seconds)was shorter than that of senior surgeons(5.65-6.19 seconds,P<0.01);mean continuous tracking duration was prolonged by 57.21-66.45 seconds compared with that of senior surgeons(P<0.01).Low nerve exposure caused most DL model false negatives,while cord-like fibrous tissue dominated false positives.All 7 harvested specimens were pathologically confirmed to contain nerve tissue,with a processing speed of 25 frames per second.CONCLUSION The model demonstrates recognition performance comparable to that of senior surgeons,with pathological confirmation.It may potentially help preserve PAN during TME and shorten the learning curve for junior surgeons.展开更多
基金supported by the National Level Project of China(No.KJSP2023020104)。
摘要Owing to the chaotic and non-integrable nature of three-body dynamics,the conventional Keplerian elements are rendered inadequate for cataloging cislunar space objects.Currently,there has been a conspicuous absence of universally recognized parameters for the characterization and cataloging of such objects,thereby posing an urgent challenge to cislunar space situational awareness.This paper proposes a novel approach to parameterize the orbits of Earth-Moon collinear libration points by leveraging the theoretical frameworks of canonical transformations.First,under the Hamiltonian-form dynamical equations of the libration point,symplectic transformations are employed to extract 3 modes of motion from locally linearized part.A subsequent canonical transformation then decouples the hyperbolic invariant manifold from the center manifold within the nonlinear remainder.Finally,6 characteristic parameters obtained via action-angle variables are established in a bijective correspondence with the state variables,where two parameters characterize the motion of the invariant manifold and four parameters characterize the motion of the central manifold.Furthermore,a distribution map of the Earth-Moon libration point orbits is drawn utilizing Poincare sections,which can be used to describe the distribution of libration point object.Simulation results demonstrate that the proposed parameters are not only applicable to orbit identification and object cataloging but also exhibit remarkable consistency and robustness against variations in observation arc length and observational errors.
基金partially supported by the Natural Science Foundation of China(Grant Number:52308457)China Postdoctoral Science Foundation(Grant Number:2024M761811)Natural Science Foundation of Shandong Province(Grant Number:ZR2023QE220).
摘要Road Abandoned Objects(RAOs)pose significant threats to traffic safety,particularly due to their small size,irregular shapes,and unpredictable distribution in complex road environments.The primary objective of this study is to develop an accurate and real-time detection framework for RAOs while maintaining low computational cost for practical deployment.To achieve this,we propose RAO-YOLO,a lightweight vision-based detection framework built upon an enhanced YOLO architecture.Specifically,a Mixed Aggregation Network(MANet)is introduced to improve multi-scale feature representation,and a Lightweight Shared Detail-Enhanced Detection(LSDD)head is designed to enhance localization accuracy for small and irregular objects.Furthermore,a Focal-MPDIoU loss function is proposed to address sample imbalance and geometric irregularity during training.Extensive experiments conducted on the RAOD dataset demonstrate that the proposed method achieves superior performance compared to state-of-the-art detectors,achieving a mAP@0.5:0.95 of 56.1%while maintaining real-time inference speed.These results validate the effectiveness of the proposed framework for practical intelligent transportation applications.
摘要To support the process of grasping objects on a tabletop for the blind or robotic arm,it is necessary to address fundamental computer vision tasks,such as detecting,recognizing,and locating objects in space,and determining the position of the grasping information.These results can then be used to guide the visually impaired or to execute grasping tasks with a robotic arm.In this paper,we collected,annotated,and published the benchmark TQUGraspingObject dataset for testing,validation,and evaluation of deep learning(DL)models for detecting,recognizing,and localizing grasping objects in 2D and 3D space,especially 3D point cloud data.Our dataset is collected in a shared room,with common everyday objects placed on the tabletop in jumbled positions by Intel RealSense D435(IR-D435).This dataset includes more than 63k RGB-D pairs and related data such as normalized 3D object point cloud,3D object point cloud segmented,coordinate system normalizationmatrix,3D object point cloud normalized,and hand pose for grasping each object.At the same time,we also conducted experiments on fourDL networks with the best performance:SSD-MobileNetV3,ResNet50-Transformer,ResNet101-Transformer,and YOLOv12.The results present that YOLOv12 has the most suitable results in detecting and recognizing objects in images.All data,annotations,toolkit,source code,point cloud data,and results are publicly available on our project website:http://gffzz188fe103f8f1460asqu99ouuwwk5x6opf.ffgz.tsg.suse.edu.cn/HuaTThanhIT2327Tqu/datasetv2.
基金supported in part by the Natural Science Foundation of Henan Province under Grant 252300423317the Science and Technology Research Project of Henan Province under Grant 262102211081+1 种基金the Key Scientific Research Projects of Colleges and Universities in Henan Province under Grant 25B510012“Pioneer”and“Leading Goose”R&DProgram of Zhejiang under grant 2026LDC01003(JT).
摘要To improve the accuracy of small object feature detection in complex backgrounds for Unmanned Aerial Vehicle(UAV)aerial photography and reduce computational complexity,we propose the lightweight UAV aerial photography small object detection method based on multi-scale feature fusion and contextual information.Firstly,by introducing the grouped content-aware reassembly(GCA)operator and designing lightweight pinwheel context convolution(LPConv),we extend the feature fusion path to the P2 layer,constructing a lightweight multi-scale feature fusion network(SG-PANet).Through the decoupling of fine-grained small object features and background interference features by the GCA operator,combined with the anisotropic receptive field constructed by LPConv,our proposed method can effectively preserve the geometric details of small objects.Furthermore,we introduce the cross-stage dense feature refinement(CSPStage)module as the pre-refining unit of the detection head,and use the full history state awareness mechanism to strengthen feature reuse and gradient propagation to solve the problem of feature degradation across layers.We utilize the Wise-IoU v3 loss function to dynamically optimize the gradient gains of high-quality and low-quality samples,thereby enhancing the detection accuracy and convergence speed of the proposed method in complex scenarios.Finally,we verified the superiority and generalization of the proposed method on the VisDrone2019 dataset and DOTAv1.5 dataset.The results show that compared with YOLOv11n,MFCI-YOLO’s detection mAP50-95 increased by 11.1%,small object mAP50 increased by 16.1%,and mAP50 reached 80.3%.It provides a practical solution for detecting small objects in dense scenes.
摘要In modern industrial production,foreign object detection in complex environments is crucial to ensure product quality and production safety.Detection systems based on deep-learning image processing algorithms often face challenges with handling high-resolution images and achieving accurate detection against complex backgrounds.To address these issues,this study employs the PatchCore unsupervised anomaly detection algorithm combined with data augmentation techniques to enhance the system’s generalization capability across varying lighting conditions,viewing angles,and object scales.The proposed method is evaluated in a complex industrial detection scenario involving the bogie of an electric multiple unit(EMU).A dataset consisting of complex backgrounds,diverse lighting conditions,and multiple viewing angles is constructed to validate the performance of the detection system in real industrial environments.Experimental results show that the proposed model achieves an average area under the receiver operating characteristic curve(AUROC)of 0.92 and an average F1 score of 0.85.Combined with data augmentation,the proposed model exhibits improvements in AUROC by 0.06 and F1 score by 0.03,demonstrating enhanced accuracy and robustness for foreign object detection in complex industrial settings.In addition,the effects of key factors on detection performance are systematically analyzed,providing practical guidance for parameter selection in real industrial applications.
摘要Structured study of spatial objects and their relationships leads to a better cognition of the geospatial information and creates the concept of context at a higher level of abstraction.This study is aimed at providing a comprehensive definition of the context for geospatial objects.A combination of binary qualitative spatial relationships(i.e.direction,distance,and topological relations)among the members of a set of spatial objects will be used accordingly.In addition,by incorporating the general concept of context,obtained from either static data(attributes in a database)or dynamic data(sensors),the compact context of spatial objects will be introduced.Our framework for presentation of the involved knowledge and conception about the objects in context is also explored using ontology and description logic because of powerful conceptualization of relationships,either spatial or non-spatial,integrally.For this purpose,the hierarchies of main structure and object properties are formed at first.The constraint and characteristics of classes,such as subclasses,equivalent classes,cardinality etc.,and object properties,such as being functional,transitive,symmetric,asymmetric,inverse functional,disjoint etc.,are discovered and presented in more detail using web ontology language in description logic mode.The implementation is then performed in the framework of semantic web and extensible markup language syntaxes.The method ultimately facilitates,spatial reasoning by effective querying in a semantic framework taking pellet reasoner and SPARQL(a recursive acronym for SPARQL Protocol and RDF Query Language).
基金National Natural Science Foundation of China(Grant Nos.62005049 and 62072110)Natural Science Foundation of Fujian Province(Grant No.2020J01451).
摘要Accurate segmentation of camouflage objects in aerial imagery is vital for improving the efficiency of UAV-based reconnaissance and rescue missions.However,camouflage object segmentation is increasingly challenging due to advances in both camouflage materials and biological mimicry.Although multispectral-RGB based technology shows promise,conventional dual-aperture multispectral-RGB imaging systems are constrained by imprecise and time-consuming registration and fusion across different modalities,limiting their performance.Here,we propose the Reconstructed Multispectral-RGB Fusion Network(RMRF-Net),which reconstructs RGB images into multispectral ones,enabling efficient multimodal segmentation using only an RGB camera.Specifically,RMRF-Net employs a divergentsimilarity feature correction strategy to minimize reconstruction errors and includes an efficient boundary-aware decoder to enhance object contours.Notably,we establish the first real-world aerial multispectral-RGB semantic segmentation of camouflage objects dataset,including 11 object categories.Experimental results demonstrate that RMRF-Net outperforms existing methods,achieving 17.38 FPS on the NVIDIA Jetson AGX Orin,with only a 0.96%drop in mIoU compared to the RTX 3090,showing its practical applicability in multimodal remote sensing.
基金supported by the National Natural Science Foundation of China(No.12472189)the Science and Technology Innovation Program of Hunan Province,China(No.2022RC11966)。
摘要In contrast to the nearly fixed flying altitude of satellite remote sensing platforms,aerial remote sensing(e.g.,unmanned aerial vehicles)often employs oblique photography at varying flying altitudes to observe objects from multiple angles and distances in real time.While the existing oriented object detection methods have already demonstrated reliable results in most satellite remote sensing scenarios and achieved high detection precision on large public datasets,such as DOTA-v1.0 and DIOR-R,these methods tend to perform suboptimally on aerial remote sensing images.This performance gap is primarily due to the following two challenges:(A)significant shape variation of objects under multi-view imaging scenarios and(B)substantial object scale variation under multi-distance imaging conditions.To address these issues,we propose the SAA-O2DINO(oriented object detection transformer with improved denoising anchor boxes and shape-adaptive assigner)method for aerial remote sensing in this paper.The proposed method is based on the recently developed AO2DINO framework.It introduces an enhanced Shape-Adaptive Assigner(SAA)that incorporates object shape information into the threshold estimation,allowing for more accurate separation of positive and negative samples,thereby improving the model's adaptability to significant shape changes across different imaging angles.Additionally,a Gradient Calibration Loss(GCL)is introduced to mitigate the problem of object scale variation.The GCL employs a gradient scaling strategy to reduce scale sensitivity during the optimisation process.We comprehensively compare the proposed method against typical oriented object detection approaches on the DOTA-v1.0 and VSAI datasets.The results show that the proposed method has substantial improvement in detection performance across all datasets,particularly for aerial remote sensing images,validating the generalisation capabilities of our model.
基金supported by the National Natural Science Foundation of China(No.62276204)the Fundamental Research Funds for the Central Universities,China(No.YJSJ24011)+1 种基金the Natural Science Basic Research Program of Shaanxi,China(Nos.2022JM-340 and 2023-JC-QN-0710)the China Postdoctoral Science Foundation(Nos.2020T130494 and 2018M633470)。
摘要Visible and infrared(RGB-IR)fusion object detection plays an important role in security,disaster relief,etc.In recent years,deep-learning-based RGB-IR fusion detection methods have been developing rapidly,but still struggle to deal with the complex and changing scenarios captured by drones,mainly due to two reasons:(A)RGB-IR fusion detectors are susceptible to inferior inputs that degrade performance and stability.(B)RGB-IR fusion detectors are susceptible to redundant features that reduce accuracy and efficiency.In this paper,an innovative RGB-IR fusion detection framework based on global-local feature optimization,named GLFDet,is proposed to improve the detection performance and efficiency of drone-captured objects.The key components of GLFDet include a Global Feature Optimization(GFO)module,a Local Feature Optimization(LFO)module and a Channel Separation Fusion(CSF)module.Specifically,GFO calculates the information content of the input image from the frequency domain and optimizes the features holistically.Then,LFO dynamically selects high-value features and filters out low-value features before fusion,which significantly improves the efficiency of fusion.Finally,CSF fuses the RGB and IR features across the corresponding channels,which avoids the rearrangement of the channel relationships and enhances the model stability.Extensive experimental results show that the proposed method achieves the best performance on three popular RGB-IR datasets Drone Vehicle,VEDAI,and LLVIP.In addition,GLFDet is more lightweight than other comparable models,making it more appealing to edge devices such as drones.The code is available at http://gffzz188fe103f8f1460asqu99ouuwwk5x6opf.ffgz.tsg.suse.edu.cn/lao chen330/GLFDet.
基金support from the National Natural Science Foundation of China(Grant Nos.52025084 and 52408420)the Beijing Natural Science Foundation(Grant No.8244058).
摘要Most image-based object detection methods employ horizontal bounding boxes(HBBs)to capture objects in tunnel images.However,these bounding boxes often fail to effectively enclose objects oriented in arbitrary directions,resulting in reduced accuracy and suboptimal detection performance.Moreover,HBBs cannot provide directional information for rotated objects.This study proposes a rotated detection method for identifying apparent defects in shield tunnels.Specifically,the oriented region-convolutional neural network(oriented R-CNN)is utilized to detect rotated objects in tunnel images.To enhance feature extraction,a novel hybrid backbone combining CNN-based networks with Swin Transformers is proposed.A feature fusion strategy is employed to integrate features extracted from both networks.Additionally,a neck network based on the bidirectional-feature pyramid network(Bi-FPN)is designed to combine multi-scale object features.The bolt hole dataset is curated to evaluate the efficacyof the proposed method.In addition,a dedicated pre-processing approach is developed for large-sized images to accommodate the rotated,dense,and small-scale characteristics of objects in tunnel images.Experimental results demonstrate that the proposed method achieves a more than 4%improvement in mAP50-95compared to other rotated detectors and a 6.6%-12.7%improvement over mainstream horizontal detectors.Furthermore,the proposed method outperforms mainstream methods by 6.5%-14.7%in detecting leakage bolt holes,underscoring its significant engineering applicability.
基金funded by the Program of Songshan Laboratory(241110210100)the National Natural Science Foundation of China(62301497)+1 种基金the Science and Technology Research Program of Henan(252102211024)the Key Research and Development Program of Henan(231111212000).
摘要The integrity of perception data transmitted over in-vehicle networks is important for the safety of autonomous driving.However,legacy protocols like the Controller Area Network(CAN)bus which lacks essential security features make In-Vehicle Networks(IVNs)vulnerable to data tampering attacks.Current research typically focuses on detecting the attack itself but ignores the information recovery from the missing data,leading to an unsafe autonomous driving system.To address the issue,we propose a 3D object recovery framework to recover the missing data caused by the tampering attack that occurred in in-vehicle networks.The proposed framework exploits both temporal and spatial context for the 3D object recovery,where a temporal branch is designed to learn the coordinate offsets of 3D objects based on historical data from previous frames,while a spatial branch employs information from the adjacent views of the attacked objects to locate the recovered objects from the overlapped regions in the current frame.By integrating the temporal and spatial clues,the framework effectively recovers the missing objects from the resting ones,thereby enhancing the immunity of in-vehicle networks for the tampering attack.Extensive experiments on the nuScenes dataset demonstrate that the proposed framework significantly improves 3D object detection performance under the attack when compared to the method without recovery.Additionally,the recovery performance becomes better as the attack intensity increases,highlighting the framework’s robustness in high-risk scenarios.The source will be available upon publication.
基金National Key Research and Development Program of China(2024YFD2400404)National Natural Science Foundation of China(62102243,42376194)Shanghai Sailing Program(21YF1417000)。
摘要[Objective]Detecting dense and small aquaculture net cages in complex backgrounds is difficult,the purpose of this study is to build a specialized dataset and design a targeted detection model that enhances recognition accuracy and robustness for practical aquaculture management.[Methods]A dataset of aquaculture net cages was constructed using highresolution remote sensing imagery collected from seven representative farming regions(Australia,Canada,Chile,Croatia,Greece,China,and the Faroe Islands),and Cage-YOLO,a deep learning model based on YOLOv5,was proposed for detecting dense and small aquaculture net cages.First,an adaptive dense perception algorithm was introduced,which automatically selects and generates feature maps that reflect the high-density distribution of small aquaculture net cages.Second,an enhanced module based on spatial pyramid pooling fast was integrated to effectively reduce background noise interference and improve global feature extraction capabilities.Finally,a mixed attention block was incorporated to further enhance the model's perception of dense and small objects.[Results and Discussions]Experimental results showed that the proposed Cage-YOLO achieved improvements over the original YOLOv5 in terms of precision,recall,and mean average precision by 5.6,21.8,and 17.4 percentage points,respectively.The model size was maintained at 16.9 MB,demonstrating both strong performance and deployment advantages.[Conclusions]This study provides a new approach for dense and small object detection and offers technical support for the intelligent management of marine cage aquaculture.
基金provided by the Science Research Project of Hebei Education Department under grant No.BJK2024115.
摘要High-resolution remote sensing images(HRSIs)are now an essential data source for gathering surface information due to advancements in remote sensing data capture technologies.However,their significant scale changes and wealth of spatial details pose challenges for semantic segmentation.While convolutional neural networks(CNNs)excel at capturing local features,they are limited in modeling long-range dependencies.Conversely,transformers utilize multihead self-attention to integrate global context effectively,but this approach often incurs a high computational cost.This paper proposes a global-local multiscale context network(GLMCNet)to extract both global and local multiscale contextual information from HRSIs.A detail-enhanced filtering module(DEFM)is proposed at the end of the encoder to refine the encoder outputs further,thereby enhancing the key details extracted by the encoder and effectively suppressing redundant information.In addition,a global-local multiscale transformer block(GLMTB)is proposed in the decoding stage to enable the modeling of rich multiscale global and local information.We also design a stair fusion mechanism to transmit deep semantic information from deep to shallow layers progressively.Finally,we propose the semantic awareness enhancement module(SAEM),which further enhances the representation of multiscale semantic features through spatial attention and covariance channel attention.Extensive ablation analyses and comparative experiments were conducted to evaluate the performance of the proposed method.Specifically,our method achieved a mean Intersection over Union(mIoU)of 86.89%on the ISPRS Potsdam dataset and 84.34%on the ISPRS Vaihingen dataset,outperforming existing models such as ABCNet and BANet.
基金supported by grants from the Hong Kong Research Grants Council(General Research Fund Grants No.14605920,14606922,14603724Collaborative Research Fund Grant No.C4023-20GF+9 种基金Research Matching Grants RMG 8601219,8601242,3110151)a grant from the Research Committee on Research Sustainability of Major Research Grants Council Funding Scheme(Grant No.3133235)of the Chinese University of Hong Kong(CUHK)grant from the 1+1+1 CUHK-CUHK(SZ)-GDSTC Joint Collaboration Fund(Grant No.4760974)grant from the Vice-Chancellor’s One-off Discretionary Fund(Smart and Sustainable Cities:City of Commons)(Grant No.4930787)of CUHKsupport from the Research Grants Council General Research Fund(Grant No.14618324)the Research Committee Direct Grant for Research(Grant No.4052336)the Strategic Partnership Award for Research Collaboration(Grant No.4750474)of the CUHKthe University Development Fund(Grant No.UDF01003932)from CUHK(SZ)grants from the“1+1+1”CUHK-CUHK(SZ)-GDSTC Joint Collaboration Fund(Grants No.2025A0505000083,2025A0505000062)supported by an RGC Postdoctoral Fellowship(Grant No.PDFS2324-4H04).
摘要Urban green space may impact human health through complex pathways and the effect can vary across different travel contexts.Revealing these disparities in health pathways between different travel contexts may provide essential and practical suggestions for sustainable developments in urban environments.In this study,we investi gated the impacts of travel contexts on people’s perceptions and evaluations of green space using a cross-sectional dataset collected in Hong Kong,China.Eight hundred participants in 4 representative communities were recruited through stratified sampling,and we identified 2,913 travel events from their two-day activity-travel diaries after rigorous cross-validation with GPS-derived trajectories.We also derived two green space exposure representa tions using fine-grained remote sensing imagery and 8 representative green space exposure indicators.Eighty logistical regression models and mixed-effects models were developed to investigate the associations with con trol of a range of potential uncertainties.Our results indicate solid and consistently positive associations between participants’measured green space exposure and perceived green space,and significant but variable effects of travel purposes,travel modes,and travel time on participants’perceptions and evaluations of green space.Walk ing significantly promotes participants’perceptions and positive evaluations of urban green space,buses are not significantly associated,and metro trains may depress the perception and evaluation.Our results provide solid evidence on how travel contexts may influence people’s perceptions and evaluations of urban green space and,thus,provide essential insights into environmental health studies and sustainable urban planning that consider green space as an important urban environmental setting.
基金supported and funded by Princess Nourah bint Abdulrahman University Researchers Supporting Project number(PNURSP2025R410),Princess Nourah bint Abdulrahman University,Riyadh,Saudi Arabia.
摘要Recognising human-object interactions(HOI)is a challenging task for traditional machine learning models,including convolutional neural networks(CNNs).Existing models show limited transferability across complex datasets such as D3D-HOI and SYSU 3D HOI.The conventional architecture of CNNs restricts their ability to handle HOI scenarios with high complexity.HOI recognition requires improved feature extraction methods to overcome the current limitations in accuracy and scalability.This work proposes a Novel quantum gate-enabled hybrid CNN(QEH-CNN)for effectiveHOI recognition.Themodel enhancesCNNperformance by integrating quantumcomputing components.The framework begins with bilateral image filtering,followed bymulti-object tracking(MOT)and Felzenszwalb superpixel segmentation.A watershed algorithm refines object boundaries by cleaning merged superpixels.Feature extraction combines a histogram of oriented gradients(HOG),Global Image Statistics for Texture(GIST)descriptors,and a novel 23-joint keypoint extractionmethod using relative joint angles and joint proximitymeasures.A fuzzy optimization process refines the extracted features before feeding them into the QEH-CNNmodel.The proposed model achieves 95.06%accuracy on the 3D-D3D-HOI dataset and 97.29%on the SYSU3DHOI dataset.Theintegration of quantum computing enhances feature optimization,leading to improved accuracy and overall model efficiency.
摘要In recent years,with the rapid advancement of artificial intelligence,object detection algorithms have made significant strides in accuracy and computational efficiency.Notably,research and applications of Anchor-Free models have opened new avenues for real-time target detection in optical remote sensing images(ORSIs).However,in the realmof adversarial attacks,developing adversarial techniques tailored to Anchor-Freemodels remains challenging.Adversarial examples generated based on Anchor-Based models often exhibit poor transferability to these new model architectures.Furthermore,the growing diversity of Anchor-Free models poses additional hurdles to achieving robust transferability of adversarial attacks.This study presents an improved cross-conv-block feature fusion You Only Look Once(YOLO)architecture,meticulously engineered to facilitate the extraction ofmore comprehensive semantic features during the backpropagation process.To address the asymmetry between densely distributed objects in ORSIs and the corresponding detector outputs,a novel dense bounding box attack strategy is proposed.This approach leverages dense target bounding boxes loss in the calculation of adversarial loss functions.Furthermore,by integrating translation-invariant(TI)and momentum-iteration(MI)adversarial methodologies,the proposed framework significantly improves the transferability of adversarial attacks.Experimental results demonstrate that our method achieves superior adversarial attack performance,with adversarial transferability rates(ATR)of 67.53%on the NWPU VHR-10 dataset and 90.71%on the HRSC2016 dataset.Compared to ensemble adversarial attack and cascaded adversarial attack approaches,our method generates adversarial examples in an average of 0.64 s,representing an approximately 14.5%improvement in efficiency under equivalent conditions.
基金Supported by National Natural Science Foundation of China (Grant No.92148206)National Key Research and Development Program of China (Grant No.2024YFC2418102)。
摘要Object tracking in 3D space is a classical problem in computer vision.In this paper,an efficient and robust X-Triplet detection method is proposed based on the support vector machine(SVM) and an adjacent matrix for locating and tracking objects through stereo vision with minimal feature points.The X-Triplet,denoted as Tri-X,is a composite marker consisting of three sequential X-corners.The definition and types of Tri-X markers are introduced at first.Then a fast and robust X-corner detector based on the block search strategy and SVM is proposed to extract X-corner candidates with sub-pixel locations and orientations.Thereafter the X-corner adjacent matrix(XAM) is constructed using the orientation angle error to describe the possibility that any X-corner pair form a valid edge vector.The Tri-X candidates are then extracted efficiently from the XAM.Finally once the Tri-X markers are detected in binocular images,their 6D pose information can be recovered through stereo matching and triangulation technique.When multiple targets are involved simultaneously,different Tri-X markers can be utilized to identify different objects.Experimental results show that the proposed method outperformed the state-of-the-art in terms of both accuracy and efficiency for Tri-X marker detection.In localization precision test,it achieved 0.1 mm error for the position and 1° error for the orientation.Our method exhibits great potential for utilization in user-defined specific tracking tasks,offering flexibility and adaptability to various tracking requirements,especially multi-tool tracking in medical robotics.
基金the National Natural Science Foundation of China(No.62271466)the Beijing Natural Science Foundation(No.4202025)+2 种基金the Beijing VanJee Technology Co.,Ltd.-Beijing Municipal Science and Technology Project(No.Z201100003920003)the Tianjin IoT Technology Enterprise Key Laboratory Research Project(No.VTJ-OT20230209-2)the Guizhou Provincial Sci-Tech Project(No.zk[2022]-012)。
摘要Intelligent transportation and autonomous driving systems have made urgent demands on the techniques with high performance on object detection in traffic scenes.This paper proposes an improved object detection model YOLO-VSF over the YOLOv4 model,which is a representative work with excellent performance among YOLO series of object detection models.The main improvement measures include:The backbone feature extraction network CSPDarknet53 of YOLOv4 is replaced with VGG16 to improve the feature extraction capability;SENet attention mechanism is incorporated to improve the salient and correlation feature representation capability;Focal Loss is integrated into the loss function to overcome the sample imbalance problem.In addition,the detection performance of small targets is improved by increasing the resolution of input images.Experimental results show that on the VanJee traffic image dataset provided by Beijing VanJee Technology Co.,Ltd.,the proposed YOLO-VSF model achieves an average mean accuracy(mAP)of 92.21 percentage points,which improves the mAP by 3.04 percentage points compared with the YOLOv4 model while maintaining the detection speed of the original model.On the UA-DETRAC dataset,the average accuracy of YOLO-VSF is close to that of the latest YOLOv7 model with the number of parameters reduced by 1.329×107.The proposed method can provide a support for object detection in traffic scenes.
基金funded by Beijing Natural Science Foundation,grant number L241078.
摘要The fasteners employed in the railway tracks are susceptible to defects arising from their intricate composition.Foreign objects are frequently observed on the track bed in an open environment.These two types of defects pose potential threats to high-speed trains,thus necessitating timely and accurate track inspection.The majority of extant automatic inspection methods are predicated on the utilization of single visible light data,and the efficacy of the algorithmic processes is influenced by complex environments.Furthermore,due to the single information dimension,the detection accuracy of defects in similar,occluded,and small object categories is low.To address the aforementioned issues,this paper proposes a track defect detectionmethod based on dynamicmulti-modal fusion and challenging object enhanced perception.First,in light of the variances in the representation dimensions ofmultimodal information,this paper proposes a dynamic weighted multi-modal feature fusion module.The fused multi-modal features are assigned weights,and thenmultiplied with the extracted single-modal features atmultiple levels,achieving adaptive adjustment of the response degree of fusion features.Second,a novel stepwise multi-scale convolution feature aggregation module is proposed for challenging objects.The proposed method employs depth separable convolution and cross-scale aggregation operations of different receptive fields to enhance feature extraction and reuse,thereby reducing the degree of progressive loss of effective information.The experimental results demonstrate the efficacy of the proposed method in comparison to eight established methods,encompassing both single-modal and multi-modal methods,as evidenced by the extensive findings within the constructed RGBD dataset.
基金Supported by The Natural Science Foundation of Fujian Province,No.2023J01122895.Institutional review board statement:This study was approved by the Ethics Committee of。
摘要BACKGROUND Recognition of pelvic autonomic nerves(PAN)during total mesorectal excision(TME)largely depends on the surgeon’s expertise,making it susceptible to misrecognition and unintentional damage.There is an urgent need for objective and real-time support methods.AIM To develop a deep learning(DL)model for precise recognition and visual annotation of 5 categories of PAN during TME.METHODS This single-center retrospective study enrolled 120 TME videos from January 2021 to January 2023.A total of 3246 high-quality images were obtained and split 9:1 into training and internal test sets.Difficult-to-recognize characteristics were summarized.An additional 20 independent TME videos from June 2023 to January 2024 were used for external validation.The DL model performance was compared with that of surgeons and verified pathologically.χ2,Fisher’s exact test and t-tests were applied(P<0.05).RESULTS The DL model achieved a precision of 0.839,a recall of 0.769,and a mean average precision at intersection over union 50 of 0.873.The overall recognition rate in external validation was 76.0%,similar to senior surgeons(73.9%,P=0.156)but superior to junior surgeons(64.9%,P=0.001).The miss rate of 5 PAN categories ranged from 12.9%to 29.6%.Initial recognition time(2.08-2.21 seconds)was shorter than that of senior surgeons(5.65-6.19 seconds,P<0.01);mean continuous tracking duration was prolonged by 57.21-66.45 seconds compared with that of senior surgeons(P<0.01).Low nerve exposure caused most DL model false negatives,while cord-like fibrous tissue dominated false positives.All 7 harvested specimens were pathologically confirmed to contain nerve tissue,with a processing speed of 25 frames per second.CONCLUSION The model demonstrates recognition performance comparable to that of senior surgeons,with pathological confirmation.It may potentially help preserve PAN during TME and shorten the learning curve for junior surgeons.