Driven by the high penetration of renewable energy,the inherent intermittency of photovoltaic(PV)generation poses severe challenges to grid stability.To manage this volatility and ensure reliable grid integration,prec...Driven by the high penetration of renewable energy,the inherent intermittency of photovoltaic(PV)generation poses severe challenges to grid stability.To manage this volatility and ensure reliable grid integration,precise PV system modeling and power forecasting have emerged as critical solutions.However,existing research predominantly focuses on algorithmic innovations and model architectures,frequently overlooking the foundational role of dataset selection.Because capturing the complex spatiotemporal dynamics of solar generation increasingly requires the integration of diverse data types,understanding how to select and fuse these multimodal sources is crucial for determining the upper bound of predictive performance.To address the persistent fragmentation of data resources in PV predictive modeling,this paper delivers a comprehensive taxonomy of publicly available benchmark datasets,establishing a roadmap for future data-driven research.We categorize these valuable resources into three core pillars:1)meteorological datasets(encompassing observational,synthetic,hybrid,and reanalysis types);2)PV generation datasets(grouped by temporal resolution);and 3)static system parameters(including plant-level geospatial data and module-level physical properties).Building upon this categorization,this review thoroughly examines multimodal data fusion strategies across various forecasting horizons and elucidates the specific data dependencies of persistence,physical,and data-driven modeling paradigms.Furthermore,we critically analyze key challenges in multi-source data fusion,particularly spatiotemporal misalignment and the lack of standardized quality control flags.Ultimately,this work provides researchers with an authoritative guide for robust data selection and model construction.展开更多
Parkinson’s disease(PD)is a debilitating neurological disorder affecting over 10 million people worldwide.PD classification models using voice signals as input are common in the literature.It is believed that using d...Parkinson’s disease(PD)is a debilitating neurological disorder affecting over 10 million people worldwide.PD classification models using voice signals as input are common in the literature.It is believed that using deep learning algorithms further enhances performance;nevertheless,it is challenging due to the nature of small-scale and imbalanced PD datasets.This paper proposed a convolutional neural network-based deep support vector machine(CNN-DSVM)to automate the feature extraction process using CNN and extend the conventional SVM to a DSVM for better classification performance in small-scale PD datasets.A customized kernel function reduces the impact of biased classification towards the majority class(healthy candidates in our consideration).An improved generative adversarial network(IGAN)was designed to generate additional training data to enhance the model’s performance.For performance evaluation,the proposed algorithm achieves a sensitivity of 97.6%and a specificity of 97.3%.The performance comparison is evaluated from five perspectives,including comparisons with different data generation algorithms,feature extraction techniques,kernel functions,and existing works.Results reveal the effectiveness of the IGAN algorithm,which improves the sensitivity and specificity by 4.05%–4.72%and 4.96%–5.86%,respectively;and the effectiveness of the CNN-DSVM algorithm,which improves the sensitivity by 1.24%–57.4%and specificity by 1.04%–163%and reduces biased detection towards the majority class.The ablation experiments confirm the effectiveness of individual components.Two future research directions have also been suggested.展开更多
Medical data has specificity compared to other fields of data,and the description of medical data characteristics is still in a qualitative stage.This study included 293 sub-datasets of 138 independent datasets.First,...Medical data has specificity compared to other fields of data,and the description of medical data characteristics is still in a qualitative stage.This study included 293 sub-datasets of 138 independent datasets.First,data preprocessing was performed using methods such as incomplete data removal,inconsistent data normalization,and data integration.Then,the characteristics of 293 research datasets were quantified using 26 indicators in three categories:simple indicators,statistical indicators,and informational indicators.Furthermore,statistical analysis was performed on the above-mentioned quantitative characteristics,and stepwise regression and decision tree methods were used for modeling learning.The characteristics of the biological and medical datasets in the study were compared with those of other fields’datasets.By comparing the results of statistical analysis and learning modeling,the study found that the sample size of medical datasets included in the UCI database analyzed in this paper is small,most within 1000.The harmonic mean or geometric mean of continuous variables is significantly higher than the data from other fields.That is to say,the scope of the continuous variable range is large.This study uses quantitative indicators to describe the characteristics of medical datasets to avoid the decrease in credibility caused by subjective analysis,and lays a foundation for further algorithm applicability research.展开更多
Accurate purchase prediction in e-commerce critically depends on the quality of behavioral features.This paper proposes a layered and interpretable feature engineering framework that organizes user signals into three ...Accurate purchase prediction in e-commerce critically depends on the quality of behavioral features.This paper proposes a layered and interpretable feature engineering framework that organizes user signals into three layers:Basic,Conversion&Stability(efficiency and volatility across actions),and Advanced Interactions&Activity(crossbehavior synergies and intensity).Using real Taobao(Alibaba’s primary e-commerce platform)logs(57,976 records for 10,203 users;25 November–03 December 2017),we conducted a hierarchical,layer-wise evaluation that holds data splits and hyperparameters fixed while varying only the feature set to quantify each layer’s marginal contribution.Across logistic regression(LR),decision tree,random forest,XGBoost,and CatBoost models with stratified 5-fold cross-validation,the performance improvedmonotonically fromBasic to Conversion&Stability to Advanced features.With LR,F1 increased from 0.613(Basic)to 0.962(Advanced);boosted models achieved high discrimination(0.995 AUC Score)and an F1 score up to 0.983.Calibration and precision–recall analyses indicated strong ranking quality and acknowledged potential dataset and period biases given the short(9-day)window.By making feature contributions measurable and reproducible,the framework complements model-centric advances and offers a transparent blueprint for production-grade behavioralmodeling.The code and processed artifacts are publicly available,and future work will extend the validation to longer,seasonal datasets and hybrid approaches that combine automated feature learning with domain-driven design.展开更多
Objective To investigate methods for constructing a high-quality instructional dataset for traditional Chinese medicine(TCM)mental disorders and to validate its efficacy.Methods We proposed the Fine-Med-Mental-T&P...Objective To investigate methods for constructing a high-quality instructional dataset for traditional Chinese medicine(TCM)mental disorders and to validate its efficacy.Methods We proposed the Fine-Med-Mental-T&P methodology for constructing high-quality instruction datasets in TCM mental disorders.This approach integrates theoretical knowledge and practical case studies through a dual-track strategy.(i)Theoretical track:textbooks and guidelines on TCM mental disorders were manually segmented.Initial responses were generated using DeepSeek-V3,followed by refinement by the Qwen3-32B model to align the expression with human preferences.A screening algorithm was then applied to select 16000 high-quality instruction pairs.(ii)Practical track:starting from over 600 real clinical case seeds,diagnostic and therapeutic instruction pairs were generated using DeepSeek-V3 and subsequently screened through manual evaluation,resulting in 4000 high-quality practiceoriented instruction pairs.The integration of both tracks yielded the Med-Mental-Instruct-T&P dataset,comprising a total of 20000 instruction pairs.To validate the dataset’s effectiveness,three experimental evaluations(both manual and automated)were conducted:(i)comparative studies to compare the performance of models fine-tuned on different datasets;(ii)benchmarking to compare against mainstream TCM-specific large language models(LLMs);(iii)data ablation study to investigate the relationship between data volume and model performance.Results Experimental results demonstrate the superior performance of T&P-model finetuned on the Med-Mental-Instruct-T&P dataset.In the comparative study,the T&P-model significantly outperformed the baseline models trained solely on self-generated or purely human-curated baseline data.This superiority was evident in both automated metrics(ROUGEL>0.55)and expert manual evaluations(scoring above 7/10 across accuracy).In benchmark comparisons,the T&P-model also excelled against existing mainstream TCM LLMs(e.g.,HuatuoGPT and ZuoyiGPT).It showed particularly strong capabilities in handling diverse clinical presentations,including challenging disorders such as insomnia and coma,showcasing its robustness and versatility.Data ablation studies showed that T&P-model performance had an overall upward trend with minor fluctuations when training data increased from 10%to 50%;beyond 50%,performance improvement slowed significantly,with metrics plateauing and approaching a saturation point.展开更多
This study introduces stacked deep learning for multilingual opinion mining.This framework incorporates RoBERTa-GRU,RoBERTa-LSTM,RoBERTa-BiGRU,and RoBERTa-BiLSTM hybrid models,optimized using the Adam optimizer.The me...This study introduces stacked deep learning for multilingual opinion mining.This framework incorporates RoBERTa-GRU,RoBERTa-LSTM,RoBERTa-BiGRU,and RoBERTa-BiLSTM hybrid models,optimized using the Adam optimizer.The methodology can handle three languages:French,English,and Arabic.This research explores the challenges of imbalanced datasets in opinion mining,employing oversampling approaches like advanced easier data augmentation,synthetic minority over-sampling technique,and generative pre-trained transformer to balance the datasets and improve classification efficiency.We used the assessment metrics of Cohen’s kappa,the receiver operating characteristic-area under curve,accuracy,and Matthews correlation coefficient,along with k-fold validation,to evaluate the sentiment analysis performance across three languages and six datasets.Moreover,we computed performance metrics for all models while scaling the dataset for training and testing.We also determined the memory usage and execution time for each model.Utilizing a stacked deep learning algorithm,the suggested methodology for multilingual opinion mining demonstrated high efficacy in extracting meaningful insights from social media data across several languages.The technique produced significant outcomes,rendering it a potentially valuable instrument for enhancing performance and customer satisfaction by identifying patterns and trends in public sentiment.展开更多
Liquid-based cytology(LBC)has become a core technology in cervical cancer screening,and artificial intelligence(AI)shows great potential in addressing issues such as the global shortage of cytopathologists and large v...Liquid-based cytology(LBC)has become a core technology in cervical cancer screening,and artificial intelligence(AI)shows great potential in addressing issues such as the global shortage of cytopathologists and large variations in diagnostic results.However,the clinical reliability of AI systems in the field of cervical cytology fundamentally depends on the quality,diversity,and representativeness of their training and validation datasets.Drawing on evidence-based medicine principles,the latest research progress in China and globally,and clinical practices,these guidelines establish standardized requirements for sample diversity and data sufficiency in cervical LBC AI datasets to reduce algorithmic bias.The objective is to improve the real-world applicability of models and ensure their safe clinical application.The guidelines strictly comply with standardized guideline development specifications,with transparent expert panel organization,systematic literature retrieval,standardized evidence grading,three-round Delphi expert consensus,external peer review,and dynamic update mechanisms.All quantitative threshold indicators in the recommendations are jointly formulated based on highquality clinical evidence and expert consensus,with clear evidence sources and consensus construction processes to enhance the transparency,credibility,and operability of the guideline.展开更多
Standardized datasets are foundational to healthcare informatization by enhancing data quality and unleashing the value of data elements.Using bibliometrics and content analysis,this study examines China's healthc...Standardized datasets are foundational to healthcare informatization by enhancing data quality and unleashing the value of data elements.Using bibliometrics and content analysis,this study examines China's healthcare dataset standards from 2011 to 2025.It analyzes their evolution across types,applications,institutions,and themes,highlighting key achievements including substantial growth in quantity,optimized typology,expansion into innovative application scenarios such as health decision support,and broadened institutional involvement.The study also identifies critical challenges,including imbalanced development,insufficient quality control,and a lack of essential metadata—such as authoritative data element mappings and privacy annotations—which hampers the delivery of intelligent services.To address these challenges,the study proposes a multi-faceted strategy focused on optimizing the standard system's architecture,enhancing quality and implementation,and advancing both data governance—through authoritative tracing and privacy protection—and intelligent service provision.These strategies aim to promote the application of dataset standards,thereby fostering and securing the development of new productive forces in healthcare.展开更多
Detecting faces under occlusion remains a significant challenge in computer vision due to variations caused by masks,sunglasses,and other obstructions.Addressing this issue is crucial for applications such as surveill...Detecting faces under occlusion remains a significant challenge in computer vision due to variations caused by masks,sunglasses,and other obstructions.Addressing this issue is crucial for applications such as surveillance,biometric authentication,and human-computer interaction.This paper provides a comprehensive review of face detection techniques developed to handle occluded faces.Studies are categorized into four main approaches:feature-based,machine learning-based,deep learning-based,and hybrid methods.We analyzed state-of-the-art studies within each category,examining their methodologies,strengths,and limitations based on widely used benchmark datasets,highlighting their adaptability to partial and severe occlusions.The review also identifies key challenges,including dataset diversity,model generalization,and computational efficiency.Our findings reveal that deep learning methods dominate recent studies,benefiting from their ability to extract hierarchical features and handle complex occlusion patterns.More recently,researchers have increasingly explored Transformer-based architectures,such as Vision Transformer(ViT)and Swin Transformer,to further improve detection robustness under challenging occlusion scenarios.In addition,hybrid approaches,which aim to combine traditional andmodern techniques,are emerging as a promising direction for improving robustness.This review provides valuable insights for researchers aiming to develop more robust face detection systems and for practitioners seeking to deploy reliable solutions in real-world,occlusionprone environments.Further improvements and the proposal of broader datasets are required to developmore scalable,robust,and efficient models that can handle complex occlusions in real-world scenarios.展开更多
The aim of this article is to explore potential directions for the development of artificial intelligence(AI).It points out that,while current AI can handle the statistical properties of complex systems,it has difficu...The aim of this article is to explore potential directions for the development of artificial intelligence(AI).It points out that,while current AI can handle the statistical properties of complex systems,it has difficulty effectively processing and fully representing their spatiotemporal complexity patterns.The article also discusses a potential path of AI development in the engineering domain.Based on the existing understanding of the principles of multilevel com-plexity,this article suggests that consistency among the logical structures of datasets,AI models,model-building software,and hardware will be an important AI development direction and is worthy of careful consideration.展开更多
Inferring phylogenetic trees from molecular sequences is a cornerstone of evolutionary biology.Many standard phylogenetic methods(such as maximum-likelihood[ML])rely on explicit models of sequence evolution and thus o...Inferring phylogenetic trees from molecular sequences is a cornerstone of evolutionary biology.Many standard phylogenetic methods(such as maximum-likelihood[ML])rely on explicit models of sequence evolution and thus often suffer from model misspecification or inadequacy.The on-rising deep learning(DL)techniques offer a powerful alternative.Deep learning employs multi-layered artificial neural networks to progressively transform input data into more abstract and complex representations.DL methods can autonomously uncover meaningful patterns from data,thereby bypassing potential biases introduced by predefined features(Franklin,2005;Murphy,2012).Recent efforts have aimed to apply deep neural networks(DNNs)to phylogenetics,with a growing number of applications in tree reconstruction(Suvorov et al.,2020;Zou et al.,2020;Nesterenko et al.,2022;Smith and Hahn,2023;Wang et al.,2023),substitution model selection(Abadi et al.,2020;Burgstaller-Muehlbacher et al.,2023),and diversification rate inference(Voznica et al.,2022;Lajaaiti et al.,2023;Lambert et al.,2023).In phylogenetic tree reconstruction,PhyDL(Zou et al.,2020)and Tree_learning(Suvorov et al.,2020)are two notable DNN-based programs designed to infer unrooted quartet trees directly from alignments of four amino acid(AA)and DNA sequences,respectively.展开更多
When dealing with imbalanced datasets,the traditional support vectormachine(SVM)tends to produce a classification hyperplane that is biased towards the majority class,which exhibits poor robustness.This paper proposes...When dealing with imbalanced datasets,the traditional support vectormachine(SVM)tends to produce a classification hyperplane that is biased towards the majority class,which exhibits poor robustness.This paper proposes a high-performance classification algorithm specifically designed for imbalanced datasets.The proposed method first uses a biased second-order cone programming support vectormachine(B-SOCP-SVM)to identify the support vectors(SVs)and non-support vectors(NSVs)in the imbalanced data.Then,it applies the synthetic minority over-sampling technique(SV-SMOTE)to oversample the support vectors of the minority class and uses the random under-sampling technique(NSV-RUS)multiple times to undersample the non-support vectors of the majority class.Combining the above-obtained minority class data set withmultiple majority class datasets can obtainmultiple new balanced data sets.Finally,SOCP-SVM is used to classify each data set,and the final result is obtained through the integrated algorithm.Experimental results demonstrate that the proposed method performs excellently on imbalanced datasets.展开更多
Data-driven autonomous driving is a hot topic in academic and industry research due to its impressive performance,flexible mobility,and reduced human intervention.However,the development of this technology relies heav...Data-driven autonomous driving is a hot topic in academic and industry research due to its impressive performance,flexible mobility,and reduced human intervention.However,the development of this technology relies heavily on large datasets that contain accurately annotated data,obtained through artificial or semi-automated strategies.Consequently,datasets play a crucial role in autonomous driving,and their characteristics significantly impact the effectiveness of algorithms.Currently,there are several diverse datasets available,such as KITTI and City Scape,that cover various tasks.However,researchers often overlook the unique features,similarities,and specificities of these datasets.Furthermore,to the best of our knowledge,there is a lack of survey articles focusing on special metrics and benchmark performance on different datasets in autonomous driving.Therefore,the purpose of this article is to analyze autonomous driving datasets,guide researchers on collecting and utilizing relevant datasets,summarize evaluation strategies,analyze benchmark performance,and provide future research points to enrich the autonomous driving community.We believe that this work will assist researchers in evaluating their data using suitable metrics and offer a fresh perspective on autonomous driving.展开更多
This paper presents a systematic survey of machine vision-based surface defect detection technologies,focusing on five core challenges in the field:interference from complex backgrounds,small object detection,class im...This paper presents a systematic survey of machine vision-based surface defect detection technologies,focusing on five core challenges in the field:interference from complex backgrounds,small object detection,class imbalance,dynamic scene modeling,and cross-scenario generalization.It reviews key technical approaches corresponding to these challenges over the past five years.Furthermore,a dataset characterization analysis framework is established around these challenges,summarizing and comparing the characteristics of over 40 publicly available datasets across more than ten scenarios,including PCB,photovoltaic,metal,and pavement surfaces.Quantitative selection metrics(such as the small target coefficient and texture complexity)are proposed for challenges like small target detection and complex backgrounds,offering a methodological guide for aligning research questions with benchmark data.Finally,the paper summarizes current limitations and provides an outlook on new paradigms driven by large-scale models and the construction of high-quality benchmark datasets,aiming to offer valuable references for both research and engineering practices in this field.展开更多
Small datasets are often challenging due to their limited sample size.This research introduces a novel solution to these problems:average linkage virtual sample generation(ALVSG).ALVSG leverages the underlying data st...Small datasets are often challenging due to their limited sample size.This research introduces a novel solution to these problems:average linkage virtual sample generation(ALVSG).ALVSG leverages the underlying data structure to create virtual samples,which can be used to augment the original dataset.The ALVSG process consists of two steps.First,an average-linkage clustering technique is applied to the dataset to create a dendrogram.The dendrogram represents the hierarchical structure of the dataset,with each merging operation regarded as a linkage.Next,the linkages are combined into an average-based dataset,which serves as a new representation of the dataset.The second step in the ALVSG process involves generating virtual samples using the average-based dataset.The research project generates a set of 100 virtual samples by uniformly distributing them within the provided boundary.These virtual samples are then added to the original dataset,creating a more extensive dataset with improved generalization performance.The efficacy of the ALVSG approach is validated through resampling experiments and t-tests conducted on two small real-world datasets.The experiments are conducted on three forecasting models:the support vector machine for regression(SVR),the deep learning model(DL),and XGBoost.The results show that the ALVSG approach outperforms the baseline methods in terms of mean square error(MSE),root mean square error(RMSE),and mean absolute error(MAE).展开更多
With the deep integration of smart manufacturing and IoT technologies,higher demands are placed on the intelligence and real-time performance of industrial equipment fault detection.For industrial fans,base bolt loose...With the deep integration of smart manufacturing and IoT technologies,higher demands are placed on the intelligence and real-time performance of industrial equipment fault detection.For industrial fans,base bolt loosening faults are difficult to identify through conventional spectrum analysis,and the extreme scarcity of fault data leads to limited training datasets,making traditional deep learning methods inaccurate in fault identification and incapable of detecting loosening severity.This paper employs Bayesian Learning by training on a small fault dataset collected from the actual operation of axial-flow fans in a factory to obtain posterior distribution.This method proposes specific data processing approaches and a configuration of Bayesian Convolutional Neural Network(BCNN).It can effectively improve the model’s generalization ability.Experimental results demonstrate high detection accuracy and alignment with real-world applications,offering practical significance and reference value for industrial fan bolt loosening detection under data-limited conditions.展开更多
Monitoring concrete cracks for structural health in civil engineering presents a significant challenge.This is primarily due to the reliance on manual investigation methods,impacts of global climatic shifts stress,and...Monitoring concrete cracks for structural health in civil engineering presents a significant challenge.This is primarily due to the reliance on manual investigation methods,impacts of global climatic shifts stress,and geohazard threats to engineering structures.To cope with this challenge,state-of-the-art Deep Learning(DL)models are utilized to predict concrete cracks and accurately identify subtle variations in crack patterns and sizes,which lighting conditions and surface textures can influence.Previous studies indicate that model accuracy may decrease when faced with obscured concrete cracks,irregular shapes,or limited datasets for real-world problem scenarios.Feature fusion enhances model performance by combining complementary information,resulting in more accurate predictions,but may increase complexity and potential information redundancy.The study presents the Fractur Encoder to Decoder(FractED)block,a novel architecture consisting of three sub-blocks:the inner block(Encoder),intermediate block(Intermediate block),and outer block(Decoder).This approach integrates fused features into the model without additional fine-tuning steps,allowing for comprehensive feature refinement and enhancement,ultimately optimizing model performance.The study investigates a DL methodology on three datasets,demonstrating its effectiveness in handling complex classification scenarios in civil engineering.The model achieved high accuracy rates,with 88.41%for multiclass(Deck,Pavement,and Walls)classification tasks,91.94%on the Pillow Dam Borehole image binary dataset,and 99.77%on the Surface Crack binary dataset.The FractED block integration ensures adaptability and scalability,making it valuable for various Artificial Intelligence(AI)applications in civil engineering.The research also provides a scientific foundation for automatizing civil engineering inspection instruments for the future.展开更多
With the continuous improvement of the performance of large language models,how to further enhance their ability in complex tasks has become a key issue.The task of abnormal text detection poses a challenge to the mod...With the continuous improvement of the performance of large language models,how to further enhance their ability in complex tasks has become a key issue.The task of abnormal text detection poses a challenge to the model in identifying non-standard semantics due to its semantic complexity and high-risk features.However,existing fine-tuning methods rely heavily on static data selection strategies,making it difficult to adapt to the dynamic evolution of model capabilities,resulting in low training efficiency.This article proposes ADS(Adaptive Dataset Selection),an adaptive framework for selecting data in anomaly text detection.ADS performs model-aware data selection prior to fine-tuning,adapting the initial state of pre-trained language models by selecting samples that are most informative for the target anomaly detection task.Empirical results on mainstream large language model architectures show that ADS significantly compresses data size while still outperforming existing static strategies and mainstream compression methods.When using only 1000 fine-tuning samples,ADS achieves a 92%F1 score,with an accuracy improvement of over 22%compared to the baseline,demonstrating excellent performance.This study proposes an efficient data selection mechanism from the perspective of model capability and dynamic adaptation of data,providing theoretical support and a practical path for fine-tuning large models in low-resource scenarios.展开更多
Climate change significantly affects environment,ecosystems,communities,and economies.These impacts often result in quick and gradual changes in water resources,environmental conditions,and weather patterns.A geograph...Climate change significantly affects environment,ecosystems,communities,and economies.These impacts often result in quick and gradual changes in water resources,environmental conditions,and weather patterns.A geographical study was conducted in Arizona State,USA,to examine monthly precipi-tation concentration rates over time.This analysis used a high-resolution 0.50×0.50 grid for monthly precip-itation data from 1961 to 2022,Provided by the Climatic Research Unit.The study aimed to analyze climatic changes affected the first and last five years of each decade,as well as the entire decade,during the specified period.GIS was used to meet the objectives of this study.Arizona experienced 51–568 mm,67–560 mm,63–622 mm,and 52–590 mm of rainfall in the sixth,seventh,eighth,and ninth decades of the second millennium,respectively.Both the first and second five year periods of each decade showed accept-able rainfall amounts despite fluctuations.However,rainfall decreased in the first and second decades of the third millennium.and in the first two years of the third decade.Rainfall amounts dropped to 42–472 mm,55–469 mm,and 74–498 mm,respectively,indicating a downward trend in precipitation.The central part of the state received the highest rainfall,while the eastern and western regions(spanning north to south)had significantly less.Over the decades of the third millennium,the average annual rainfall every five years was relatively low,showing a declining trend due to severe climate changes,generally ranging between 35 mm and 498 mm.The central regions consistently received more rainfall than the eastern and western outskirts.Arizona is currently experiencing a decrease in rainfall due to climate change,a situation that could deterio-rate further.This highlights the need to optimize the use of existing rainfall and explore alternative water sources.展开更多
基金supported by Smart Grid-National Science and Technology Major Project(No.2025ZD0803600,2025ZD0803601)National Natural Science Foundation of China(No.52307133)+1 种基金Tianjin Metrology Science and Technology Project(No.2024TJMT028)Tianjin Transportation Technology Project(No.2025-76).
摘要Driven by the high penetration of renewable energy,the inherent intermittency of photovoltaic(PV)generation poses severe challenges to grid stability.To manage this volatility and ensure reliable grid integration,precise PV system modeling and power forecasting have emerged as critical solutions.However,existing research predominantly focuses on algorithmic innovations and model architectures,frequently overlooking the foundational role of dataset selection.Because capturing the complex spatiotemporal dynamics of solar generation increasingly requires the integration of diverse data types,understanding how to select and fuse these multimodal sources is crucial for determining the upper bound of predictive performance.To address the persistent fragmentation of data resources in PV predictive modeling,this paper delivers a comprehensive taxonomy of publicly available benchmark datasets,establishing a roadmap for future data-driven research.We categorize these valuable resources into three core pillars:1)meteorological datasets(encompassing observational,synthetic,hybrid,and reanalysis types);2)PV generation datasets(grouped by temporal resolution);and 3)static system parameters(including plant-level geospatial data and module-level physical properties).Building upon this categorization,this review thoroughly examines multimodal data fusion strategies across various forecasting horizons and elucidates the specific data dependencies of persistence,physical,and data-driven modeling paradigms.Furthermore,we critically analyze key challenges in multi-source data fusion,particularly spatiotemporal misalignment and the lack of standardized quality control flags.Ultimately,this work provides researchers with an authoritative guide for robust data selection and model construction.
基金The work described in this paper was fully supported by a grant from Hong Kong Metropolitan University(RIF/2021/05).
摘要Parkinson’s disease(PD)is a debilitating neurological disorder affecting over 10 million people worldwide.PD classification models using voice signals as input are common in the literature.It is believed that using deep learning algorithms further enhances performance;nevertheless,it is challenging due to the nature of small-scale and imbalanced PD datasets.This paper proposed a convolutional neural network-based deep support vector machine(CNN-DSVM)to automate the feature extraction process using CNN and extend the conventional SVM to a DSVM for better classification performance in small-scale PD datasets.A customized kernel function reduces the impact of biased classification towards the majority class(healthy candidates in our consideration).An improved generative adversarial network(IGAN)was designed to generate additional training data to enhance the model’s performance.For performance evaluation,the proposed algorithm achieves a sensitivity of 97.6%and a specificity of 97.3%.The performance comparison is evaluated from five perspectives,including comparisons with different data generation algorithms,feature extraction techniques,kernel functions,and existing works.Results reveal the effectiveness of the IGAN algorithm,which improves the sensitivity and specificity by 4.05%–4.72%and 4.96%–5.86%,respectively;and the effectiveness of the CNN-DSVM algorithm,which improves the sensitivity by 1.24%–57.4%and specificity by 1.04%–163%and reduces biased detection towards the majority class.The ablation experiments confirm the effectiveness of individual components.Two future research directions have also been suggested.
基金funded by the Qingdao Huanghai University Doctoral Research Foundation Project,grant number 2023boshi02,and Qingdao Huanghai University scientific research project,grant number KYH2025001.
摘要Medical data has specificity compared to other fields of data,and the description of medical data characteristics is still in a qualitative stage.This study included 293 sub-datasets of 138 independent datasets.First,data preprocessing was performed using methods such as incomplete data removal,inconsistent data normalization,and data integration.Then,the characteristics of 293 research datasets were quantified using 26 indicators in three categories:simple indicators,statistical indicators,and informational indicators.Furthermore,statistical analysis was performed on the above-mentioned quantitative characteristics,and stepwise regression and decision tree methods were used for modeling learning.The characteristics of the biological and medical datasets in the study were compared with those of other fields’datasets.By comparing the results of statistical analysis and learning modeling,the study found that the sample size of medical datasets included in the UCI database analyzed in this paper is small,most within 1000.The harmonic mean or geometric mean of continuous variables is significantly higher than the data from other fields.That is to say,the scope of the continuous variable range is large.This study uses quantitative indicators to describe the characteristics of medical datasets to avoid the decrease in credibility caused by subjective analysis,and lays a foundation for further algorithm applicability research.
基金supported by the research fund of Hanyang University(HY-202500000001616).
摘要Accurate purchase prediction in e-commerce critically depends on the quality of behavioral features.This paper proposes a layered and interpretable feature engineering framework that organizes user signals into three layers:Basic,Conversion&Stability(efficiency and volatility across actions),and Advanced Interactions&Activity(crossbehavior synergies and intensity).Using real Taobao(Alibaba’s primary e-commerce platform)logs(57,976 records for 10,203 users;25 November–03 December 2017),we conducted a hierarchical,layer-wise evaluation that holds data splits and hyperparameters fixed while varying only the feature set to quantify each layer’s marginal contribution.Across logistic regression(LR),decision tree,random forest,XGBoost,and CatBoost models with stratified 5-fold cross-validation,the performance improvedmonotonically fromBasic to Conversion&Stability to Advanced features.With LR,F1 increased from 0.613(Basic)to 0.962(Advanced);boosted models achieved high discrimination(0.995 AUC Score)and an F1 score up to 0.983.Calibration and precision–recall analyses indicated strong ranking quality and acknowledged potential dataset and period biases given the short(9-day)window.By making feature contributions measurable and reproducible,the framework complements model-centric advances and offers a transparent blueprint for production-grade behavioralmodeling.The code and processed artifacts are publicly available,and future work will extend the validation to longer,seasonal datasets and hybrid approaches that combine automated feature learning with domain-driven design.
基金Key Scientific Research Project of the Hunan Provincial Department of Education(23A312).
摘要Objective To investigate methods for constructing a high-quality instructional dataset for traditional Chinese medicine(TCM)mental disorders and to validate its efficacy.Methods We proposed the Fine-Med-Mental-T&P methodology for constructing high-quality instruction datasets in TCM mental disorders.This approach integrates theoretical knowledge and practical case studies through a dual-track strategy.(i)Theoretical track:textbooks and guidelines on TCM mental disorders were manually segmented.Initial responses were generated using DeepSeek-V3,followed by refinement by the Qwen3-32B model to align the expression with human preferences.A screening algorithm was then applied to select 16000 high-quality instruction pairs.(ii)Practical track:starting from over 600 real clinical case seeds,diagnostic and therapeutic instruction pairs were generated using DeepSeek-V3 and subsequently screened through manual evaluation,resulting in 4000 high-quality practiceoriented instruction pairs.The integration of both tracks yielded the Med-Mental-Instruct-T&P dataset,comprising a total of 20000 instruction pairs.To validate the dataset’s effectiveness,three experimental evaluations(both manual and automated)were conducted:(i)comparative studies to compare the performance of models fine-tuned on different datasets;(ii)benchmarking to compare against mainstream TCM-specific large language models(LLMs);(iii)data ablation study to investigate the relationship between data volume and model performance.Results Experimental results demonstrate the superior performance of T&P-model finetuned on the Med-Mental-Instruct-T&P dataset.In the comparative study,the T&P-model significantly outperformed the baseline models trained solely on self-generated or purely human-curated baseline data.This superiority was evident in both automated metrics(ROUGEL>0.55)and expert manual evaluations(scoring above 7/10 across accuracy).In benchmark comparisons,the T&P-model also excelled against existing mainstream TCM LLMs(e.g.,HuatuoGPT and ZuoyiGPT).It showed particularly strong capabilities in handling diverse clinical presentations,including challenging disorders such as insomnia and coma,showcasing its robustness and versatility.Data ablation studies showed that T&P-model performance had an overall upward trend with minor fluctuations when training data increased from 10%to 50%;beyond 50%,performance improvement slowed significantly,with metrics plateauing and approaching a saturation point.
摘要This study introduces stacked deep learning for multilingual opinion mining.This framework incorporates RoBERTa-GRU,RoBERTa-LSTM,RoBERTa-BiGRU,and RoBERTa-BiLSTM hybrid models,optimized using the Adam optimizer.The methodology can handle three languages:French,English,and Arabic.This research explores the challenges of imbalanced datasets in opinion mining,employing oversampling approaches like advanced easier data augmentation,synthetic minority over-sampling technique,and generative pre-trained transformer to balance the datasets and improve classification efficiency.We used the assessment metrics of Cohen’s kappa,the receiver operating characteristic-area under curve,accuracy,and Matthews correlation coefficient,along with k-fold validation,to evaluate the sentiment analysis performance across three languages and six datasets.Moreover,we computed performance metrics for all models while scaling the dataset for training and testing.We also determined the memory usage and execution time for each model.Utilizing a stacked deep learning algorithm,the suggested methodology for multilingual opinion mining demonstrated high efficacy in extracting meaningful insights from social media data across several languages.The technique produced significant outcomes,rendering it a potentially valuable instrument for enhancing performance and customer satisfaction by identifying patterns and trends in public sentiment.
摘要Liquid-based cytology(LBC)has become a core technology in cervical cancer screening,and artificial intelligence(AI)shows great potential in addressing issues such as the global shortage of cytopathologists and large variations in diagnostic results.However,the clinical reliability of AI systems in the field of cervical cytology fundamentally depends on the quality,diversity,and representativeness of their training and validation datasets.Drawing on evidence-based medicine principles,the latest research progress in China and globally,and clinical practices,these guidelines establish standardized requirements for sample diversity and data sufficiency in cervical LBC AI datasets to reduce algorithmic bias.The objective is to improve the real-world applicability of models and ensure their safe clinical application.The guidelines strictly comply with standardized guideline development specifications,with transparent expert panel organization,systematic literature retrieval,standardized evidence grading,three-round Delphi expert consensus,external peer review,and dynamic update mechanisms.All quantitative threshold indicators in the recommendations are jointly formulated based on highquality clinical evidence and expert consensus,with clear evidence sources and consensus construction processes to enhance the transparency,credibility,and operability of the guideline.
摘要Standardized datasets are foundational to healthcare informatization by enhancing data quality and unleashing the value of data elements.Using bibliometrics and content analysis,this study examines China's healthcare dataset standards from 2011 to 2025.It analyzes their evolution across types,applications,institutions,and themes,highlighting key achievements including substantial growth in quantity,optimized typology,expansion into innovative application scenarios such as health decision support,and broadened institutional involvement.The study also identifies critical challenges,including imbalanced development,insufficient quality control,and a lack of essential metadata—such as authoritative data element mappings and privacy annotations—which hampers the delivery of intelligent services.To address these challenges,the study proposes a multi-faceted strategy focused on optimizing the standard system's architecture,enhancing quality and implementation,and advancing both data governance—through authoritative tracing and privacy protection—and intelligent service provision.These strategies aim to promote the application of dataset standards,thereby fostering and securing the development of new productive forces in healthcare.
基金funded by A’Sharqiyah University,Sultanate of Oman,under Research Project grant number(BFP/RGP/ICT/22/490).
摘要Detecting faces under occlusion remains a significant challenge in computer vision due to variations caused by masks,sunglasses,and other obstructions.Addressing this issue is crucial for applications such as surveillance,biometric authentication,and human-computer interaction.This paper provides a comprehensive review of face detection techniques developed to handle occluded faces.Studies are categorized into four main approaches:feature-based,machine learning-based,deep learning-based,and hybrid methods.We analyzed state-of-the-art studies within each category,examining their methodologies,strengths,and limitations based on widely used benchmark datasets,highlighting their adaptability to partial and severe occlusions.The review also identifies key challenges,including dataset diversity,model generalization,and computational efficiency.Our findings reveal that deep learning methods dominate recent studies,benefiting from their ability to extract hierarchical features and handle complex occlusion patterns.More recently,researchers have increasingly explored Transformer-based architectures,such as Vision Transformer(ViT)and Swin Transformer,to further improve detection robustness under challenging occlusion scenarios.In addition,hybrid approaches,which aim to combine traditional andmodern techniques,are emerging as a promising direction for improving robustness.This review provides valuable insights for researchers aiming to develop more robust face detection systems and for practitioners seeking to deploy reliable solutions in real-world,occlusionprone environments.Further improvements and the proposal of broader datasets are required to developmore scalable,robust,and efficient models that can handle complex occlusions in real-world scenarios.
摘要The aim of this article is to explore potential directions for the development of artificial intelligence(AI).It points out that,while current AI can handle the statistical properties of complex systems,it has difficulty effectively processing and fully representing their spatiotemporal complexity patterns.The article also discusses a potential path of AI development in the engineering domain.Based on the existing understanding of the principles of multilevel com-plexity,this article suggests that consistency among the logical structures of datasets,AI models,model-building software,and hardware will be an important AI development direction and is worthy of careful consideration.
基金supported by the National Key R&D Program of China(2022YFD1401600)the National Science Foundation for Distinguished Young Scholars of Zhejang Province,China(LR23C140001)supported by the Key Area Research and Development Program of Guangdong Province,China(2018B020205003 and 2020B0202090001).
摘要Inferring phylogenetic trees from molecular sequences is a cornerstone of evolutionary biology.Many standard phylogenetic methods(such as maximum-likelihood[ML])rely on explicit models of sequence evolution and thus often suffer from model misspecification or inadequacy.The on-rising deep learning(DL)techniques offer a powerful alternative.Deep learning employs multi-layered artificial neural networks to progressively transform input data into more abstract and complex representations.DL methods can autonomously uncover meaningful patterns from data,thereby bypassing potential biases introduced by predefined features(Franklin,2005;Murphy,2012).Recent efforts have aimed to apply deep neural networks(DNNs)to phylogenetics,with a growing number of applications in tree reconstruction(Suvorov et al.,2020;Zou et al.,2020;Nesterenko et al.,2022;Smith and Hahn,2023;Wang et al.,2023),substitution model selection(Abadi et al.,2020;Burgstaller-Muehlbacher et al.,2023),and diversification rate inference(Voznica et al.,2022;Lajaaiti et al.,2023;Lambert et al.,2023).In phylogenetic tree reconstruction,PhyDL(Zou et al.,2020)and Tree_learning(Suvorov et al.,2020)are two notable DNN-based programs designed to infer unrooted quartet trees directly from alignments of four amino acid(AA)and DNA sequences,respectively.
基金supported by the Natural Science Basic Research Program of Shaanxi(Program No.2024JC-YBMS-026).
摘要When dealing with imbalanced datasets,the traditional support vectormachine(SVM)tends to produce a classification hyperplane that is biased towards the majority class,which exhibits poor robustness.This paper proposes a high-performance classification algorithm specifically designed for imbalanced datasets.The proposed method first uses a biased second-order cone programming support vectormachine(B-SOCP-SVM)to identify the support vectors(SVs)and non-support vectors(NSVs)in the imbalanced data.Then,it applies the synthetic minority over-sampling technique(SV-SMOTE)to oversample the support vectors of the minority class and uses the random under-sampling technique(NSV-RUS)multiple times to undersample the non-support vectors of the majority class.Combining the above-obtained minority class data set withmultiple majority class datasets can obtainmultiple new balanced data sets.Finally,SOCP-SVM is used to classify each data set,and the final result is obtained through the integrated algorithm.Experimental results demonstrate that the proposed method performs excellently on imbalanced datasets.
基金supported by the Joint Funds of the National Natural Science Foundation of China(U24B20162)the National Natural Science Foundation of China(62373356)。
摘要Data-driven autonomous driving is a hot topic in academic and industry research due to its impressive performance,flexible mobility,and reduced human intervention.However,the development of this technology relies heavily on large datasets that contain accurately annotated data,obtained through artificial or semi-automated strategies.Consequently,datasets play a crucial role in autonomous driving,and their characteristics significantly impact the effectiveness of algorithms.Currently,there are several diverse datasets available,such as KITTI and City Scape,that cover various tasks.However,researchers often overlook the unique features,similarities,and specificities of these datasets.Furthermore,to the best of our knowledge,there is a lack of survey articles focusing on special metrics and benchmark performance on different datasets in autonomous driving.Therefore,the purpose of this article is to analyze autonomous driving datasets,guide researchers on collecting and utilizing relevant datasets,summarize evaluation strategies,analyze benchmark performance,and provide future research points to enrich the autonomous driving community.We believe that this work will assist researchers in evaluating their data using suitable metrics and offer a fresh perspective on autonomous driving.
基金supported in part by the Natural Science Foundation of Shaanxi Province of China under Grant 2024JC-YBQN-0695.
摘要This paper presents a systematic survey of machine vision-based surface defect detection technologies,focusing on five core challenges in the field:interference from complex backgrounds,small object detection,class imbalance,dynamic scene modeling,and cross-scenario generalization.It reviews key technical approaches corresponding to these challenges over the past five years.Furthermore,a dataset characterization analysis framework is established around these challenges,summarizing and comparing the characteristics of over 40 publicly available datasets across more than ten scenarios,including PCB,photovoltaic,metal,and pavement surfaces.Quantitative selection metrics(such as the small target coefficient and texture complexity)are proposed for challenges like small target detection and complex backgrounds,offering a methodological guide for aligning research questions with benchmark data.Finally,the paper summarizes current limitations and provides an outlook on new paradigms driven by large-scale models and the construction of high-quality benchmark datasets,aiming to offer valuable references for both research and engineering practices in this field.
摘要Small datasets are often challenging due to their limited sample size.This research introduces a novel solution to these problems:average linkage virtual sample generation(ALVSG).ALVSG leverages the underlying data structure to create virtual samples,which can be used to augment the original dataset.The ALVSG process consists of two steps.First,an average-linkage clustering technique is applied to the dataset to create a dendrogram.The dendrogram represents the hierarchical structure of the dataset,with each merging operation regarded as a linkage.Next,the linkages are combined into an average-based dataset,which serves as a new representation of the dataset.The second step in the ALVSG process involves generating virtual samples using the average-based dataset.The research project generates a set of 100 virtual samples by uniformly distributing them within the provided boundary.These virtual samples are then added to the original dataset,creating a more extensive dataset with improved generalization performance.The efficacy of the ALVSG approach is validated through resampling experiments and t-tests conducted on two small real-world datasets.The experiments are conducted on three forecasting models:the support vector machine for regression(SVR),the deep learning model(DL),and XGBoost.The results show that the ALVSG approach outperforms the baseline methods in terms of mean square error(MSE),root mean square error(RMSE),and mean absolute error(MAE).
基金funded by the Zhejiang Provincial Key Science and Technology“LingYan”Project Foundation,grant number 2023C01145Zhejiang Gongshang University Higher Education Research Projects,grant number Xgy22028.
摘要With the deep integration of smart manufacturing and IoT technologies,higher demands are placed on the intelligence and real-time performance of industrial equipment fault detection.For industrial fans,base bolt loosening faults are difficult to identify through conventional spectrum analysis,and the extreme scarcity of fault data leads to limited training datasets,making traditional deep learning methods inaccurate in fault identification and incapable of detecting loosening severity.This paper employs Bayesian Learning by training on a small fault dataset collected from the actual operation of axial-flow fans in a factory to obtain posterior distribution.This method proposes specific data processing approaches and a configuration of Bayesian Convolutional Neural Network(BCNN).It can effectively improve the model’s generalization ability.Experimental results demonstrate high detection accuracy and alignment with real-world applications,offering practical significance and reference value for industrial fan bolt loosening detection under data-limited conditions.
基金supported by Fundamental Research Funds of Central University Research Grant no:B240201122-Muhammad Ishfaque under the Post-Doctoral Research Program of Hohai University,Nanjing,Jiangsu Province of China.
摘要Monitoring concrete cracks for structural health in civil engineering presents a significant challenge.This is primarily due to the reliance on manual investigation methods,impacts of global climatic shifts stress,and geohazard threats to engineering structures.To cope with this challenge,state-of-the-art Deep Learning(DL)models are utilized to predict concrete cracks and accurately identify subtle variations in crack patterns and sizes,which lighting conditions and surface textures can influence.Previous studies indicate that model accuracy may decrease when faced with obscured concrete cracks,irregular shapes,or limited datasets for real-world problem scenarios.Feature fusion enhances model performance by combining complementary information,resulting in more accurate predictions,but may increase complexity and potential information redundancy.The study presents the Fractur Encoder to Decoder(FractED)block,a novel architecture consisting of three sub-blocks:the inner block(Encoder),intermediate block(Intermediate block),and outer block(Decoder).This approach integrates fused features into the model without additional fine-tuning steps,allowing for comprehensive feature refinement and enhancement,ultimately optimizing model performance.The study investigates a DL methodology on three datasets,demonstrating its effectiveness in handling complex classification scenarios in civil engineering.The model achieved high accuracy rates,with 88.41%for multiclass(Deck,Pavement,and Walls)classification tasks,91.94%on the Pillow Dam Borehole image binary dataset,and 99.77%on the Surface Crack binary dataset.The FractED block integration ensures adaptability and scalability,making it valuable for various Artificial Intelligence(AI)applications in civil engineering.The research also provides a scientific foundation for automatizing civil engineering inspection instruments for the future.
摘要With the continuous improvement of the performance of large language models,how to further enhance their ability in complex tasks has become a key issue.The task of abnormal text detection poses a challenge to the model in identifying non-standard semantics due to its semantic complexity and high-risk features.However,existing fine-tuning methods rely heavily on static data selection strategies,making it difficult to adapt to the dynamic evolution of model capabilities,resulting in low training efficiency.This article proposes ADS(Adaptive Dataset Selection),an adaptive framework for selecting data in anomaly text detection.ADS performs model-aware data selection prior to fine-tuning,adapting the initial state of pre-trained language models by selecting samples that are most informative for the target anomaly detection task.Empirical results on mainstream large language model architectures show that ADS significantly compresses data size while still outperforming existing static strategies and mainstream compression methods.When using only 1000 fine-tuning samples,ADS achieves a 92%F1 score,with an accuracy improvement of over 22%compared to the baseline,demonstrating excellent performance.This study proposes an efficient data selection mechanism from the perspective of model capability and dynamic adaptation of data,providing theoretical support and a practical path for fine-tuning large models in low-resource scenarios.
摘要Climate change significantly affects environment,ecosystems,communities,and economies.These impacts often result in quick and gradual changes in water resources,environmental conditions,and weather patterns.A geographical study was conducted in Arizona State,USA,to examine monthly precipi-tation concentration rates over time.This analysis used a high-resolution 0.50×0.50 grid for monthly precip-itation data from 1961 to 2022,Provided by the Climatic Research Unit.The study aimed to analyze climatic changes affected the first and last five years of each decade,as well as the entire decade,during the specified period.GIS was used to meet the objectives of this study.Arizona experienced 51–568 mm,67–560 mm,63–622 mm,and 52–590 mm of rainfall in the sixth,seventh,eighth,and ninth decades of the second millennium,respectively.Both the first and second five year periods of each decade showed accept-able rainfall amounts despite fluctuations.However,rainfall decreased in the first and second decades of the third millennium.and in the first two years of the third decade.Rainfall amounts dropped to 42–472 mm,55–469 mm,and 74–498 mm,respectively,indicating a downward trend in precipitation.The central part of the state received the highest rainfall,while the eastern and western regions(spanning north to south)had significantly less.Over the decades of the third millennium,the average annual rainfall every five years was relatively low,showing a declining trend due to severe climate changes,generally ranging between 35 mm and 498 mm.The central regions consistently received more rainfall than the eastern and western outskirts.Arizona is currently experiencing a decrease in rainfall due to climate change,a situation that could deterio-rate further.This highlights the need to optimize the use of existing rainfall and explore alternative water sources.