BACKGROUND Ischemic heart disease(IHD)impacts the quality of life and has the highest mortality rate of cardiovascular diseases globally.AIM To compare variations in the parameters of the single-lead electrocardiogram...BACKGROUND Ischemic heart disease(IHD)impacts the quality of life and has the highest mortality rate of cardiovascular diseases globally.AIM To compare variations in the parameters of the single-lead electrocardiogram(ECG)during resting conditions and physical exertion in individuals diagnosed with IHD and those without the condition using vasodilator-induced stress computed tomography(CT)myocardial perfusion imaging as the diagnostic reference standard.METHODS This single center observational study included 80 participants.The participants were aged≥40 years and given an informed written consent to participate in the study.Both groups,G1(n=31)with and G2(n=49)without post stress induced myocardial perfusion defect,passed cardiologist consultation,anthropometric measurements,blood pressure and pulse rate measurement,echocardiography,cardio-ankle vascular index,bicycle ergometry,recording 3-min single-lead ECG(Cardio-Qvark)before and just after bicycle ergometry followed by performing CT myocardial perfusion.The LASSO regression with nested cross-validation was used to find the association between Cardio-Qvark parameters and the existence of the perfusion defect.Statistical processing was performed with the R programming language v4.2,Python v.3.10[^R],and Statistica 12 program.RESULTS Bicycle ergometry yielded an area under the receiver operating characteristic curve of 50.7%[95%confidence interval(CI):0.388-0.625],specificity of 53.1%(95%CI:0.392-0.673),and sensitivity of 48.4%(95%CI:0.306-0.657).In contrast,the Cardio-Qvark test performed notably better with an area under the receiver operating characteristic curve of 67%(95%CI:0.530-0.801),specificity of 75.5%(95%CI:0.628-0.88),and sensitivity of 51.6%(95%CI:0.333-0.695).CONCLUSION The single-lead ECG has a relatively higher diagnostic accuracy compared with bicycle ergometry by using machine learning models,but the difference was not statistically significant.However,further investigations are required to uncover the hidden capabilities of single-lead ECG in IHD diagnosis.展开更多
Understanding spatial heterogeneity in groundwater responses to multiple factors is critical for water resource management in coastal cities.Daily groundwater depth(GWD)data from 43 wells(2018-2022)were collected in t...Understanding spatial heterogeneity in groundwater responses to multiple factors is critical for water resource management in coastal cities.Daily groundwater depth(GWD)data from 43 wells(2018-2022)were collected in three coastal cities in Jiangsu Province,China.Seasonal and Trend decomposition using Loess(STL)together with wavelet analysis and empirical mode decomposition were applied to identify tide-influenced wells while remaining wells were grouped by hierarchical clustering analysis(HCA).Machine learning models were developed to predict GWD,then their response to natural conditions and human activities was assessed by the Shapley Additive exPlanations(SHAP)method.Results showed that eXtreme Gradient Boosting(XGB)was superior to other models in terms of prediction performance and computational efficiency(R2>0.95).GWD in Yancheng and southern Lianyungang were greater than those in Nantong,exhibiting larger fluctuations.Groundwater within 5 km of the coastline was affected by tides,with more pronounced effects in agricultural areas compared to urban areas.Shallow groundwater(3-7 m depth)responded immediately(0-1 day)to rainfall,primarily influenced by farmland and topography(slope and distance from rivers).Rainfall recharge to groundwater peaked at 50%farmland coverage,but this effect was suppressed by high temperatures(>30℃)which intensified as distance from rivers increased,especially in forest and grassland.Deep groundwater(>10 m)showed delayed responses to rainfall(1-4 days)and temperature(10-15 days),with GDP as the primary influence,followed by agricultural irrigation and population density.Farmland helped to maintain stable GWD in low population density regions,while excessive farmland coverage(>90%)led to overexploitation.In the early stages of GDP development,increased industrial and agricultural water demand led to GWD decline,but as GDP levels significantly improved,groundwater consumption pressure gradually eased.This methodological framework is applicable not only to coastal cities in China but also could be extended to coastal regions worldwide.展开更多
The backwater effect caused by tributary inflow can significantly elevate the water level profile upstream of a confluence point.However,the influence of mainstream and confluence discharges on the backwater effect in...The backwater effect caused by tributary inflow can significantly elevate the water level profile upstream of a confluence point.However,the influence of mainstream and confluence discharges on the backwater effect in a river reach remains unclear.In this study,various hydrological data collected from the Jingjiang Reach of the Yangtze River in China were statistically analyzed to determine the backwater degree and range with three representative mainstream discharges.The results indicated that the backwater degree increased with mainstream discharge,and a positive relationship was observed between the runoff ratio and backwater degree at specific representative mainstream discharges.Following the operation of the Three Gorges Project,the backwater effect in the Jingjiang Reach diminished.For instance,mean backwater degrees for low,moderate,and high mainstream discharges were recorded as 0.83 m,1.61 m,and 2.41 m during the period from 1990 to 2002,whereas these values decreased to 0.30 m,0.95 m,and 2.08 m from 2009 to 2020.The backwater range extended upstream as mainstream discharge increased from 7000 m3/s to 30000 m3/s.Moreover,a random forest-based machine learning model was used to quantify the backwater effect with varying mainstream and confluence discharges,accounting for the impacts of mainstream discharge,confluence discharge,and channel degradation in the Jingjiang Reach.At the Jianli Hydrological Station,a decrease in mainstream discharge during flood seasons resulted in a 7%–15%increase in monthly mean backwater degree,while an increase in mainstream discharge during dry seasons led to a 1%–15%decrease in monthly mean backwater degree.Furthermore,increasing confluence discharge from Dongting Lake during June to July and September to November resulted in an 11%–42%increase in monthly mean backwater degree.Continuous channel degradation in the Jingjiang Reach contributed to a 6%–19%decrease in monthly mean backwater degree.Under the influence of these factors,the monthly mean backwater degree in 2017 varied from a decrease of 53%to an increase of 37%compared to corresponding values in 1991.展开更多
This research investigates the influence of indoor and outdoor factors on photovoltaic(PV)power generation at Utrecht University to accurately predict PV system performance by identifying critical impact factors and i...This research investigates the influence of indoor and outdoor factors on photovoltaic(PV)power generation at Utrecht University to accurately predict PV system performance by identifying critical impact factors and improving renewable energy efficiency.To predict plant efficiency,nineteen variables are analyzed,consisting of nine indoor photovoltaic panel characteristics(Open Circuit Voltage(Voc),Short Circuit Current(Isc),Maximum Power(Pmpp),Maximum Voltage(Umpp),Maximum Current(Impp),Filling Factor(FF),Parallel Resistance(Rp),Series Resistance(Rs),Module Temperature)and ten environmental factors(Air Temperature,Air Humidity,Dew Point,Air Pressure,Irradiation,Irradiation Propagation,Wind Speed,Wind Speed Propagation,Wind Direction,Wind Direction Propagation).This study provides a new perspective not previously addressed in the literature.In this study,different machine learning methods such as Multilayer Perceptron(MLP),Multivariate Adaptive Regression Spline(MARS),Multiple Linear Regression(MLR),and Random Forest(RF)models are used to predict power values using data from installed PVpanels.Panel values obtained under real field conditions were used to train the models,and the results were compared.The Multilayer Perceptron(MLP)model was achieved with the highest classification accuracy of 0.990%.The machine learning models used for solar energy forecasting show high performance and produce results close to actual values.Models like Multi-Layer Perceptron(MLP)and Random Forest(RF)can be used in diverse locations based on load demand.展开更多
The investigation by Zhu et al on the assessment of cellular proliferation markers to assist clinical decision-making in patients with hepatocellular carcinoma(HCC)using a machine learning model-based approach is a sc...The investigation by Zhu et al on the assessment of cellular proliferation markers to assist clinical decision-making in patients with hepatocellular carcinoma(HCC)using a machine learning model-based approach is a scientific approach.This study looked into the possibilities of using a Ki-67(a marker for cell proliferation)expression-based machine learning model to help doctors make decisions about treatment options for patients with HCC before surgery.The study used reconstructed tomography images of 164 patients with confirmed HCC from the intratumoral and peritumoral regions.The features were chosen using various statistical methods,including least absolute shrinkage and selection operator regression.Also,a nomogram was made using Radscore and clinical risk factors.It was tested for its ability to predict receiver operating characteristic curves and calibration curves,and its clinical benefits were found using decision curve analysis.The calibration curve demonstrated excellent consistency between predicted and actual probability,and the decision curve confirmed its clinical benefit.The proposed model is helpful for treating patients with HCC because the predicted and actual probabilities are very close to each other,as shown by the decision curve analysis.Further prospective studies are required,incorporating a multicenter and large sample size design,additional relevant exclusion criteria,information on tumors(size,number,and grade),and cancer stage to strengthen the clinical benefit in patients with HCC.展开更多
Floods and storm surges pose significant threats to coastal regions worldwide,demanding timely and accurate early warning systems(EWS)for disaster preparedness.Traditional numerical and statistical methods often fall ...Floods and storm surges pose significant threats to coastal regions worldwide,demanding timely and accurate early warning systems(EWS)for disaster preparedness.Traditional numerical and statistical methods often fall short in capturing complex,nonlinear,and real-time environmental dynamics.In recent years,machine learning(ML)and deep learning(DL)techniques have emerged as promising alternatives for enhancing the accuracy,speed,and scalability of EWS.This review critically evaluates the evolution of ML models—such as Artificial Neural Networks(ANN),Convolutional Neural Networks(CNN),and Long Short-Term Memory(LSTM)—in coastal flood prediction,highlighting their architectures,data requirements,performance metrics,and implementation challenges.A unique contribution of this work is the synthesis of real-time deployment challenges including latency,edge-cloud tradeoffs,and policy-level integration,areas often overlooked in prior literature.Furthermore,the review presents a comparative framework of model performance across different geographic and hydrologic settings,offering actionable insights for researchers and practitioners.Limitations of current AI-driven models,such as interpretability,data scarcity,and generalization across regions,are discussed in detail.Finally,the paper outlines future research directions including hybrid modelling,transfer learning,explainable AI,and policy-aware alert systems.By bridging technical performance and operational feasibility,this review aims to guide the development of next-generation intelligent EWS for resilient and adaptive coastal management.展开更多
To perform landslide susceptibility prediction(LSP),it is important to select appropriate mapping unit and landslide-related conditioning factors.The efficient and automatic multi-scale segmentation(MSS)method propose...To perform landslide susceptibility prediction(LSP),it is important to select appropriate mapping unit and landslide-related conditioning factors.The efficient and automatic multi-scale segmentation(MSS)method proposed by the authors promotes the application of slope units.However,LSP modeling based on these slope units has not been performed.Moreover,the heterogeneity of conditioning factors in slope units is neglected,leading to incomplete input variables of LSP modeling.In this study,the slope units extracted by the MSS method are used to construct LSP modeling,and the heterogeneity of conditioning factors is represented by the internal variations of conditioning factors within slope unit using the descriptive statistics features of mean,standard deviation and range.Thus,slope units-based machine learning models considering internal variations of conditioning factors(variant slope-machine learning)are proposed.The Chongyi County is selected as the case study and is divided into 53,055 slope units.Fifteen original slope unit-based conditioning factors are expanded to 38 slope unit-based conditioning factors through considering their internal variations.Random forest(RF)and multi-layer perceptron(MLP)machine learning models are used to construct variant Slope-RF and Slope-MLP models.Meanwhile,the Slope-RF and Slope-MLP models without considering the internal variations of conditioning factors,and conventional grid units-based machine learning(Grid-RF and MLP)models are built for comparisons through the LSP performance assessments.Results show that the variant Slopemachine learning models have higher LSP performances than Slope-machine learning models;LSP results of variant Slope-machine learning models have stronger directivity and practical application than Grid-machine learning models.It is concluded that slope units extracted by MSS method can be appropriate for LSP modeling,and the heterogeneity of conditioning factors within slope units can more comprehensively reflect the relationships between conditioning factors and landslides.The research results have important reference significance for land use and landslide prevention.展开更多
Landslide susceptibility assessment is crucial in predicting landslide occurrence and potential risks.However,traditional methods usually emphasize on larger regions of landsliding and rely on relatively static enviro...Landslide susceptibility assessment is crucial in predicting landslide occurrence and potential risks.However,traditional methods usually emphasize on larger regions of landsliding and rely on relatively static environmental conditions,which exposes the hysteresis of landslide susceptibility assessment in refined-scale and temporal dynamic changes.This study presents an improved landslide susceptibility assessment approach by integrating machine learning models based on random forest(RF),logical regression(LR),and gradient boosting decision tree(GBDT)with interferometric synthetic aperture radar(InSAR)technology and comparing them to their respective original models.The results demonstrated that the combined approach improves prediction accuracy and reduces the false negative and false positive errors.The LR-InSAR model showed the best performance in dynamic landslide susceptibility assessment at both regional and smaller scale,particularly when identifying areas of high and very high susceptibility.Modeling results were verified using data from field investigations including unmanned aerial vehicle(UAV)flights.This study is of great significance to accurately assess dynamic landslide susceptibility and to help reduce and prevent landslide risk.展开更多
The Indian Himalayan region is frequently experiencing climate change-induced landslides.Thus,landslide susceptibility assessment assumes greater significance for lessening the impact of a landslide hazard.This paper ...The Indian Himalayan region is frequently experiencing climate change-induced landslides.Thus,landslide susceptibility assessment assumes greater significance for lessening the impact of a landslide hazard.This paper makes an attempt to assess landslide susceptibility in Shimla district of the northwest Indian Himalayan region.It examined the effectiveness of random forest(RF),multilayer perceptron(MLP),sequential minimal optimization regression(SMOreg)and bagging ensemble(B-RF,BSMOreg,B-MLP)models.A landslide inventory map comprising 1052 locations of past landslide occurrences was classified into training(70%)and testing(30%)datasets.The site-specific influencing factors were selected by employing a multicollinearity test.The relationship between past landslide occurrences and influencing factors was established using the frequency ratio method.The effectiveness of machine learning models was verified through performance assessors.The landslide susceptibility maps were validated by the area under the receiver operating characteristic curves(ROC-AUC),accuracy,precision,recall and F1-score.The key performance metrics and map validation demonstrated that the BRF model(correlation coefficient:0.988,mean absolute error:0.010,root mean square error:0.058,relative absolute error:2.964,ROC-AUC:0.947,accuracy:0.778,precision:0.819,recall:0.917 and F-1 score:0.865)outperformed the single classifiers and other bagging ensemble models for landslide susceptibility.The results show that the largest area was found under the very high susceptibility zone(33.87%),followed by the low(27.30%),high(20.68%)and moderate(18.16%)susceptibility zones.The factors,namely average annual rainfall,slope,lithology,soil texture and earthquake magnitude have been identified as the influencing factors for very high landslide susceptibility.Soil texture,lineament density and elevation have been attributed to high and moderate susceptibility.Thus,the study calls for devising suitable landslide mitigation measures in the study area.Structural measures,an immediate response system,community participation and coordination among stakeholders may help lessen the detrimental impact of landslides.The findings from this study could aid decision-makers in mitigating future catastrophes and devising suitable strategies in other geographical regions with similar geological characteristics.展开更多
In recent years evidence has emerged suggesting that Mini-basketball training program(MBTP)can be an effec-tive intervention method to improve social communication(SC)impairments and restricted and repetitive beha-vio...In recent years evidence has emerged suggesting that Mini-basketball training program(MBTP)can be an effec-tive intervention method to improve social communication(SC)impairments and restricted and repetitive beha-viors(RRBs)in preschool children suffering from autism spectrum disorder(ASD).However,there is a considerable degree if interindividual variability concerning these social outcomes and thus not all preschool chil-dren with ASD profit from a MBTP intervention to the same extent.In order to make more accurate predictions which preschool children with ASD can benefit from an MBTP intervention or which preschool children with ASD need additional interventions to achieve behavioral improvements,further research is required.This study aimed to investigate which individual factors of preschool children with ASD can predict MBTP intervention out-comes concerning SC impairments and RRBs.Then,test the performance of machine learning models in predict-ing intervention outcomes based on these factors.Participants were 26 preschool children with ASD who enrolled in a quasi-experiment and received MBTP intervention.Baseline demographic variables(e.g.,age,body,mass index[BMI]),indicators of physicalfitness(e.g.,handgrip strength,balance performance),performance in execu-tive function,severity of ASD symptoms,level of SC impairments,and severity of RRBs were obtained to predict treatment outcomes after MBTP intervention.Machine learning models were established based on support vector machine algorithm were implemented.For comparison,we also employed multiple linear regression models in statistics.Ourfindings suggest that in preschool children with ASD symptomatic severity(r=0.712,p<0.001)and baseline SC impairments(r=0.713,p<0.001)are predictors for intervention outcomes of SC impair-ments.Furthermore,BMI(r=-0.430,p=0.028),symptomatic severity(r=0.656,p<0.001),baseline SC impair-ments(r=0.504,p=0.009)and baseline RRBs(r=0.647,p<0.001)can predict intervention outcomes of RRBs.Statistical models predicted 59.6%of variance in post-treatment SC impairments(MSE=0.455,RMSE=0.675,R2=0.596)and 58.9%of variance in post-treatment RRBs(MSE=0.464,RMSE=0.681,R2=0.589).Machine learning models predicted 83%of variance in post-treatment SC impairments(MSE=0.188,RMSE=0.434,R2=0.83)and 85.9%of variance in post-treatment RRBs(MSE=0.051,RMSE=0.226,R2=0.859),which were better than statistical models.Ourfindings suggest that baseline characteristics such as symptomatic severity of 144 IJMHP,2022,vol.24,no.2 ASD symptoms and SC impairments are important predictors determining MBTP intervention-induced improvements concerning SC impairments and RBBs.Furthermore,the current study revealed that machine learning models can successfully be applied to predict the MBTP intervention-related outcomes in preschool chil-dren with ASD,and performed better than statistical models.Ourfindings can help to inform which preschool children with ASD are most likely to benefit from an MBTP intervention,and they might provide a reference for the development of personalized intervention programs for preschool children with ASD.展开更多
Unconventional reservoirs have become the main alternative for increasing oil and gas reserves around the world. Owing to their ultralow permeability properties and special pore structure, hydraulic fracturing technol...Unconventional reservoirs have become the main alternative for increasing oil and gas reserves around the world. Owing to their ultralow permeability properties and special pore structure, hydraulic fracturing technology is necessary to realize the efficient development and economic management of unconventional resources. To maximize the production capacity of wells, several fracture parameters, including fracture number, length, width, conductivity, and spacing, need to be optimized effectively. The optimization of hydraulic fracture parameters in shale gas reservoirs generally demands intensive computations owing to the necessity of numerous physicalmodel simulations. This study proposes a machine learning (ML)–assisted global optimization framework to rapidly obtain optimal fracture parameters. We employed three supervised ML models, including the radialbasis function, K-nearest neighbor, and multilayer perceptron, to emulate the relationship between fracture parameters and shale gas productivity for multistage fractured horizontal wells. Firstly, several forward shale gas simulations with embedded discrete fracture models generate training samples. Then, the samples are divided into training and testing samples to train these ML models and optimize network hyper parameters, respectively. Finally, the trained ML models are combined with an intelligent differential evolution algorithm to optimize the fracture parameters. This novel method has been applied to a naturally fractured reservoir model based on the real-field Barnett shale formation. The obtained results are compared with those of conventional optimizations with high-fidelity models. The results confirm the superiority of the proposed method owing to its very low computational cost. The use of ML modeling technology and an intelligent optimization algorithm could greatly contribute to simulation optimization and design, prompting progress in the intelligent development of unconventional oil and gas reservoirs in China.展开更多
Machine learning models were used to improve the accuracy of China Meteorological Administration Multisource Precipitation Analysis System(CMPAS)in complex terrain areas by combining rain gauge precipitation with topo...Machine learning models were used to improve the accuracy of China Meteorological Administration Multisource Precipitation Analysis System(CMPAS)in complex terrain areas by combining rain gauge precipitation with topographic factors like altitude,slope,slope direction,slope variability,surface roughness,and meteorological factors like temperature and wind speed.The results of the correction demonstrated that the ensemble learning method has a considerably corrective effect and the three methods(Random Forest,AdaBoost,and Bagging)adopted in the study had similar results.The mean bias between CMPAS and 85%of automatic weather stations has dropped by more than 30%.The plateau region displays the largest accuracy increase,the winter season shows the greatest error reduction,and decreasing precipitation improves the correction outcome.Additionally,the heavy precipitation process’precision has improved to some degree.For individual stations,the revised CMPAS error fluctuation range is significantly reduced.展开更多
BACKGROUND Colorectal cancer significantly impacts global health,with unplanned reoperations post-surgery being key determinants of patient outcomes.Existing predictive models for these reoperations lack precision in ...BACKGROUND Colorectal cancer significantly impacts global health,with unplanned reoperations post-surgery being key determinants of patient outcomes.Existing predictive models for these reoperations lack precision in integrating complex clinical data.AIM To develop and validate a machine learning model for predicting unplanned reoperation risk in colorectal cancer patients.METHODS Data of patients treated for colorectal cancer(n=2044)at the First Affiliated Hospital of Wenzhou Medical University and Wenzhou Central Hospital from March 2020 to March 2022 were retrospectively collected.Patients were divided into an experimental group(n=60)and a control group(n=1984)according to unplanned reoperation occurrence.Patients were also divided into a training group and a validation group(7:3 ratio).We used three different machine learning methods to screen characteristic variables.A nomogram was created based on multifactor logistic regression,and the model performance was assessed using receiver operating characteristic curve,calibration curve,Hosmer-Lemeshow test,and decision curve analysis.The risk scores of the two groups were calculated and compared to validate the model.RESULTS More patients in the experimental group were≥60 years old,male,and had a history of hypertension,laparotomy,and hypoproteinemia,compared to the control group.Multiple logistic regression analysis confirmed the following as independent risk factors for unplanned reoperation(P<0.05):Prognostic Nutritional Index value,history of laparotomy,hypertension,or stroke,hypoproteinemia,age,tumor-node-metastasis staging,surgical time,gender,and American Society of Anesthesiologists classification.Receiver operating characteristic curve analysis showed that the model had good discrimination and clinical utility.CONCLUSION This study used a machine learning approach to build a model that accurately predicts the risk of postoperative unplanned reoperation in patients with colorectal cancer,which can improve treatment decisions and prognosis.展开更多
BACKGROUND Liver transplantation(LT)is a life-saving intervention for patients with end-stage liver disease.However,the equitable allocation of scarce donor organs remains a formidable challenge.Prognostic tools are p...BACKGROUND Liver transplantation(LT)is a life-saving intervention for patients with end-stage liver disease.However,the equitable allocation of scarce donor organs remains a formidable challenge.Prognostic tools are pivotal in identifying the most suitable transplant candidates.Traditionally,scoring systems like the model for end-stage liver disease have been instrumental in this process.Nevertheless,the landscape of prognostication is undergoing a transformation with the integration of machine learning(ML)and artificial intelligence models.AIM To assess the utility of ML models in prognostication for LT,comparing their performance and reliability to established traditional scoring systems.METHODS Following the Preferred Reporting Items for Systematic Reviews and Meta-Analysis guidelines,we conducted a thorough and standardized literature search using the PubMed/MEDLINE database.Our search imposed no restrictions on publication year,age,or gender.Exclusion criteria encompassed non-English studies,review articles,case reports,conference papers,studies with missing data,or those exhibiting evident methodological flaws.RESULTS Our search yielded a total of 64 articles,with 23 meeting the inclusion criteria.Among the selected studies,60.8%originated from the United States and China combined.Only one pediatric study met the criteria.Notably,91%of the studies were published within the past five years.ML models consistently demonstrated satisfactory to excellent area under the receiver operating characteristic curve values(ranging from 0.6 to 1)across all studies,surpassing the performance of traditional scoring systems.Random forest exhibited superior predictive capabilities for 90-d mortality following LT,sepsis,and acute kidney injury(AKI).In contrast,gradient boosting excelled in predicting the risk of graft-versus-host disease,pneumonia,and AKI.CONCLUSION This study underscores the potential of ML models in guiding decisions related to allograft allocation and LT,marking a significant evolution in the field of prognostication.展开更多
Background:With high colorectal cancer(CRC)incidence,accurate early differentiation of precancerous polyps is critical for prognosis;while the gold-standard colonoscopy-biopsy is limited by invasiveness,cost and poor ...Background:With high colorectal cancer(CRC)incidence,accurate early differentiation of precancerous polyps is critical for prognosis;while the gold-standard colonoscopy-biopsy is limited by invasiveness,cost and poor scalability,and routine blood tests lack efficiency with simple indicator combinations,machine learning's feature-mining capacity offers a solution.This study aimed to develop and evaluate a machine learning system for differentiating patients with CRC and colorectal polyps using routine blood indices.Methods:A retrospective analysis was conducted on the clinical data of 284 patients with CRC and 79 patients with colorectal polyps who were diagnosed at the Chinese PLA General Hospital from October 2021 to February 2024.The extreme gradient boosting(XGBoost)algorithm was used to establish a machine learning model using demographic characteristics and routine blood indices.The Shapley additive explanation method was used to evaluate feature importance.Results:The constructed XGBoost model achieved high levels in differentiating CRC and colorectal polyps,with a precision of 0.906,a recall of 0.817,an accuracy of 0.791,and an area under the receiver operating characteristic curve of 0.869.The Shapley additive explanation showed that the top five important features were fibrinogen,carcinoembryonic antigen,plasma thrombin time,ferritin,and D-dimer.Conclusion:The XGBoost machine learning model based on blood indices has certain application value in the differential diagnosis of CRC and colorectal polyps,providing a new efficient tool for auxiliary diagnosis to assist clinical decision-making.展开更多
The sustainable development of nuclear energy requires efficient and reusable materials for uranium extraction from seawater,where uranyl species exist at ultra-trace levels within complex coordination environments.To...The sustainable development of nuclear energy requires efficient and reusable materials for uranium extraction from seawater,where uranyl species exist at ultra-trace levels within complex coordination environments.To overcome the limitations of empirical,trial-and-error design,this study establishes U-Predict v1.0,a structure-performance database for uranium adsorbents,and introduces an interpretable machine learning framework for adsorption performance prediction and mechanistic analysis.Trained on over 220 samples and 54 structural descriptors,the Light Gradient Boosting Machine(LightGBM)model with engineered features and optimized hyperparameters achieved a training R2 of 0.9887 and a test R2 of 0.7501,indicating robust predictive performance under heterogeneous and literature-derived conditions.SHapley Additive exPlanations(SHAP)interpretation revealed that structural and environmental variables,particularly functional group chemistry,solution pH,and surface area,jointly govern uranium the maximum adsorption capacity(qmax)through nonlinear interactions,emphasizing the dual control of material composition and adsorption environment.Experimental validation using a representative high-performance covalent organic frameworks-based adsorbent confirmed the model's predictive reliability,with the measured qmax of 342.2 mg g-1 showing reasonable agreement with the predicted value(449.3 mg g-1;23.8%deviation).This study establishes an explainable AI paradigm that bridges modeling and experimentation,providing a transferable foundation to support the intelligent discovery and sustainable design of next-generation uranium adsorbents for high-efficiency extraction.展开更多
ARC inoculant(A,aflatoxin prevention and control;R,Rhizobia nodulation induction;C,Coupling)is a brandnew inoculant with coupling function that enhances legume quality and nitrogen fixation.Comprehensive characterizat...ARC inoculant(A,aflatoxin prevention and control;R,Rhizobia nodulation induction;C,Coupling)is a brandnew inoculant with coupling function that enhances legume quality and nitrogen fixation.Comprehensive characterization of its key functional strains is critical for establishing a quality-control framework for the inoculant's formulation.Here,we constructed a characteristic spectral dataset comprising over 63,000 single-cell Raman spectra of the constituent strains by employing Ramanome technology.Six machine learning-based predictive models were developed and compared for the constituent strains,while the Linear Discriminant Analysis(LDA)model demonstrated the best performance,with a classification accuracy exceeding 92.4%.This work provides a unique spectral fingerprint for ARC inoculant and will directly aid its application in sustainable agricultural production.展开更多
Understanding soil depth is fundamental for evaluating a soil's ability to support vegetation,retain water and nutrients,sequester carbon,and sustain infrastructure.Despite its importance,the spatial variability o...Understanding soil depth is fundamental for evaluating a soil's ability to support vegetation,retain water and nutrients,sequester carbon,and sustain infrastructure.Despite its importance,the spatial variability of soil depth remains poorly understood,especially in mountainous terrains with limited sampling.Digital Soil Mapping(DSM),empowered by Machine Learning(ML)algorithms,offers a practical solution for estimating soil properties across complex landscapes.This study aimed to develop a predictive soil depth map for Nagaland,a hilly state in Northeastern India,using ML techniques.A dataset comprising soil depth observations from 70locations and 37 environmental covariates was analysed.Four ML algorithms Random Forest(RF),Support Vector Machine(SVM),Extreme Gradient Boosting(XGB),and K-Nearest Neighbours(KNN)were employed.Covariates derived from MODIS,Worldclim,and digital elevation models(DEM)were used to model soil-landscape relationships.Among all predictors,BIO15(precipitation seasonality)and NDVI(Normalized Difference Vegetation Index)emerged as key variables influencing soil depth variation.The soil depth predictions across models ranged between 44 and 158 cm,with XGB showing the best performance(R2=0.32±0.24;RMSE=22.63±4.74 cm;MAE=18.59±3.99 cm),followed by RF,while SVM and KNN were comparatively less accurate.The spatial distribution revealed that deeper soils were concentrated in the central and northern plains and valleys,while shallower soils were predominantly found in the steep,dissected Southeastern Hill Region.The findings highlight the efficacy of ML-driven DSM in mountainous,data-scarce environments,offering a valuable tool for land use planning,soil conservation,and ecological management in the Eastern Himalayas.展开更多
Over the past decades,surrogate model-aided reliability analysis approaches grounded in active learning have undergone extensive development.However,Gaussian process models like Kriging suffer from severe computationa...Over the past decades,surrogate model-aided reliability analysis approaches grounded in active learning have undergone extensive development.However,Gaussian process models like Kriging suffer from severe computational burdens when handling high-dimensional problems or large samples.Conversely,machine learning algorithms such as extreme learning machines exhibit high computational efficiency but lack variance output and stability,making them difficult to employ for adaptive active learning strategies.To address these limitations,this study proposes a population Monte Carlo method based on an adaptive closed neuron extreme learning machine.First,a closed neuron strategy uses a consistency metric to screen and retain neurons containing the most informative features.This preserves the fast analytical solution advantage of extreme learning machines while significantly improving the reconstruction accuracy and stability of the true limit state surface.Second,to overcome the lack of variance in the output,an ensemble model is constructed.By calculating predictive mean and standard deviation,a learning function is formulated for efficient adaptive sample enrichment.Finally,utilizing the adaptive importance sampling mechanism of the population Monte Carlo framework,the auxiliary density function is optimized to progressively shift the sampling center toward high contribution failure regions.Four engineering examples confirm that the proposed method achieves exceptional computational efficiency and high accuracy for complex reliability analysis involving extremely small failure probabilities.展开更多
Knowledge of the domain of applicability of a machine learning model is essential to ensuring accurate and reliable model predictions.In this work,we develop a new and general approach of assessing model domain and de...Knowledge of the domain of applicability of a machine learning model is essential to ensuring accurate and reliable model predictions.In this work,we develop a new and general approach of assessing model domain and demonstrate that our approach provides accurate and meaningful domain designation across multiple model types and material property data sets.Our approach assesses the distance between data in feature space using kernel density estimation,where this distance provides an effective tool for domain determination.We show that chemical groups considered unrelated based on chemical knowledge exhibit significant dissimilarities by our measure.We also show that high measures of dissimilarity are associated with poor model performance(i.e.,high residual magnitudes)and poor estimates of model uncertainty(i.e.,unreliable uncertainty estimation).Automated tools are provided to enable researchers to establish acceptable dissimilarity thresholds to identify whether new predictions of their own machine learning models are in-domain versus out-of-domain.展开更多
基金Supported by Government Assignment,No.1023022600020-6RSF Grant,No.24-15-00549Ministry of Science and Higher Education of the Russian Federation within the Framework of State Support for the Creation and Development of World-Class Research Center,No.075-15-2022-304.
摘要BACKGROUND Ischemic heart disease(IHD)impacts the quality of life and has the highest mortality rate of cardiovascular diseases globally.AIM To compare variations in the parameters of the single-lead electrocardiogram(ECG)during resting conditions and physical exertion in individuals diagnosed with IHD and those without the condition using vasodilator-induced stress computed tomography(CT)myocardial perfusion imaging as the diagnostic reference standard.METHODS This single center observational study included 80 participants.The participants were aged≥40 years and given an informed written consent to participate in the study.Both groups,G1(n=31)with and G2(n=49)without post stress induced myocardial perfusion defect,passed cardiologist consultation,anthropometric measurements,blood pressure and pulse rate measurement,echocardiography,cardio-ankle vascular index,bicycle ergometry,recording 3-min single-lead ECG(Cardio-Qvark)before and just after bicycle ergometry followed by performing CT myocardial perfusion.The LASSO regression with nested cross-validation was used to find the association between Cardio-Qvark parameters and the existence of the perfusion defect.Statistical processing was performed with the R programming language v4.2,Python v.3.10[^R],and Statistica 12 program.RESULTS Bicycle ergometry yielded an area under the receiver operating characteristic curve of 50.7%[95%confidence interval(CI):0.388-0.625],specificity of 53.1%(95%CI:0.392-0.673),and sensitivity of 48.4%(95%CI:0.306-0.657).In contrast,the Cardio-Qvark test performed notably better with an area under the receiver operating characteristic curve of 67%(95%CI:0.530-0.801),specificity of 75.5%(95%CI:0.628-0.88),and sensitivity of 51.6%(95%CI:0.333-0.695).CONCLUSION The single-lead ECG has a relatively higher diagnostic accuracy compared with bicycle ergometry by using machine learning models,but the difference was not statistically significant.However,further investigations are required to uncover the hidden capabilities of single-lead ECG in IHD diagnosis.
基金supported by the Natural Science Foundation of Jiangsu province,China(BK20240937)the Belt and Road Special Foundation of the National Key Laboratory of Water Disaster Prevention(2022491411,2021491811)the Basal Research Fund of Central Public Welfare Scientific Institution of Nanjing Hydraulic Research Institute(Y223006).
摘要Understanding spatial heterogeneity in groundwater responses to multiple factors is critical for water resource management in coastal cities.Daily groundwater depth(GWD)data from 43 wells(2018-2022)were collected in three coastal cities in Jiangsu Province,China.Seasonal and Trend decomposition using Loess(STL)together with wavelet analysis and empirical mode decomposition were applied to identify tide-influenced wells while remaining wells were grouped by hierarchical clustering analysis(HCA).Machine learning models were developed to predict GWD,then their response to natural conditions and human activities was assessed by the Shapley Additive exPlanations(SHAP)method.Results showed that eXtreme Gradient Boosting(XGB)was superior to other models in terms of prediction performance and computational efficiency(R2>0.95).GWD in Yancheng and southern Lianyungang were greater than those in Nantong,exhibiting larger fluctuations.Groundwater within 5 km of the coastline was affected by tides,with more pronounced effects in agricultural areas compared to urban areas.Shallow groundwater(3-7 m depth)responded immediately(0-1 day)to rainfall,primarily influenced by farmland and topography(slope and distance from rivers).Rainfall recharge to groundwater peaked at 50%farmland coverage,but this effect was suppressed by high temperatures(>30℃)which intensified as distance from rivers increased,especially in forest and grassland.Deep groundwater(>10 m)showed delayed responses to rainfall(1-4 days)and temperature(10-15 days),with GDP as the primary influence,followed by agricultural irrigation and population density.Farmland helped to maintain stable GWD in low population density regions,while excessive farmland coverage(>90%)led to overexploitation.In the early stages of GDP development,increased industrial and agricultural water demand led to GWD decline,but as GDP levels significantly improved,groundwater consumption pressure gradually eased.This methodological framework is applicable not only to coastal cities in China but also could be extended to coastal regions worldwide.
基金supported by the National Key Research and Development Program of China(Grant No.2023YFC3209504)the National Natural Science Foundation of China(Grants No.U2040215 and 52479075)the Natural Science Foundation of Hubei Province(Grant No.2021CFA029).
摘要The backwater effect caused by tributary inflow can significantly elevate the water level profile upstream of a confluence point.However,the influence of mainstream and confluence discharges on the backwater effect in a river reach remains unclear.In this study,various hydrological data collected from the Jingjiang Reach of the Yangtze River in China were statistically analyzed to determine the backwater degree and range with three representative mainstream discharges.The results indicated that the backwater degree increased with mainstream discharge,and a positive relationship was observed between the runoff ratio and backwater degree at specific representative mainstream discharges.Following the operation of the Three Gorges Project,the backwater effect in the Jingjiang Reach diminished.For instance,mean backwater degrees for low,moderate,and high mainstream discharges were recorded as 0.83 m,1.61 m,and 2.41 m during the period from 1990 to 2002,whereas these values decreased to 0.30 m,0.95 m,and 2.08 m from 2009 to 2020.The backwater range extended upstream as mainstream discharge increased from 7000 m3/s to 30000 m3/s.Moreover,a random forest-based machine learning model was used to quantify the backwater effect with varying mainstream and confluence discharges,accounting for the impacts of mainstream discharge,confluence discharge,and channel degradation in the Jingjiang Reach.At the Jianli Hydrological Station,a decrease in mainstream discharge during flood seasons resulted in a 7%–15%increase in monthly mean backwater degree,while an increase in mainstream discharge during dry seasons led to a 1%–15%decrease in monthly mean backwater degree.Furthermore,increasing confluence discharge from Dongting Lake during June to July and September to November resulted in an 11%–42%increase in monthly mean backwater degree.Continuous channel degradation in the Jingjiang Reach contributed to a 6%–19%decrease in monthly mean backwater degree.Under the influence of these factors,the monthly mean backwater degree in 2017 varied from a decrease of 53%to an increase of 37%compared to corresponding values in 1991.
摘要This research investigates the influence of indoor and outdoor factors on photovoltaic(PV)power generation at Utrecht University to accurately predict PV system performance by identifying critical impact factors and improving renewable energy efficiency.To predict plant efficiency,nineteen variables are analyzed,consisting of nine indoor photovoltaic panel characteristics(Open Circuit Voltage(Voc),Short Circuit Current(Isc),Maximum Power(Pmpp),Maximum Voltage(Umpp),Maximum Current(Impp),Filling Factor(FF),Parallel Resistance(Rp),Series Resistance(Rs),Module Temperature)and ten environmental factors(Air Temperature,Air Humidity,Dew Point,Air Pressure,Irradiation,Irradiation Propagation,Wind Speed,Wind Speed Propagation,Wind Direction,Wind Direction Propagation).This study provides a new perspective not previously addressed in the literature.In this study,different machine learning methods such as Multilayer Perceptron(MLP),Multivariate Adaptive Regression Spline(MARS),Multiple Linear Regression(MLR),and Random Forest(RF)models are used to predict power values using data from installed PVpanels.Panel values obtained under real field conditions were used to train the models,and the results were compared.The Multilayer Perceptron(MLP)model was achieved with the highest classification accuracy of 0.990%.The machine learning models used for solar energy forecasting show high performance and produce results close to actual values.Models like Multi-Layer Perceptron(MLP)and Random Forest(RF)can be used in diverse locations based on load demand.
摘要The investigation by Zhu et al on the assessment of cellular proliferation markers to assist clinical decision-making in patients with hepatocellular carcinoma(HCC)using a machine learning model-based approach is a scientific approach.This study looked into the possibilities of using a Ki-67(a marker for cell proliferation)expression-based machine learning model to help doctors make decisions about treatment options for patients with HCC before surgery.The study used reconstructed tomography images of 164 patients with confirmed HCC from the intratumoral and peritumoral regions.The features were chosen using various statistical methods,including least absolute shrinkage and selection operator regression.Also,a nomogram was made using Radscore and clinical risk factors.It was tested for its ability to predict receiver operating characteristic curves and calibration curves,and its clinical benefits were found using decision curve analysis.The calibration curve demonstrated excellent consistency between predicted and actual probability,and the decision curve confirmed its clinical benefit.The proposed model is helpful for treating patients with HCC because the predicted and actual probabilities are very close to each other,as shown by the decision curve analysis.Further prospective studies are required,incorporating a multicenter and large sample size design,additional relevant exclusion criteria,information on tumors(size,number,and grade),and cancer stage to strengthen the clinical benefit in patients with HCC.
摘要Floods and storm surges pose significant threats to coastal regions worldwide,demanding timely and accurate early warning systems(EWS)for disaster preparedness.Traditional numerical and statistical methods often fall short in capturing complex,nonlinear,and real-time environmental dynamics.In recent years,machine learning(ML)and deep learning(DL)techniques have emerged as promising alternatives for enhancing the accuracy,speed,and scalability of EWS.This review critically evaluates the evolution of ML models—such as Artificial Neural Networks(ANN),Convolutional Neural Networks(CNN),and Long Short-Term Memory(LSTM)—in coastal flood prediction,highlighting their architectures,data requirements,performance metrics,and implementation challenges.A unique contribution of this work is the synthesis of real-time deployment challenges including latency,edge-cloud tradeoffs,and policy-level integration,areas often overlooked in prior literature.Furthermore,the review presents a comparative framework of model performance across different geographic and hydrologic settings,offering actionable insights for researchers and practitioners.Limitations of current AI-driven models,such as interpretability,data scarcity,and generalization across regions,are discussed in detail.Finally,the paper outlines future research directions including hybrid modelling,transfer learning,explainable AI,and policy-aware alert systems.By bridging technical performance and operational feasibility,this review aims to guide the development of next-generation intelligent EWS for resilient and adaptive coastal management.
基金funded by the Natural Science Foundation of China(Grant Nos.41807285,41972280 and 52179103).
摘要To perform landslide susceptibility prediction(LSP),it is important to select appropriate mapping unit and landslide-related conditioning factors.The efficient and automatic multi-scale segmentation(MSS)method proposed by the authors promotes the application of slope units.However,LSP modeling based on these slope units has not been performed.Moreover,the heterogeneity of conditioning factors in slope units is neglected,leading to incomplete input variables of LSP modeling.In this study,the slope units extracted by the MSS method are used to construct LSP modeling,and the heterogeneity of conditioning factors is represented by the internal variations of conditioning factors within slope unit using the descriptive statistics features of mean,standard deviation and range.Thus,slope units-based machine learning models considering internal variations of conditioning factors(variant slope-machine learning)are proposed.The Chongyi County is selected as the case study and is divided into 53,055 slope units.Fifteen original slope unit-based conditioning factors are expanded to 38 slope unit-based conditioning factors through considering their internal variations.Random forest(RF)and multi-layer perceptron(MLP)machine learning models are used to construct variant Slope-RF and Slope-MLP models.Meanwhile,the Slope-RF and Slope-MLP models without considering the internal variations of conditioning factors,and conventional grid units-based machine learning(Grid-RF and MLP)models are built for comparisons through the LSP performance assessments.Results show that the variant Slopemachine learning models have higher LSP performances than Slope-machine learning models;LSP results of variant Slope-machine learning models have stronger directivity and practical application than Grid-machine learning models.It is concluded that slope units extracted by MSS method can be appropriate for LSP modeling,and the heterogeneity of conditioning factors within slope units can more comprehensively reflect the relationships between conditioning factors and landslides.The research results have important reference significance for land use and landslide prevention.
基金the National Natural Science Foundation of China(Grant No.42271078)Key Research and Development Program of Shaanxi(2024SF-YBXM-669)。
摘要Landslide susceptibility assessment is crucial in predicting landslide occurrence and potential risks.However,traditional methods usually emphasize on larger regions of landsliding and rely on relatively static environmental conditions,which exposes the hysteresis of landslide susceptibility assessment in refined-scale and temporal dynamic changes.This study presents an improved landslide susceptibility assessment approach by integrating machine learning models based on random forest(RF),logical regression(LR),and gradient boosting decision tree(GBDT)with interferometric synthetic aperture radar(InSAR)technology and comparing them to their respective original models.The results demonstrated that the combined approach improves prediction accuracy and reduces the false negative and false positive errors.The LR-InSAR model showed the best performance in dynamic landslide susceptibility assessment at both regional and smaller scale,particularly when identifying areas of high and very high susceptibility.Modeling results were verified using data from field investigations including unmanned aerial vehicle(UAV)flights.This study is of great significance to accurately assess dynamic landslide susceptibility and to help reduce and prevent landslide risk.
摘要The Indian Himalayan region is frequently experiencing climate change-induced landslides.Thus,landslide susceptibility assessment assumes greater significance for lessening the impact of a landslide hazard.This paper makes an attempt to assess landslide susceptibility in Shimla district of the northwest Indian Himalayan region.It examined the effectiveness of random forest(RF),multilayer perceptron(MLP),sequential minimal optimization regression(SMOreg)and bagging ensemble(B-RF,BSMOreg,B-MLP)models.A landslide inventory map comprising 1052 locations of past landslide occurrences was classified into training(70%)and testing(30%)datasets.The site-specific influencing factors were selected by employing a multicollinearity test.The relationship between past landslide occurrences and influencing factors was established using the frequency ratio method.The effectiveness of machine learning models was verified through performance assessors.The landslide susceptibility maps were validated by the area under the receiver operating characteristic curves(ROC-AUC),accuracy,precision,recall and F1-score.The key performance metrics and map validation demonstrated that the BRF model(correlation coefficient:0.988,mean absolute error:0.010,root mean square error:0.058,relative absolute error:2.964,ROC-AUC:0.947,accuracy:0.778,precision:0.819,recall:0.917 and F-1 score:0.865)outperformed the single classifiers and other bagging ensemble models for landslide susceptibility.The results show that the largest area was found under the very high susceptibility zone(33.87%),followed by the low(27.30%),high(20.68%)and moderate(18.16%)susceptibility zones.The factors,namely average annual rainfall,slope,lithology,soil texture and earthquake magnitude have been identified as the influencing factors for very high landslide susceptibility.Soil texture,lineament density and elevation have been attributed to high and moderate susceptibility.Thus,the study calls for devising suitable landslide mitigation measures in the study area.Structural measures,an immediate response system,community participation and coordination among stakeholders may help lessen the detrimental impact of landslides.The findings from this study could aid decision-makers in mitigating future catastrophes and devising suitable strategies in other geographical regions with similar geological characteristics.
基金supported by grants from the National Natural Science Foundation of China(31771243)the Fok Ying Tong Education Foundation(141113)to Aiguo Chen.
摘要In recent years evidence has emerged suggesting that Mini-basketball training program(MBTP)can be an effec-tive intervention method to improve social communication(SC)impairments and restricted and repetitive beha-viors(RRBs)in preschool children suffering from autism spectrum disorder(ASD).However,there is a considerable degree if interindividual variability concerning these social outcomes and thus not all preschool chil-dren with ASD profit from a MBTP intervention to the same extent.In order to make more accurate predictions which preschool children with ASD can benefit from an MBTP intervention or which preschool children with ASD need additional interventions to achieve behavioral improvements,further research is required.This study aimed to investigate which individual factors of preschool children with ASD can predict MBTP intervention out-comes concerning SC impairments and RRBs.Then,test the performance of machine learning models in predict-ing intervention outcomes based on these factors.Participants were 26 preschool children with ASD who enrolled in a quasi-experiment and received MBTP intervention.Baseline demographic variables(e.g.,age,body,mass index[BMI]),indicators of physicalfitness(e.g.,handgrip strength,balance performance),performance in execu-tive function,severity of ASD symptoms,level of SC impairments,and severity of RRBs were obtained to predict treatment outcomes after MBTP intervention.Machine learning models were established based on support vector machine algorithm were implemented.For comparison,we also employed multiple linear regression models in statistics.Ourfindings suggest that in preschool children with ASD symptomatic severity(r=0.712,p<0.001)and baseline SC impairments(r=0.713,p<0.001)are predictors for intervention outcomes of SC impair-ments.Furthermore,BMI(r=-0.430,p=0.028),symptomatic severity(r=0.656,p<0.001),baseline SC impair-ments(r=0.504,p=0.009)and baseline RRBs(r=0.647,p<0.001)can predict intervention outcomes of RRBs.Statistical models predicted 59.6%of variance in post-treatment SC impairments(MSE=0.455,RMSE=0.675,R2=0.596)and 58.9%of variance in post-treatment RRBs(MSE=0.464,RMSE=0.681,R2=0.589).Machine learning models predicted 83%of variance in post-treatment SC impairments(MSE=0.188,RMSE=0.434,R2=0.83)and 85.9%of variance in post-treatment RRBs(MSE=0.051,RMSE=0.226,R2=0.859),which were better than statistical models.Ourfindings suggest that baseline characteristics such as symptomatic severity of 144 IJMHP,2022,vol.24,no.2 ASD symptoms and SC impairments are important predictors determining MBTP intervention-induced improvements concerning SC impairments and RBBs.Furthermore,the current study revealed that machine learning models can successfully be applied to predict the MBTP intervention-related outcomes in preschool chil-dren with ASD,and performed better than statistical models.Ourfindings can help to inform which preschool children with ASD are most likely to benefit from an MBTP intervention,and they might provide a reference for the development of personalized intervention programs for preschool children with ASD.
基金supported by Science Foundation of China University of Petroleum,Beijing(No.2462021BJRC005)Key Technologies of Mahu Conglomerate Reservoir(ZLZX2020-01-04).
摘要Unconventional reservoirs have become the main alternative for increasing oil and gas reserves around the world. Owing to their ultralow permeability properties and special pore structure, hydraulic fracturing technology is necessary to realize the efficient development and economic management of unconventional resources. To maximize the production capacity of wells, several fracture parameters, including fracture number, length, width, conductivity, and spacing, need to be optimized effectively. The optimization of hydraulic fracture parameters in shale gas reservoirs generally demands intensive computations owing to the necessity of numerous physicalmodel simulations. This study proposes a machine learning (ML)–assisted global optimization framework to rapidly obtain optimal fracture parameters. We employed three supervised ML models, including the radialbasis function, K-nearest neighbor, and multilayer perceptron, to emulate the relationship between fracture parameters and shale gas productivity for multistage fractured horizontal wells. Firstly, several forward shale gas simulations with embedded discrete fracture models generate training samples. Then, the samples are divided into training and testing samples to train these ML models and optimize network hyper parameters, respectively. Finally, the trained ML models are combined with an intelligent differential evolution algorithm to optimize the fracture parameters. This novel method has been applied to a naturally fractured reservoir model based on the real-field Barnett shale formation. The obtained results are compared with those of conventional optimizations with high-fidelity models. The results confirm the superiority of the proposed method owing to its very low computational cost. The use of ML modeling technology and an intelligent optimization algorithm could greatly contribute to simulation optimization and design, prompting progress in the intelligent development of unconventional oil and gas reservoirs in China.
基金Program of Science and Technology Department of Sichuan Province(2022YFS0541-02)Program of Heavy Rain and Drought-flood Disasters in Plateau and Basin Key Laboratory of Sichuan Province(SCQXKJQN202121)Innovative Development Program of the China Meteorological Administration(CXFZ2021Z007)。
摘要Machine learning models were used to improve the accuracy of China Meteorological Administration Multisource Precipitation Analysis System(CMPAS)in complex terrain areas by combining rain gauge precipitation with topographic factors like altitude,slope,slope direction,slope variability,surface roughness,and meteorological factors like temperature and wind speed.The results of the correction demonstrated that the ensemble learning method has a considerably corrective effect and the three methods(Random Forest,AdaBoost,and Bagging)adopted in the study had similar results.The mean bias between CMPAS and 85%of automatic weather stations has dropped by more than 30%.The plateau region displays the largest accuracy increase,the winter season shows the greatest error reduction,and decreasing precipitation improves the correction outcome.Additionally,the heavy precipitation process’precision has improved to some degree.For individual stations,the revised CMPAS error fluctuation range is significantly reduced.
基金This study has been reviewed and approved by the Clinical Research Ethics Committee of Wenzhou Central Hospital and the First Hospital Affiliated to Wenzhou Medical University,No.KY2024-R016.
摘要BACKGROUND Colorectal cancer significantly impacts global health,with unplanned reoperations post-surgery being key determinants of patient outcomes.Existing predictive models for these reoperations lack precision in integrating complex clinical data.AIM To develop and validate a machine learning model for predicting unplanned reoperation risk in colorectal cancer patients.METHODS Data of patients treated for colorectal cancer(n=2044)at the First Affiliated Hospital of Wenzhou Medical University and Wenzhou Central Hospital from March 2020 to March 2022 were retrospectively collected.Patients were divided into an experimental group(n=60)and a control group(n=1984)according to unplanned reoperation occurrence.Patients were also divided into a training group and a validation group(7:3 ratio).We used three different machine learning methods to screen characteristic variables.A nomogram was created based on multifactor logistic regression,and the model performance was assessed using receiver operating characteristic curve,calibration curve,Hosmer-Lemeshow test,and decision curve analysis.The risk scores of the two groups were calculated and compared to validate the model.RESULTS More patients in the experimental group were≥60 years old,male,and had a history of hypertension,laparotomy,and hypoproteinemia,compared to the control group.Multiple logistic regression analysis confirmed the following as independent risk factors for unplanned reoperation(P<0.05):Prognostic Nutritional Index value,history of laparotomy,hypertension,or stroke,hypoproteinemia,age,tumor-node-metastasis staging,surgical time,gender,and American Society of Anesthesiologists classification.Receiver operating characteristic curve analysis showed that the model had good discrimination and clinical utility.CONCLUSION This study used a machine learning approach to build a model that accurately predicts the risk of postoperative unplanned reoperation in patients with colorectal cancer,which can improve treatment decisions and prognosis.
摘要BACKGROUND Liver transplantation(LT)is a life-saving intervention for patients with end-stage liver disease.However,the equitable allocation of scarce donor organs remains a formidable challenge.Prognostic tools are pivotal in identifying the most suitable transplant candidates.Traditionally,scoring systems like the model for end-stage liver disease have been instrumental in this process.Nevertheless,the landscape of prognostication is undergoing a transformation with the integration of machine learning(ML)and artificial intelligence models.AIM To assess the utility of ML models in prognostication for LT,comparing their performance and reliability to established traditional scoring systems.METHODS Following the Preferred Reporting Items for Systematic Reviews and Meta-Analysis guidelines,we conducted a thorough and standardized literature search using the PubMed/MEDLINE database.Our search imposed no restrictions on publication year,age,or gender.Exclusion criteria encompassed non-English studies,review articles,case reports,conference papers,studies with missing data,or those exhibiting evident methodological flaws.RESULTS Our search yielded a total of 64 articles,with 23 meeting the inclusion criteria.Among the selected studies,60.8%originated from the United States and China combined.Only one pediatric study met the criteria.Notably,91%of the studies were published within the past five years.ML models consistently demonstrated satisfactory to excellent area under the receiver operating characteristic curve values(ranging from 0.6 to 1)across all studies,surpassing the performance of traditional scoring systems.Random forest exhibited superior predictive capabilities for 90-d mortality following LT,sepsis,and acute kidney injury(AKI).In contrast,gradient boosting excelled in predicting the risk of graft-versus-host disease,pneumonia,and AKI.CONCLUSION This study underscores the potential of ML models in guiding decisions related to allograft allocation and LT,marking a significant evolution in the field of prognostication.
摘要Background:With high colorectal cancer(CRC)incidence,accurate early differentiation of precancerous polyps is critical for prognosis;while the gold-standard colonoscopy-biopsy is limited by invasiveness,cost and poor scalability,and routine blood tests lack efficiency with simple indicator combinations,machine learning's feature-mining capacity offers a solution.This study aimed to develop and evaluate a machine learning system for differentiating patients with CRC and colorectal polyps using routine blood indices.Methods:A retrospective analysis was conducted on the clinical data of 284 patients with CRC and 79 patients with colorectal polyps who were diagnosed at the Chinese PLA General Hospital from October 2021 to February 2024.The extreme gradient boosting(XGBoost)algorithm was used to establish a machine learning model using demographic characteristics and routine blood indices.The Shapley additive explanation method was used to evaluate feature importance.Results:The constructed XGBoost model achieved high levels in differentiating CRC and colorectal polyps,with a precision of 0.906,a recall of 0.817,an accuracy of 0.791,and an area under the receiver operating characteristic curve of 0.869.The Shapley additive explanation showed that the top five important features were fibrinogen,carcinoembryonic antigen,plasma thrombin time,ferritin,and D-dimer.Conclusion:The XGBoost machine learning model based on blood indices has certain application value in the differential diagnosis of CRC and colorectal polyps,providing a new efficient tool for auxiliary diagnosis to assist clinical decision-making.
基金supported by the National Natural Science Foundation of China(U21A20290)。
摘要The sustainable development of nuclear energy requires efficient and reusable materials for uranium extraction from seawater,where uranyl species exist at ultra-trace levels within complex coordination environments.To overcome the limitations of empirical,trial-and-error design,this study establishes U-Predict v1.0,a structure-performance database for uranium adsorbents,and introduces an interpretable machine learning framework for adsorption performance prediction and mechanistic analysis.Trained on over 220 samples and 54 structural descriptors,the Light Gradient Boosting Machine(LightGBM)model with engineered features and optimized hyperparameters achieved a training R2 of 0.9887 and a test R2 of 0.7501,indicating robust predictive performance under heterogeneous and literature-derived conditions.SHapley Additive exPlanations(SHAP)interpretation revealed that structural and environmental variables,particularly functional group chemistry,solution pH,and surface area,jointly govern uranium the maximum adsorption capacity(qmax)through nonlinear interactions,emphasizing the dual control of material composition and adsorption environment.Experimental validation using a representative high-performance covalent organic frameworks-based adsorbent confirmed the model's predictive reliability,with the measured qmax of 342.2 mg g-1 showing reasonable agreement with the predicted value(449.3 mg g-1;23.8%deviation).This study establishes an explainable AI paradigm that bridges modeling and experimentation,providing a transferable foundation to support the intelligent discovery and sustainable design of next-generation uranium adsorbents for high-efficiency extraction.
基金financially supported by the Special Funds of the National Natural Science Foundation of China(32441047)the Major Project of Hubei Province Science&Technology(2023BBA002)+1 种基金the“Pioneer”and“Leading Goose”R&D Program of Zhejiang(2024SSYS0103)the Major Scientific and Technological Tasks of the Chinese Academy of Agricultural Sciences(CAAS-ZDRW202416)。
摘要ARC inoculant(A,aflatoxin prevention and control;R,Rhizobia nodulation induction;C,Coupling)is a brandnew inoculant with coupling function that enhances legume quality and nitrogen fixation.Comprehensive characterization of its key functional strains is critical for establishing a quality-control framework for the inoculant's formulation.Here,we constructed a characteristic spectral dataset comprising over 63,000 single-cell Raman spectra of the constituent strains by employing Ramanome technology.Six machine learning-based predictive models were developed and compared for the constituent strains,while the Linear Discriminant Analysis(LDA)model demonstrated the best performance,with a classification accuracy exceeding 92.4%.This work provides a unique spectral fingerprint for ARC inoculant and will directly aid its application in sustainable agricultural production.
基金the financial support from the Indian Council of Forestry Research and Education。
摘要Understanding soil depth is fundamental for evaluating a soil's ability to support vegetation,retain water and nutrients,sequester carbon,and sustain infrastructure.Despite its importance,the spatial variability of soil depth remains poorly understood,especially in mountainous terrains with limited sampling.Digital Soil Mapping(DSM),empowered by Machine Learning(ML)algorithms,offers a practical solution for estimating soil properties across complex landscapes.This study aimed to develop a predictive soil depth map for Nagaland,a hilly state in Northeastern India,using ML techniques.A dataset comprising soil depth observations from 70locations and 37 environmental covariates was analysed.Four ML algorithms Random Forest(RF),Support Vector Machine(SVM),Extreme Gradient Boosting(XGB),and K-Nearest Neighbours(KNN)were employed.Covariates derived from MODIS,Worldclim,and digital elevation models(DEM)were used to model soil-landscape relationships.Among all predictors,BIO15(precipitation seasonality)and NDVI(Normalized Difference Vegetation Index)emerged as key variables influencing soil depth variation.The soil depth predictions across models ranged between 44 and 158 cm,with XGB showing the best performance(R2=0.32±0.24;RMSE=22.63±4.74 cm;MAE=18.59±3.99 cm),followed by RF,while SVM and KNN were comparatively less accurate.The spatial distribution revealed that deeper soils were concentrated in the central and northern plains and valleys,while shallower soils were predominantly found in the steep,dissected Southeastern Hill Region.The findings highlight the efficacy of ML-driven DSM in mountainous,data-scarce environments,offering a valuable tool for land use planning,soil conservation,and ecological management in the Eastern Himalayas.
摘要Over the past decades,surrogate model-aided reliability analysis approaches grounded in active learning have undergone extensive development.However,Gaussian process models like Kriging suffer from severe computational burdens when handling high-dimensional problems or large samples.Conversely,machine learning algorithms such as extreme learning machines exhibit high computational efficiency but lack variance output and stability,making them difficult to employ for adaptive active learning strategies.To address these limitations,this study proposes a population Monte Carlo method based on an adaptive closed neuron extreme learning machine.First,a closed neuron strategy uses a consistency metric to screen and retain neurons containing the most informative features.This preserves the fast analytical solution advantage of extreme learning machines while significantly improving the reconstruction accuracy and stability of the true limit state surface.Second,to overcome the lack of variance in the output,an ensemble model is constructed.By calculating predictive mean and standard deviation,a learning function is formulated for efficient adaptive sample enrichment.Finally,utilizing the adaptive importance sampling mechanism of the population Monte Carlo framework,the auxiliary density function is optimized to progressively shift the sampling center toward high contribution failure regions.Four engineering examples confirm that the proposed method achieves exceptional computational efficiency and high accuracy for complex reliability analysis involving extremely small failure probabilities.
基金the Bridge to the Doctorate:Wisconsin Louis Stokes Alliance for Minority Participation National Science Foundation(NSF)award number HRD-1612530the University of Wisconsin-Madison Graduate Engineering Research Scholars(GERS)fellowship program,and the PPG Coating Innovation Center for financial support for the initial part of this work.The other authors gratefully acknowledge support from the NSF Collaborative Research:Framework:Machine Learning Materials Innovation Infrastructure award number 1931306+1 种基金Lane E.Schultz also acknowledges this award for support for the latter part of this work.Machine learning was performed with the computational resources provided by XSEDE 2.0:Integrating,Enabling and Enhancing National Cyberinfrastructure with Expanding Community Involvement Grant ACI-1548562We thank former and current members of the Informatics Skunkworks at the University of Wisconsin-Madison for their contributions to early aspects of this work:Angelo Cortez,Evelin Yin,Jodie Felice Ritchie,Stanley Tzeng,Avi Sharma,Linxiu Zeng,and Vidit Agrawal.
摘要Knowledge of the domain of applicability of a machine learning model is essential to ensuring accurate and reliable model predictions.In this work,we develop a new and general approach of assessing model domain and demonstrate that our approach provides accurate and meaningful domain designation across multiple model types and material property data sets.Our approach assesses the distance between data in feature space using kernel density estimation,where this distance provides an effective tool for domain determination.We show that chemical groups considered unrelated based on chemical knowledge exhibit significant dissimilarities by our measure.We also show that high measures of dissimilarity are associated with poor model performance(i.e.,high residual magnitudes)and poor estimates of model uncertainty(i.e.,unreliable uncertainty estimation).Automated tools are provided to enable researchers to establish acceptable dissimilarity thresholds to identify whether new predictions of their own machine learning models are in-domain versus out-of-domain.