tRNA-derived small RNAs(tsRNAs),as a class of regulatory small noncoding RNA,have been implicated in a wide variety of human diseases.Large amounts of tsRNA–disease associations have been identified in recent years f...tRNA-derived small RNAs(tsRNAs),as a class of regulatory small noncoding RNA,have been implicated in a wide variety of human diseases.Large amounts of tsRNA–disease associations have been identified in recent years from accumulating studies.However,repositories for cataloging the detailed information on tsRNA–disease associations are scarce.In this study,we provide a tsRNADisease database by integrating experimentally and computationally supported tsRNA–disease associations from manual curation of literatures and other related resources.tsRNADisease contains 5571 manually curated associations between 4759 tsRNAs and 166 diseases with experimental evidence from 346 studies.In addition,it also contains 5013 predicted associations between 1297 tsRNAs and 111 diseases.tsRNADisease provides a user-friendly interface to browse,retrieve,and download data conveniently.This database can improve our understanding of tsRNA deregulation in diseases and serve as a valuable resource for investigating the mechanism of disease-related tsRNAs.tsRNADisease is freely available at http://gffzz9c504e06f78b4edahonvwv5bwfxx96okp.ffgz.tsg.suse.edu.cn.展开更多
Data assimilation algorithms have been demonstrated to increase the accuracy of predictions in airfoil flow fields.However,slight changes in airfoil geometry and Reynolds number(Re)variations could lead to differences...Data assimilation algorithms have been demonstrated to increase the accuracy of predictions in airfoil flow fields.However,slight changes in airfoil geometry and Reynolds number(Re)variations could lead to differences in aerodynamic characteristics and stall behavior,consequently affecting assimilation outcomes.Hence,this research uses the ensemble Kalman filter(EnKF)algorithm.The aerodynamic characteristics of two wind turbine airfoils obtained through wind tunnel experiments were investigated under varying degrees of stall by recalibrating the constants in the(S-A)model.The impacts of the airfoil thickness,Re variation,and Gurney flap installation on the assimilation results were subsequently examined.Verifying the applicability of the constants obtained via data assimilation under varying conditions might offer opportunities to reduce the demand for computational resources.The assimilation results indicate that at a Re on the order of magnitude of 105,the original model tends to delay flow separation as the Re increases.Consequently,the recalibrated constant Cb1 generally decreases with increasing Re.Despite belonging to the same airfoil family,discrepancies in the flow separation behavior predicted by the original model resulted in variations in the recalibrated constants.The constants derived from the thinner airfoil induce premature flow separation in the thicker YA-30 airfoil under stall conditions.When assimilated constants are applied to flow field calculations under analogous stall conditions,constants from another condition may demonstrate an optimization effect and substitute the self-assimilated constants,provided that simulations using default constants for both conditions consistently exhibit an experimental separation trend.However,practical implementation requires caution due to the risk of overadjustment.展开更多
The distribution and transport of carbon dioxide(CO2)in the middle and upper atmosphere are closely linked to atmospheric dynamical processes,but the influence of planetary waves on CO2 transport remains unclear...The distribution and transport of carbon dioxide(CO2)in the middle and upper atmosphere are closely linked to atmospheric dynamical processes,but the influence of planetary waves on CO2 transport remains unclear.This study aims to investigate the impact of quasi-2-day waves(Q2DWs)on CO2 transport in the stratosphere-mesosphere during the sudden stratospheric warming(SSW)period.On the basis of data assimilation(DA)using the Whole Atmosphere Community Climate Model(WACCM)+Next-generation Ensemble Data Assimilation System(NEDAS)from December 2018 to February 2019,we used the transformed Eulerian mean framework to conduct a diagnostic analysis on Q2DW propagation and amplification.Additionally,we identified unstable regions in the stratosphere between 60°N and 80°N.This configuration facilitates the amplification of Q2DWs through enhanced baroclinic and barotropic instabilities.The results demonstrate that the dynamic features exhibit distinct propagation and amplification characteristics.We also found distinct Q2DWs in CO2 at the middle latitudes of the southern hemisphere,which were mainly induced by the Q2DWs in temperature.Further analysis showed that the anomalous vertical motion during the SSW period could lead to the enhancement of CO2 vertical gradients,which also contributed to the strengthening of Q2DWs in CO2.展开更多
Intelligent Connected Vehicles(ICVs)generate massive heterogeneous multi-modal data during operation,and due to the limited computing resources on board,graded data encryption protection is of great significance for b...Intelligent Connected Vehicles(ICVs)generate massive heterogeneous multi-modal data during operation,and due to the limited computing resources on board,graded data encryption protection is of great significance for balancing data security and efficient utilization.However,the current data grading processes struggle to address the evolving inference attacks and dynamic operational environments,and traditional grading approaches relying on static expert judgment or information-theoretic metrics.To bridge this gap,this paper proposes a novel inference strength-driven data grading framework,where inference strength quantifies the susceptibility of one dataset to infer another through adversarial reasoning.The framework employs a systematic methodology combining graph theory,optimization,and Large Language Models to construct an inference library and calculate inference strength.The framework also provides a PageRank-based algorithm to generate interpretable data grading lists for both static policy and vehicle-end application,prioritizing core data protection while respecting computational constraints.Validated on the Audi A2D2 dataset and real vehicle controller,our approach demonstrates improved protection utility compared to default grading baselines.The results highlight its potential to enhance data security in ICVs through prioritized protection of core data under computational constraints.展开更多
The full waveform inversion(FWI)utilizes full wave field data to invert subsurface parameters and is considered one of the most promising data-driven tools for obtaining high precision velocity models.However,the succ...The full waveform inversion(FWI)utilizes full wave field data to invert subsurface parameters and is considered one of the most promising data-driven tools for obtaining high precision velocity models.However,the successful application of FWI in geophysical explo ration remains limited,primarily due to the cycle-skipping issue caused by the absence of low-frequency data,which is one of the main reasons for FWI failures.Incorporating prior regularization constraints FWI can effectively compensate for the lacking low-frequency components and constrain the iterative updates of FWI toward the desired direction,offering a natural advantage in addressing this challenge.However,the weights of the prior information terms are still determined empirically,which introduces significant subjectivity and randomness to the inversion results.To solve this issue,we propose an adaptive method to determine the weight factor based on posterior probability distribution within the Bayesian theoretical framework.This factor adaptively adjusts during each iteration to balance the contributions of the data error term and the prior information term in FWI,which can effectively mitigate the cycle-skipping problem and alleviating the nonlinearity of the inversion process.Numerical examples from the Overthrust model and the Marmousi model show that our method not only enhance the accuracy of FWI,but also demonstrate strong noise resistance.展开更多
With continuous advancement in geological studies,three-dimensional(3D)geological modeling technology based on big data and artificial intelligence(AI)has become a prominent focus in the interdisciplinary field of ear...With continuous advancement in geological studies,three-dimensional(3D)geological modeling technology based on big data and artificial intelligence(AI)has become a prominent focus in the interdisciplinary field of earth and information sciences.Through the analysis and comparison of existing 3D modeling methods,this study introduces a high-precision 3D geological modeling approach.The new method leverages advanced computing technologies,including multisource heterogeneous data processing,integrated model databases,seamless splicing of local models,and cluster analysis.Furthermore,it enables unified 3D visualization of subsurface and surface conditions,providing new insights into disaster prevention,mitigation,and intelligent mineral exploration.To validate its practicality,this study conducts 3D geological modeling of the X area within the Sichuan Basin,China.The data sources include remote sensing images,geological maps,geophysical data,borehole data,and X-ray fluorescence(XRF)spectroscopy data.Preliminary exploration in the X area has successfully identified new mineralization belts,verifying the feasibility and effectiveness of the new 3D geological modeling method that integrates big data processing and intelligent techniques.展开更多
This study investigates the pivotal role of data clustering in both data science and management,focusing on core methodologies,tools,and diverse applications.It examines traditional clustering techniques such as parti...This study investigates the pivotal role of data clustering in both data science and management,focusing on core methodologies,tools,and diverse applications.It examines traditional clustering techniques such as partitional and hierarchical methods,alongside more advanced approaches,including data stream,density-based,graphbased,and model-based clustering,which are essential for processing complex and structured datasets.The study highlights fundamental principles,presents commonly adapted tools and frameworks,outlines the clustering workflow within data science,and discusses major implementation challenges.Beyond technical applications,this study emphasizes how clustering supports managerial tasks and decision-making through a comprehensive survey of recent literature.By bridging analytical techniques with real-world business needs,clustering remains an essential tool in both data science and management.The study concludes by outlining future research directions,underscoring the role of clustering in driving innovation and enabling informed strategic and operational decisions.展开更多
Effective fault diagnosis is crucial for the reliable running of Electromechanical coupling Systems(EMS),yet hampered by insufficient entity fault data.Digital Twin(DT)technology offers the potential for virtual fault...Effective fault diagnosis is crucial for the reliable running of Electromechanical coupling Systems(EMS),yet hampered by insufficient entity fault data.Digital Twin(DT)technology offers the potential for virtual fault data augmentation and fault diagnosis improvement.However,there is still a lack of an effective and systematic methodology,to decouple complicated EMS entities and construct their full-system DT.To address this,a hierarchical collaborative DT construction framework is proposed for fault data augmentation of EMS.Specially,we decouple EMS entity into the triplet representations of element,data,and relationship,which establish the profound understanding of coupling characteristics from multiple modalities.Furthermore,we develop a hierarchical DT modeling method to mirror these complicated couplings as four-level sub-DTs of space,behavior,process,and status.Each level of sub-DT utilizes the data-mechanism combined technique to balance modeling adaptability and precision.Finally,these heterogeneous sub-DTs are integrated as full-system DT driven by collaborative orchestration algorithm,which achieves the global consistency mirror with the real fault manifestation under diverse fault modes.Experiments on a multi-coupled electromechanical fault test bench validate our framework.Results exhibit the average improvements of 17.29%and 9.97%in accuracy of data augmentation fault classification,confirming its superiority and effectiveness.展开更多
Airborne Mobile Networks(AMNs)require high-precision situational awareness in enclosed spaces for critical tasks like border surveillance and maritime monitoring.Compared to vision and wearable technologies,WiFi signa...Airborne Mobile Networks(AMNs)require high-precision situational awareness in enclosed spaces for critical tasks like border surveillance and maritime monitoring.Compared to vision and wearable technologies,WiFi signals in AMN environments are a promising sensing medium due to their non-intrusiveness,cost-effectiveness,and privacy benefits.However,the development of WiFi-based sensing in AMNs is hindered by the scarcity of high-quality large-scale data.While Data Augmentation(DA)can address this scarcity,traditional time-series DA may distort WiFi Channel State Information(CSI)'s time–frequency properties,and deep generative models suffer from mode collapse,spectral distortion,and high computational costs.To overcome these limitations,we propose OT-ADG,an adaptive data generation algorithm based on Optimal Transport(OT)theory to synthesize high-fidelity WiFi sensing samples.It dynamically models subcarrierspecific energy distributions and computes optimal transport plans using Sinkhorn's algorithm with entropy regularization.Furthermore,we present ViFi,the first large-scale and scenario-rich WiFi sensing dataset to fill the gap in AMN applications,encompassing 26 real-world scenarios with 20640 CSI samples and synchronized videos.Experimental results show OT-ADG's superior performance,with maximum improvements of up to 14.88%over baseline recognition methods,while outperforming existing DA approaches.Moreover,the robust performance of multiple WiFi-based HAR models validates the effectiveness of the ViFi dataset.展开更多
The Ok null test can not only assess whether the cosmic curvature is zero—thereby,if true,reducing degeneracies between cosmic curvature and other cosmological parameters—but also provide a model-independent chec...The Ok null test can not only assess whether the cosmic curvature is zero—thereby,if true,reducing degeneracies between cosmic curvature and other cosmological parameters—but also provide a model-independent check of compatibility between different data sets.However,traditional implementations often require absolute distance data from Type Ia supernovae(SNe Ia)or baryon acoustic oscillation(BAO)measurements,limiting their applicability because such absolute distance data are usually not accessible.The BAO Alcock-Paczynski(AP)parameter FAP is a measurement of a distance ratio,making the Dark Energy Spectroscopic Instrument(DESI)AP measurements particularly well suited for the Oknull test,as no absolute distance measurements are required.We propose a novel null test of cosmic curvature tailored to DESI BAO data that combines FAPwith ratios such as D′V/DVor D′M/DM.Crucially,this construction eliminates the need for absolute distance measurements.We further develop multi-task Gaussian processes to perform the null test.This approach can also be applied to a joint DESI BAO and SNe Ia dataset,and we find that DESI BAO and SNe Ia data are compatible.Although there is~2σ evidence of nonzero curvature at low redshift z■0.5,this result is not conclusive,largely due to the lack of observational data in the corresponding redshift range.展开更多
Real-world studies(RWSs)have emerged as a transformative force in oncology research,complementing traditional randomized controlled trials(RCTs)by providing comprehensive insights into cancer care within routine clini...Real-world studies(RWSs)have emerged as a transformative force in oncology research,complementing traditional randomized controlled trials(RCTs)by providing comprehensive insights into cancer care within routine clinical settings.This review examines the evolving landscape of RWSs in oncology,focusing on their implementation,methodological considerations,and impact on precision medicine.We systematically analyze how RWSs leverage diverse data sources,including electronic health records(EHRs),insurance claims,and patient registries,to generate evidence that bridges the gap between controlled clinical trials and real-world clinical practice.The review underscores the key contributions of RWSs,including capturing therapeutic outcomes in traditionally underrepresented populations,expanding drug indications,and evaluating long-term safety and effectiveness in routine clinical settings.While acknowledging significant challenges,including data quality variability and privacy concerns,we discuss how emerging technologies like artificial intelligence are helping to address these limitations.The integration of RWSs with traditional clinical research is revolutionizing the paradigm of precision oncology and enabling more personalized treatment approaches based on real-world evidence.展开更多
This study presents a method to correct the lithology of mud-logging profile with logging data based on neural network,which aims to solve the problems of time-consuming,high labor intensity and great infl uence of hu...This study presents a method to correct the lithology of mud-logging profile with logging data based on neural network,which aims to solve the problems of time-consuming,high labor intensity and great infl uence of human factors in the process of traditional lithology correction of mud-logging profi le.Firstly,the lithology of mud-logging profi le is processed by digital technology and converted into digital curve which is consistent with the logging sampling interval,and the logging lithology curve is calculated by using the optimal logging method.Then,combining automatic depth-correction technology with manual correction methods,the lithology of mud-logging profi le is corrected for depth.On the basis of lithology depth-correction of mudlogging profile,the multi-layer perceptron(MLP)neural network is used to learn logging data and realize accurate identification of multiple lithologies,so as to construct a high-precision logging profile lithology curve and provide accurate basis for lithology correction of mud-logging profile.The effectiveness and accuracy of the proposed method are verifi ed by practical application cases.The corrected lithology of mudlogging profi le is highly consistent with the lithology of logging profi le,which provides a solid foundation for subsequent geological interpretation,reservoir evaluation and oil and gas resource assessment.This study not only improves the effi ciency of mud-logging data processing,but also ensures that the needs of exploration and exploitation work are met in a timely manner,which has important theoretical signifi cance and application value.展开更多
The discovery of topological materials has advanced rapidly due to high-throughput computation and machine learning,but research progress is hampered by inconsistent classification standards and fragmented data resour...The discovery of topological materials has advanced rapidly due to high-throughput computation and machine learning,but research progress is hampered by inconsistent classification standards and fragmented data resources.Existing databases differ in computational methods,material coverage,and labeling criteria,making it difficult to compare findings across studies.To overcome these challenges,we present a unified topological materials dataset that systematically combines and reconciles two major databases:Materiae and the Topological Materials Database.This dataset provides consistent topological classifications for 35608 materials,accessible through the Materials Galaxy platform for interactive exploration and available for bulk download via MatElab.We describe the featurization methodology that converts crystal structures into 4710 machine-learning-ready descriptors and present a comprehensive analysis of topological material distributions.This work serves as a complete guide for accessing,utilizing,and interpreting this unified resource,designed to enable reproducible machine learning applications and accelerate the discovery of topological materials.展开更多
In this study,buoy measurements of sea surface wind fields(10 m above the sea surface)were collected at 12 stations in the Bohai and Yellow Seas over a period of 11 years(2011-2021).An evaluation was conducted of the ...In this study,buoy measurements of sea surface wind fields(10 m above the sea surface)were collected at 12 stations in the Bohai and Yellow Seas over a period of 11 years(2011-2021).An evaluation was conducted of the 10 m wind performance of the Climate Forecast System Reanalysis version 2(CFSv2),Modern-Era Retrospective Analysis for Research and Applications version 2(MERRA2),and the fifth-generation European Centre for Medium-Range Weather Forecasts reanalysis products(ERA5)in reproducing wind fields over the Bohai and Yellow Seas.The mean BIAS,root mean square error(RMSE),and correlation coefficient(R)were calculated and used as performance indicators.All three reanalysis products demonstrate commendable performance.The correlation coefficients for wind speed and wind direction exceed 0.83 and 0.90,respectively.Of the three products,the fifth-generation European Centre for Medium-Range Weather Forecasts reanalysis demonstrates the highest level of performance.The model demonstrated the lowest average BIAS and RMSE for wind speed,and the highest R for wind direction.The findings indicate that product performance is subject to variation in both season and wind speed.展开更多
To overcome the limitations of ground-based telescopes in spatial resolution and imaging quality, recent researchhas concentrated on image super-resolution (SR) reconstruction. This approach aims to enhance ground-bas...To overcome the limitations of ground-based telescopes in spatial resolution and imaging quality, recent researchhas concentrated on image super-resolution (SR) reconstruction. This approach aims to enhance ground-basedimages by recovering “space-based-like” images with higher spatial resolution and more detailed structuralinformation, without the need for expensive space-based observations. The training of SR models typically relieson large-scale, high-quality paired datasets of ground-based and space-based images. However, the acquisition ofsuch data is highly costly, which significantly limits the widespread application of these models. Based on this,this paper proposes using realistic synthetic galaxy images generated by IllustrisTNG cosmological simulation toreplace real space-based images and constructs a training dataset for image SR tasks, TNG-RealSR. To verify thevalidity of this dataset, we selected five mainstream lightweight SR models for evaluation. The results show thatTNG-RealSR exhibits good applicability in the task of galaxy image SR, providing reliable data support forrelated research. To the best of our knowledge, this work is the first to apply cosmological simulation-generatedsynthetic galaxy images to the field of galaxy image SR, demonstrating the feasibility of deep learning methodsbased on synthetic data for astronomical image enhancement tasks. This provides a low-cost and efficient solutionfor improving data quality in future large-scale surveys. The dataset is available online at http://gffzz188fe103f8f1460asonvwv5bwfxx96okp.ffgz.tsg.suse.edu.cn/jiaweimmiao/TNG-RealSR.展开更多
Speech Emotion Recognition(SER)is a critical component of affective computing with broad applications in human–computer interaction,mental health monitoring,and intelligent multimedia systems.However,SER remains chal...Speech Emotion Recognition(SER)is a critical component of affective computing with broad applications in human–computer interaction,mental health monitoring,and intelligent multimedia systems.However,SER remains challenging due to the emotional ambiguity,lack of labeled data,class imbalance,and speaker variability.This study presents an effective SER framework that integrates contrastive representation learning,optimized spectrogram-based data augmentation,and selective synthetic data generation by using TimeGAN to enhance emotion classification performance.Contrastive learning enables the model to better discriminate acoustically similar emotions while Optuna automatically tunes augmentation strategies such as noise injection,time shifting,and time-frequency masking.Unlike existing approaches that apply synthetic generation uniformly across all classes,the proposed method targets only confusing or under-represented emotion classes to preserve the inter-class separability.A CNN-BiLSTM architecture is used to extract spectral and temporal information of the speech.The framework is evaluated with benchmark SER datasets—EMO-DB and RAVDESS—under speaker independent protocols.Experimental results demonstrate improved accuracy,robustness,and generalization under limited and imbalanced data conditions,supported by confusion matrices,UMAP,and t-SNE visualizations.展开更多
Timely and accurate forecasting of crop yields is critical for food management and trade.However,only limited research has explored the impact of integrating crop phenotypic parameters(CPPs)with unmanned aerial vehicl...Timely and accurate forecasting of crop yields is critical for food management and trade.However,only limited research has explored the impact of integrating crop phenotypic parameters(CPPs)with unmanned aerial vehicle(UAV)data across different phenological stages on maize yield prediction.The extent to which multi-temporal data enhances the accuracy and reliability of yield projections compared to mono-temporal data has yet to be systematically investigated.To attain the proper balance between accuracy and cost in crop yield estimation,this study proposed a structured framework for identifying the optimal phenological periods for summer maize yield prediction using UAV-based multispectral data.Three classical methods of custom mean decrease accuracy(C-MDA),optimal parameters-based geographical detector(OPGD),and grey relational analysis(GRA)were first used to sort and screen both the CPPs and vegetation indices(VIs)derived from UAV-based information over six growth stages.Ridge regression models based on multi-temporal data combinations and mono-temporal data were established separately,and their performance in yield prediction were compared to identify the optimal phenological stages and the corresponding key factors.Our results showed that C-MDA was much better at factor screening and ranking compared to OPGD and GRA.The green normalized difference vegetation index(GNDVI),normalized difference vegetation index(NDVI),and normalized difference red edge index(NDRE)emerged as the topperforming VIs,while the leaf area index(LAI)and above ground biomass(AGB)proved to be the most effective CPPs.When predicting yield using only mono-temporal data,the dough stage delivered the highest predictive accuracy(R2=0.871,RMSE=0.407 t ha-1),while the tasseling stage was the earliest that achieved yield estimates with acceptable precision(R2=0.810,RMSE=0.493 t ha-1).In contrast,the integration of UAV data from different crop growth stages markedly enhanced the accuracy of yield estimation.Combinations of data from the tasseling,silking,and dough stages were recommended as the best option(R2=0.942,RMSE=0.291 t ha-1).These findings indicate that the precise estimation of maize yields in smallholder fields may be attainable,and present both substantial theoretical insights and practical benefits for the advancement of precision agriculture.展开更多
Training software models for crop disease diagnosis requires large image datasets to achieve high accuracy.We describe a lesion information transfer diffusion model,LesionDiff,for generating image data that augments a...Training software models for crop disease diagnosis requires large image datasets to achieve high accuracy.We describe a lesion information transfer diffusion model,LesionDiff,for generating image data that augments a real-world disease lesion image dataset.An information preprocessing module identifies lesion areas on leaves,an enhancement module captures diverse visual and semantic lesion features,and a generation module fills missing regions in masked disease images by synthesizing lesion phenotypes.This augmentation increased the average diagnostic accuracy of a test dataset by more than 3%.展开更多
The traditional method of performance degradation prediction and maintenance of rolling bearings only considers a single sensor signal,which makes it difficult to automatically partition degradation stages and prone t...The traditional method of performance degradation prediction and maintenance of rolling bearings only considers a single sensor signal,which makes it difficult to automatically partition degradation stages and prone to over-detection.A new method of performance degradation evaluation and maintenance of rolling bearings based on data-level fusion,adaptive health state partitioning,and state maintenance is proposed.Firstly,considering the degradation and impact in the process of bearing deterioration,the multi-sensor signals are dynamically weighted to achieve data-level fusion.Secondly,a bearing health index was established based on fast spectral correlation,Wasserstein distance,and linear rectification techniques.On this basis,by combining the Bayesian information criterion and the elbow rule,the precise division of rolling bearing health state is realized through hidden Markov model regression.Then,random forest was used to classify and predict the data to verify the validity of the proposed data fusion method and health indicator.Finally,condition-based maintenance strategy based on the fourth moment,stress-strength interference model,and Gamma process is proposed to avoid excessive detection and reduce maintenance costs.Through accelerated degradation experiments and field validation tests on the rolling bearing test data set of Xi’an Jiaotong University and FEMTO(PRONOSTIA),the accuracy and superiority of the proposed method in the prediction and maintenance of bearing health state are verified.展开更多
基金supported by the National Natural Science Foundation of China(91959106)the Foundation of the Shanghai Municipal Education Commission(24RGZNC02)+4 种基金Shanghai Key Laboratory of Intelligent Information Processing,Fudan University(IIPL-2025-RD3-02)Key University Science Research Project of Anhui Province(2023AH030108)Climbing Peak Training Program for Innovative Technology team of Yijishan Hospital,Wannan Medical College(PF201904)Peak Training Program for Scientific Research of Yijishan Hospital,Wannan Medical College(GF2019G15)the talent project of the First Affiliated Hospital of Wannan Medical College(Yijishan Hospital of Wannan Medical College)(YR202422).
摘要tRNA-derived small RNAs(tsRNAs),as a class of regulatory small noncoding RNA,have been implicated in a wide variety of human diseases.Large amounts of tsRNA–disease associations have been identified in recent years from accumulating studies.However,repositories for cataloging the detailed information on tsRNA–disease associations are scarce.In this study,we provide a tsRNADisease database by integrating experimentally and computationally supported tsRNA–disease associations from manual curation of literatures and other related resources.tsRNADisease contains 5571 manually curated associations between 4759 tsRNAs and 166 diseases with experimental evidence from 346 studies.In addition,it also contains 5013 predicted associations between 1297 tsRNAs and 111 diseases.tsRNADisease provides a user-friendly interface to browse,retrieve,and download data conveniently.This database can improve our understanding of tsRNA deregulation in diseases and serve as a valuable resource for investigating the mechanism of disease-related tsRNAs.tsRNADisease is freely available at http://gffzz9c504e06f78b4edahonvwv5bwfxx96okp.ffgz.tsg.suse.edu.cn.
基金supported by the Natural Science Foundation of Jiangsu Higher Education Institutions of China(Grant No.25KJB480015)the Qing Lan Project of Jiangsu Higher Education Institutions+2 种基金the China Postdoctoral Science Foundation(Grant No.2023M742958)the Excellent Doctor of Yangzhou“Lvyang Jinfeng Plan”(Grant No.YZLYJFJH2021YXNS132)the Philosophy and Social Science Project of Jiangsu Provincial Education Department(Grant No.2025SJYB1556)。
摘要Data assimilation algorithms have been demonstrated to increase the accuracy of predictions in airfoil flow fields.However,slight changes in airfoil geometry and Reynolds number(Re)variations could lead to differences in aerodynamic characteristics and stall behavior,consequently affecting assimilation outcomes.Hence,this research uses the ensemble Kalman filter(EnKF)algorithm.The aerodynamic characteristics of two wind turbine airfoils obtained through wind tunnel experiments were investigated under varying degrees of stall by recalibrating the constants in the(S-A)model.The impacts of the airfoil thickness,Re variation,and Gurney flap installation on the assimilation results were subsequently examined.Verifying the applicability of the constants obtained via data assimilation under varying conditions might offer opportunities to reduce the demand for computational resources.The assimilation results indicate that at a Re on the order of magnitude of 105,the original model tends to delay flow separation as the Re increases.Consequently,the recalibrated constant Cb1 generally decreases with increasing Re.Despite belonging to the same airfoil family,discrepancies in the flow separation behavior predicted by the original model resulted in variations in the recalibrated constants.The constants derived from the thinner airfoil induce premature flow separation in the thicker YA-30 airfoil under stall conditions.When assimilated constants are applied to flow field calculations under analogous stall conditions,constants from another condition may demonstrate an optimization effect and substitute the self-assimilated constants,provided that simulations using default constants for both conditions consistently exhibit an experimental separation trend.However,practical implementation requires caution due to the risk of overadjustment.
基金supported by the National Natural Science Foundation of China(Grant Nos.42374195,42404168,and 42188101)a fellowship from the China National Postdoctoral Program for Innovative Talents(Grant No.BX20230273)+1 种基金the Hubei Provincial Natural Science Foundation of China(Grant No.2024AFB097)the Postdoctoral Project of Hubei Province(Grant No.2024HBBHCXA054).
摘要The distribution and transport of carbon dioxide(CO2)in the middle and upper atmosphere are closely linked to atmospheric dynamical processes,but the influence of planetary waves on CO2 transport remains unclear.This study aims to investigate the impact of quasi-2-day waves(Q2DWs)on CO2 transport in the stratosphere-mesosphere during the sudden stratospheric warming(SSW)period.On the basis of data assimilation(DA)using the Whole Atmosphere Community Climate Model(WACCM)+Next-generation Ensemble Data Assimilation System(NEDAS)from December 2018 to February 2019,we used the transformed Eulerian mean framework to conduct a diagnostic analysis on Q2DW propagation and amplification.Additionally,we identified unstable regions in the stratosphere between 60°N and 80°N.This configuration facilitates the amplification of Q2DWs through enhanced baroclinic and barotropic instabilities.The results demonstrate that the dynamic features exhibit distinct propagation and amplification characteristics.We also found distinct Q2DWs in CO2 at the middle latitudes of the southern hemisphere,which were mainly induced by the Q2DWs in temperature.Further analysis showed that the anomalous vertical motion during the SSW period could lead to the enhancement of CO2 vertical gradients,which also contributed to the strengthening of Q2DWs in CO2.
基金supported by National Key R&D Program of China(No.2022YFB3104900)the National Natural Science Foundation of China(No.U22A202101)State Key Laboratory of Intelligent Vehicle Safety Technology program(No.IVSTSKL-202439)。
摘要Intelligent Connected Vehicles(ICVs)generate massive heterogeneous multi-modal data during operation,and due to the limited computing resources on board,graded data encryption protection is of great significance for balancing data security and efficient utilization.However,the current data grading processes struggle to address the evolving inference attacks and dynamic operational environments,and traditional grading approaches relying on static expert judgment or information-theoretic metrics.To bridge this gap,this paper proposes a novel inference strength-driven data grading framework,where inference strength quantifies the susceptibility of one dataset to infer another through adversarial reasoning.The framework employs a systematic methodology combining graph theory,optimization,and Large Language Models to construct an inference library and calculate inference strength.The framework also provides a PageRank-based algorithm to generate interpretable data grading lists for both static policy and vehicle-end application,prioritizing core data protection while respecting computational constraints.Validated on the Audi A2D2 dataset and real vehicle controller,our approach demonstrates improved protection utility compared to default grading baselines.The results highlight its potential to enhance data security in ICVs through prioritized protection of core data under computational constraints.
基金partly supported by National Natural Science Foundation of China(42274154)the Fund of State Key Laboratory of Deep Oil and Gas,China University of Petroleum(East China),China(SKLDOG2024-ZYTS-03)。
摘要The full waveform inversion(FWI)utilizes full wave field data to invert subsurface parameters and is considered one of the most promising data-driven tools for obtaining high precision velocity models.However,the successful application of FWI in geophysical explo ration remains limited,primarily due to the cycle-skipping issue caused by the absence of low-frequency data,which is one of the main reasons for FWI failures.Incorporating prior regularization constraints FWI can effectively compensate for the lacking low-frequency components and constrain the iterative updates of FWI toward the desired direction,offering a natural advantage in addressing this challenge.However,the weights of the prior information terms are still determined empirically,which introduces significant subjectivity and randomness to the inversion results.To solve this issue,we propose an adaptive method to determine the weight factor based on posterior probability distribution within the Bayesian theoretical framework.This factor adaptively adjusts during each iteration to balance the contributions of the data error term and the prior information term in FWI,which can effectively mitigate the cycle-skipping problem and alleviating the nonlinearity of the inversion process.Numerical examples from the Overthrust model and the Marmousi model show that our method not only enhance the accuracy of FWI,but also demonstrate strong noise resistance.
基金supported by a project of the China Geological Survey(DD20240126).
摘要With continuous advancement in geological studies,three-dimensional(3D)geological modeling technology based on big data and artificial intelligence(AI)has become a prominent focus in the interdisciplinary field of earth and information sciences.Through the analysis and comparison of existing 3D modeling methods,this study introduces a high-precision 3D geological modeling approach.The new method leverages advanced computing technologies,including multisource heterogeneous data processing,integrated model databases,seamless splicing of local models,and cluster analysis.Furthermore,it enables unified 3D visualization of subsurface and surface conditions,providing new insights into disaster prevention,mitigation,and intelligent mineral exploration.To validate its practicality,this study conducts 3D geological modeling of the X area within the Sichuan Basin,China.The data sources include remote sensing images,geological maps,geophysical data,borehole data,and X-ray fluorescence(XRF)spectroscopy data.Preliminary exploration in the X area has successfully identified new mineralization belts,verifying the feasibility and effectiveness of the new 3D geological modeling method that integrates big data processing and intelligent techniques.
摘要This study investigates the pivotal role of data clustering in both data science and management,focusing on core methodologies,tools,and diverse applications.It examines traditional clustering techniques such as partitional and hierarchical methods,alongside more advanced approaches,including data stream,density-based,graphbased,and model-based clustering,which are essential for processing complex and structured datasets.The study highlights fundamental principles,presents commonly adapted tools and frameworks,outlines the clustering workflow within data science,and discusses major implementation challenges.Beyond technical applications,this study emphasizes how clustering supports managerial tasks and decision-making through a comprehensive survey of recent literature.By bridging analytical techniques with real-world business needs,clustering remains an essential tool in both data science and management.The study concludes by outlining future research directions,underscoring the role of clustering in driving innovation and enabling informed strategic and operational decisions.
基金supported by the National Key R&D Programof China under Grant STI 2030—Major Projects(No.2021ZD0201300)the National Natural Science Foundation of China(Nos.52472442,72471013)+1 种基金the Zhejiang Provincial Natural Science Foundation(Grant No.LMS26E050036)the Research Start-up Funds of Hangzhou International Innovation Institute of Beihang University,China(Nos.2024KQ069,2024KQ036,2024KQ035 and 2025BKZ055)。
摘要Effective fault diagnosis is crucial for the reliable running of Electromechanical coupling Systems(EMS),yet hampered by insufficient entity fault data.Digital Twin(DT)technology offers the potential for virtual fault data augmentation and fault diagnosis improvement.However,there is still a lack of an effective and systematic methodology,to decouple complicated EMS entities and construct their full-system DT.To address this,a hierarchical collaborative DT construction framework is proposed for fault data augmentation of EMS.Specially,we decouple EMS entity into the triplet representations of element,data,and relationship,which establish the profound understanding of coupling characteristics from multiple modalities.Furthermore,we develop a hierarchical DT modeling method to mirror these complicated couplings as four-level sub-DTs of space,behavior,process,and status.Each level of sub-DT utilizes the data-mechanism combined technique to balance modeling adaptability and precision.Finally,these heterogeneous sub-DTs are integrated as full-system DT driven by collaborative orchestration algorithm,which achieves the global consistency mirror with the real fault manifestation under diverse fault modes.Experiments on a multi-coupled electromechanical fault test bench validate our framework.Results exhibit the average improvements of 17.29%and 9.97%in accuracy of data augmentation fault classification,confirming its superiority and effectiveness.
基金supported by the National Science Foundation of China(Nos.62027826,62502066)the Fundamental Research Funds for the Central Universities of China(No.DUT25RC(3)044)the Natural Science Foundation Project of Liaoning Province of China(No.2025-BS-0003)。
摘要Airborne Mobile Networks(AMNs)require high-precision situational awareness in enclosed spaces for critical tasks like border surveillance and maritime monitoring.Compared to vision and wearable technologies,WiFi signals in AMN environments are a promising sensing medium due to their non-intrusiveness,cost-effectiveness,and privacy benefits.However,the development of WiFi-based sensing in AMNs is hindered by the scarcity of high-quality large-scale data.While Data Augmentation(DA)can address this scarcity,traditional time-series DA may distort WiFi Channel State Information(CSI)'s time–frequency properties,and deep generative models suffer from mode collapse,spectral distortion,and high computational costs.To overcome these limitations,we propose OT-ADG,an adaptive data generation algorithm based on Optimal Transport(OT)theory to synthesize high-fidelity WiFi sensing samples.It dynamically models subcarrierspecific energy distributions and computes optimal transport plans using Sinkhorn's algorithm with entropy regularization.Furthermore,we present ViFi,the first large-scale and scenario-rich WiFi sensing dataset to fill the gap in AMN applications,encompassing 26 real-world scenarios with 20640 CSI samples and synchronized videos.Experimental results show OT-ADG's superior performance,with maximum improvements of up to 14.88%over baseline recognition methods,while outperforming existing DA approaches.Moreover,the robust performance of multiple WiFi-based HAR models validates the effectiveness of the ViFi dataset.
基金supported in part by the National Natural Science Foundation of China(Grant Nos.12588101,12535002,12175184,12433001,and 12205015)。
摘要The Ok null test can not only assess whether the cosmic curvature is zero—thereby,if true,reducing degeneracies between cosmic curvature and other cosmological parameters—but also provide a model-independent check of compatibility between different data sets.However,traditional implementations often require absolute distance data from Type Ia supernovae(SNe Ia)or baryon acoustic oscillation(BAO)measurements,limiting their applicability because such absolute distance data are usually not accessible.The BAO Alcock-Paczynski(AP)parameter FAP is a measurement of a distance ratio,making the Dark Energy Spectroscopic Instrument(DESI)AP measurements particularly well suited for the Oknull test,as no absolute distance measurements are required.We propose a novel null test of cosmic curvature tailored to DESI BAO data that combines FAPwith ratios such as D′V/DVor D′M/DM.Crucially,this construction eliminates the need for absolute distance measurements.We further develop multi-task Gaussian processes to perform the null test.This approach can also be applied to a joint DESI BAO and SNe Ia dataset,and we find that DESI BAO and SNe Ia data are compatible.Although there is~2σ evidence of nonzero curvature at low redshift z■0.5,this result is not conclusive,largely due to the lack of observational data in the corresponding redshift range.
基金supported by the Zhejiang Provincial Natural Science Foundation(No.ZCLY24H1601)the National Natural Science Foundation of China(No.82403697)+1 种基金the Medical and Health Science and Technology Project of Zhejiang Province(No.2025KY411)the National Key R&D Program of China(No.2022YFC2505100).
摘要Real-world studies(RWSs)have emerged as a transformative force in oncology research,complementing traditional randomized controlled trials(RCTs)by providing comprehensive insights into cancer care within routine clinical settings.This review examines the evolving landscape of RWSs in oncology,focusing on their implementation,methodological considerations,and impact on precision medicine.We systematically analyze how RWSs leverage diverse data sources,including electronic health records(EHRs),insurance claims,and patient registries,to generate evidence that bridges the gap between controlled clinical trials and real-world clinical practice.The review underscores the key contributions of RWSs,including capturing therapeutic outcomes in traditionally underrepresented populations,expanding drug indications,and evaluating long-term safety and effectiveness in routine clinical settings.While acknowledging significant challenges,including data quality variability and privacy concerns,we discuss how emerging technologies like artificial intelligence are helping to address these limitations.The integration of RWSs with traditional clinical research is revolutionizing the paradigm of precision oncology and enabling more personalized treatment approaches based on real-world evidence.
摘要This study presents a method to correct the lithology of mud-logging profile with logging data based on neural network,which aims to solve the problems of time-consuming,high labor intensity and great infl uence of human factors in the process of traditional lithology correction of mud-logging profi le.Firstly,the lithology of mud-logging profi le is processed by digital technology and converted into digital curve which is consistent with the logging sampling interval,and the logging lithology curve is calculated by using the optimal logging method.Then,combining automatic depth-correction technology with manual correction methods,the lithology of mud-logging profi le is corrected for depth.On the basis of lithology depth-correction of mudlogging profile,the multi-layer perceptron(MLP)neural network is used to learn logging data and realize accurate identification of multiple lithologies,so as to construct a high-precision logging profile lithology curve and provide accurate basis for lithology correction of mud-logging profile.The effectiveness and accuracy of the proposed method are verifi ed by practical application cases.The corrected lithology of mudlogging profi le is highly consistent with the lithology of logging profi le,which provides a solid foundation for subsequent geological interpretation,reservoir evaluation and oil and gas resource assessment.This study not only improves the effi ciency of mud-logging data processing,but also ensures that the needs of exploration and exploitation work are met in a timely manner,which has important theoretical signifi cance and application value.
基金financial support from the National Key Research and Development Program of China(Grant No.2022YFA1403800)the National Natural Science Foundation of China(Grant Nos.12188101 and 11925408)+1 种基金the Chinese Academy of Sciences(Grant No.XDB33000000)support from the New Cornerstone Science Foundation through the XPLORER PRIZE。
摘要The discovery of topological materials has advanced rapidly due to high-throughput computation and machine learning,but research progress is hampered by inconsistent classification standards and fragmented data resources.Existing databases differ in computational methods,material coverage,and labeling criteria,making it difficult to compare findings across studies.To overcome these challenges,we present a unified topological materials dataset that systematically combines and reconciles two major databases:Materiae and the Topological Materials Database.This dataset provides consistent topological classifications for 35608 materials,accessible through the Materials Galaxy platform for interactive exploration and available for bulk download via MatElab.We describe the featurization methodology that converts crystal structures into 4710 machine-learning-ready descriptors and present a comprehensive analysis of topological material distributions.This work serves as a complete guide for accessing,utilizing,and interpreting this unified resource,designed to enable reproducible machine learning applications and accelerate the discovery of topological materials.
基金supported by the Three Year Action Project for Science and Technology Innovation of Beihai Bureau(Nos.2023B06-YJC and 2023B17-YJC)。
摘要In this study,buoy measurements of sea surface wind fields(10 m above the sea surface)were collected at 12 stations in the Bohai and Yellow Seas over a period of 11 years(2011-2021).An evaluation was conducted of the 10 m wind performance of the Climate Forecast System Reanalysis version 2(CFSv2),Modern-Era Retrospective Analysis for Research and Applications version 2(MERRA2),and the fifth-generation European Centre for Medium-Range Weather Forecasts reanalysis products(ERA5)in reproducing wind fields over the Bohai and Yellow Seas.The mean BIAS,root mean square error(RMSE),and correlation coefficient(R)were calculated and used as performance indicators.All three reanalysis products demonstrate commendable performance.The correlation coefficients for wind speed and wind direction exceed 0.83 and 0.90,respectively.Of the three products,the fifth-generation European Centre for Medium-Range Weather Forecasts reanalysis demonstrates the highest level of performance.The model demonstrated the lowest average BIAS and RMSE for wind speed,and the highest R for wind direction.The findings indicate that product performance is subject to variation in both season and wind speed.
基金supported by the National Natural Science Foundation of China(NSFC,grant No.U1731128).
摘要To overcome the limitations of ground-based telescopes in spatial resolution and imaging quality, recent researchhas concentrated on image super-resolution (SR) reconstruction. This approach aims to enhance ground-basedimages by recovering “space-based-like” images with higher spatial resolution and more detailed structuralinformation, without the need for expensive space-based observations. The training of SR models typically relieson large-scale, high-quality paired datasets of ground-based and space-based images. However, the acquisition ofsuch data is highly costly, which significantly limits the widespread application of these models. Based on this,this paper proposes using realistic synthetic galaxy images generated by IllustrisTNG cosmological simulation toreplace real space-based images and constructs a training dataset for image SR tasks, TNG-RealSR. To verify thevalidity of this dataset, we selected five mainstream lightweight SR models for evaluation. The results show thatTNG-RealSR exhibits good applicability in the task of galaxy image SR, providing reliable data support forrelated research. To the best of our knowledge, this work is the first to apply cosmological simulation-generatedsynthetic galaxy images to the field of galaxy image SR, demonstrating the feasibility of deep learning methodsbased on synthetic data for astronomical image enhancement tasks. This provides a low-cost and efficient solutionfor improving data quality in future large-scale surveys. The dataset is available online at http://gffzz188fe103f8f1460asonvwv5bwfxx96okp.ffgz.tsg.suse.edu.cn/jiaweimmiao/TNG-RealSR.
基金supported by Princess Nourah bint Abdulrahman University,Riyadh,Saudi Arabia through the Researchers Supporting Project PNURSP2026R760.
摘要Speech Emotion Recognition(SER)is a critical component of affective computing with broad applications in human–computer interaction,mental health monitoring,and intelligent multimedia systems.However,SER remains challenging due to the emotional ambiguity,lack of labeled data,class imbalance,and speaker variability.This study presents an effective SER framework that integrates contrastive representation learning,optimized spectrogram-based data augmentation,and selective synthetic data generation by using TimeGAN to enhance emotion classification performance.Contrastive learning enables the model to better discriminate acoustically similar emotions while Optuna automatically tunes augmentation strategies such as noise injection,time shifting,and time-frequency masking.Unlike existing approaches that apply synthetic generation uniformly across all classes,the proposed method targets only confusing or under-represented emotion classes to preserve the inter-class separability.A CNN-BiLSTM architecture is used to extract spectral and temporal information of the speech.The framework is evaluated with benchmark SER datasets—EMO-DB and RAVDESS—under speaker independent protocols.Experimental results demonstrate improved accuracy,robustness,and generalization under limited and imbalanced data conditions,supported by confusion matrices,UMAP,and t-SNE visualizations.
基金funded by the National Natural Science Foundation of China(U2243235 and 52309060)。
摘要Timely and accurate forecasting of crop yields is critical for food management and trade.However,only limited research has explored the impact of integrating crop phenotypic parameters(CPPs)with unmanned aerial vehicle(UAV)data across different phenological stages on maize yield prediction.The extent to which multi-temporal data enhances the accuracy and reliability of yield projections compared to mono-temporal data has yet to be systematically investigated.To attain the proper balance between accuracy and cost in crop yield estimation,this study proposed a structured framework for identifying the optimal phenological periods for summer maize yield prediction using UAV-based multispectral data.Three classical methods of custom mean decrease accuracy(C-MDA),optimal parameters-based geographical detector(OPGD),and grey relational analysis(GRA)were first used to sort and screen both the CPPs and vegetation indices(VIs)derived from UAV-based information over six growth stages.Ridge regression models based on multi-temporal data combinations and mono-temporal data were established separately,and their performance in yield prediction were compared to identify the optimal phenological stages and the corresponding key factors.Our results showed that C-MDA was much better at factor screening and ranking compared to OPGD and GRA.The green normalized difference vegetation index(GNDVI),normalized difference vegetation index(NDVI),and normalized difference red edge index(NDRE)emerged as the topperforming VIs,while the leaf area index(LAI)and above ground biomass(AGB)proved to be the most effective CPPs.When predicting yield using only mono-temporal data,the dough stage delivered the highest predictive accuracy(R2=0.871,RMSE=0.407 t ha-1),while the tasseling stage was the earliest that achieved yield estimates with acceptable precision(R2=0.810,RMSE=0.493 t ha-1).In contrast,the integration of UAV data from different crop growth stages markedly enhanced the accuracy of yield estimation.Combinations of data from the tasseling,silking,and dough stages were recommended as the best option(R2=0.942,RMSE=0.291 t ha-1).These findings indicate that the precise estimation of maize yields in smallholder fields may be attainable,and present both substantial theoretical insights and practical benefits for the advancement of precision agriculture.
基金supported by the National Key Research and Development Program of China(2024YFD2001100,2024YFE0214300)the National Natural Science Foundation of China(62162008)+3 种基金Guizhou Provincial Science and Technology Projects([2024]002,CXTD[2023]027)Guizhou Province Youth Science and Technology Talent Project([2024]317)Guiyang Guian Science and Technology Talent Training Project([2024]2-15)the Guizhou Provincial Graduate Research Fund Project(2024YJSKYJJ096)。
摘要Training software models for crop disease diagnosis requires large image datasets to achieve high accuracy.We describe a lesion information transfer diffusion model,LesionDiff,for generating image data that augments a real-world disease lesion image dataset.An information preprocessing module identifies lesion areas on leaves,an enhancement module captures diverse visual and semantic lesion features,and a generation module fills missing regions in masked disease images by synthesizing lesion phenotypes.This augmentation increased the average diagnostic accuracy of a test dataset by more than 3%.
基金supported by the Key Program of Natural Science Foundation of Tianjin(Grant No.21JCZDJC00770)the Tianjin Metrology Technology Project(Grant No.2024TJMT049).
摘要The traditional method of performance degradation prediction and maintenance of rolling bearings only considers a single sensor signal,which makes it difficult to automatically partition degradation stages and prone to over-detection.A new method of performance degradation evaluation and maintenance of rolling bearings based on data-level fusion,adaptive health state partitioning,and state maintenance is proposed.Firstly,considering the degradation and impact in the process of bearing deterioration,the multi-sensor signals are dynamically weighted to achieve data-level fusion.Secondly,a bearing health index was established based on fast spectral correlation,Wasserstein distance,and linear rectification techniques.On this basis,by combining the Bayesian information criterion and the elbow rule,the precise division of rolling bearing health state is realized through hidden Markov model regression.Then,random forest was used to classify and predict the data to verify the validity of the proposed data fusion method and health indicator.Finally,condition-based maintenance strategy based on the fourth moment,stress-strength interference model,and Gamma process is proposed to avoid excessive detection and reduce maintenance costs.Through accelerated degradation experiments and field validation tests on the rolling bearing test data set of Xi’an Jiaotong University and FEMTO(PRONOSTIA),the accuracy and superiority of the proposed method in the prediction and maintenance of bearing health state are verified.