The National Data Administration announced the establishment of the“data infrastructure technology community”and launched the“data standard semantic service platform”on April 28,which aims to accelerate the nation...The National Data Administration announced the establishment of the“data infrastructure technology community”and launched the“data standard semantic service platform”on April 28,which aims to accelerate the national data infrastructure and data standardization process during the 15th Five-Year Plan period(2026-2030).展开更多
Earthquakes are highly destructive spatio-temporal phenomena whose analysis is essential for disaster preparedness and risk mitigation.Modern seismological research produces vast volumes of heterogeneous data from sei...Earthquakes are highly destructive spatio-temporal phenomena whose analysis is essential for disaster preparedness and risk mitigation.Modern seismological research produces vast volumes of heterogeneous data from seismic networks,satellite observations,and geospatial repositories,creating the need for scalable infrastructures capable of integrating and analyzing such data to support intelligent decision-making.Data warehousing technologies provide a robust foundation for this purpose;however,existing earthquake-oriented data warehouses remain limited,often relying on simplified schemas,domain-specific analytics,or cataloguing efforts.This paper presents the design and implementation of a spatio-temporal data warehouse for seismic activity.The framework integrates spatial and temporal dimensions in a unified schema and introduces a novel array-based approach for managing many-to-many relationships between facts and dimensions without intermediate bridge tables.A comparative evaluation against a conventional bridge-table schema demonstrates that the array-based design improves fact-centric query performance,while the bridge-table schema remains advantageous for dimension-centric queries.To reconcile these trade-offs,a hybrid schema is proposed that retains both representations,ensuring balanced efficiency across heterogeneous workloads.The proposed framework demonstrates how spatio-temporal data warehousing can address schema complexity,improve query performance,and support multidimensional visualization.In doing so,it provides a foundation for integrating seismic analysis into broader big data-driven intelligent decision systems for disaster resilience,risk mitigation,and emergency management.展开更多
With continuous advancement in geological studies,three-dimensional(3D)geological modeling technology based on big data and artificial intelligence(AI)has become a prominent focus in the interdisciplinary field of ear...With continuous advancement in geological studies,three-dimensional(3D)geological modeling technology based on big data and artificial intelligence(AI)has become a prominent focus in the interdisciplinary field of earth and information sciences.Through the analysis and comparison of existing 3D modeling methods,this study introduces a high-precision 3D geological modeling approach.The new method leverages advanced computing technologies,including multisource heterogeneous data processing,integrated model databases,seamless splicing of local models,and cluster analysis.Furthermore,it enables unified 3D visualization of subsurface and surface conditions,providing new insights into disaster prevention,mitigation,and intelligent mineral exploration.To validate its practicality,this study conducts 3D geological modeling of the X area within the Sichuan Basin,China.The data sources include remote sensing images,geological maps,geophysical data,borehole data,and X-ray fluorescence(XRF)spectroscopy data.Preliminary exploration in the X area has successfully identified new mineralization belts,verifying the feasibility and effectiveness of the new 3D geological modeling method that integrates big data processing and intelligent techniques.展开更多
The Ok null test can not only assess whether the cosmic curvature is zero—thereby,if true,reducing degeneracies between cosmic curvature and other cosmological parameters—but also provide a model-independent chec...The Ok null test can not only assess whether the cosmic curvature is zero—thereby,if true,reducing degeneracies between cosmic curvature and other cosmological parameters—but also provide a model-independent check of compatibility between different data sets.However,traditional implementations often require absolute distance data from Type Ia supernovae(SNe Ia)or baryon acoustic oscillation(BAO)measurements,limiting their applicability because such absolute distance data are usually not accessible.The BAO Alcock-Paczynski(AP)parameter FAP is a measurement of a distance ratio,making the Dark Energy Spectroscopic Instrument(DESI)AP measurements particularly well suited for the Oknull test,as no absolute distance measurements are required.We propose a novel null test of cosmic curvature tailored to DESI BAO data that combines FAPwith ratios such as D′V/DVor D′M/DM.Crucially,this construction eliminates the need for absolute distance measurements.We further develop multi-task Gaussian processes to perform the null test.This approach can also be applied to a joint DESI BAO and SNe Ia dataset,and we find that DESI BAO and SNe Ia data are compatible.Although there is~2σ evidence of nonzero curvature at low redshift z■0.5,this result is not conclusive,largely due to the lack of observational data in the corresponding redshift range.展开更多
Airborne Mobile Networks(AMNs)require high-precision situational awareness in enclosed spaces for critical tasks like border surveillance and maritime monitoring.Compared to vision and wearable technologies,WiFi signa...Airborne Mobile Networks(AMNs)require high-precision situational awareness in enclosed spaces for critical tasks like border surveillance and maritime monitoring.Compared to vision and wearable technologies,WiFi signals in AMN environments are a promising sensing medium due to their non-intrusiveness,cost-effectiveness,and privacy benefits.However,the development of WiFi-based sensing in AMNs is hindered by the scarcity of high-quality large-scale data.While Data Augmentation(DA)can address this scarcity,traditional time-series DA may distort WiFi Channel State Information(CSI)'s time–frequency properties,and deep generative models suffer from mode collapse,spectral distortion,and high computational costs.To overcome these limitations,we propose OT-ADG,an adaptive data generation algorithm based on Optimal Transport(OT)theory to synthesize high-fidelity WiFi sensing samples.It dynamically models subcarrierspecific energy distributions and computes optimal transport plans using Sinkhorn's algorithm with entropy regularization.Furthermore,we present ViFi,the first large-scale and scenario-rich WiFi sensing dataset to fill the gap in AMN applications,encompassing 26 real-world scenarios with 20640 CSI samples and synchronized videos.Experimental results show OT-ADG's superior performance,with maximum improvements of up to 14.88%over baseline recognition methods,while outperforming existing DA approaches.Moreover,the robust performance of multiple WiFi-based HAR models validates the effectiveness of the ViFi dataset.展开更多
Parkinson’s disease(PD)is a debilitating neurological disorder affecting over 10 million people worldwide.PD classification models using voice signals as input are common in the literature.It is believed that using d...Parkinson’s disease(PD)is a debilitating neurological disorder affecting over 10 million people worldwide.PD classification models using voice signals as input are common in the literature.It is believed that using deep learning algorithms further enhances performance;nevertheless,it is challenging due to the nature of small-scale and imbalanced PD datasets.This paper proposed a convolutional neural network-based deep support vector machine(CNN-DSVM)to automate the feature extraction process using CNN and extend the conventional SVM to a DSVM for better classification performance in small-scale PD datasets.A customized kernel function reduces the impact of biased classification towards the majority class(healthy candidates in our consideration).An improved generative adversarial network(IGAN)was designed to generate additional training data to enhance the model’s performance.For performance evaluation,the proposed algorithm achieves a sensitivity of 97.6%and a specificity of 97.3%.The performance comparison is evaluated from five perspectives,including comparisons with different data generation algorithms,feature extraction techniques,kernel functions,and existing works.Results reveal the effectiveness of the IGAN algorithm,which improves the sensitivity and specificity by 4.05%–4.72%and 4.96%–5.86%,respectively;and the effectiveness of the CNN-DSVM algorithm,which improves the sensitivity by 1.24%–57.4%and specificity by 1.04%–163%and reduces biased detection towards the majority class.The ablation experiments confirm the effectiveness of individual components.Two future research directions have also been suggested.展开更多
Biomedical data is surging due to technological innovations and integration of multidisciplinary data,posing challenges to data management.This article summarizes the policies,data collection efforts,platform construc...Biomedical data is surging due to technological innovations and integration of multidisciplinary data,posing challenges to data management.This article summarizes the policies,data collection efforts,platform construction,and applications of biomedical data in China,aiming to identify key issues and needs,enhance the capacity-building of platform construction,unleash the value of data,and leverage the advantages of China's vast amount of data.展开更多
tRNA-derived small RNAs(tsRNAs),as a class of regulatory small noncoding RNA,have been implicated in a wide variety of human diseases.Large amounts of tsRNA–disease associations have been identified in recent years f...tRNA-derived small RNAs(tsRNAs),as a class of regulatory small noncoding RNA,have been implicated in a wide variety of human diseases.Large amounts of tsRNA–disease associations have been identified in recent years from accumulating studies.However,repositories for cataloging the detailed information on tsRNA–disease associations are scarce.In this study,we provide a tsRNADisease database by integrating experimentally and computationally supported tsRNA–disease associations from manual curation of literatures and other related resources.tsRNADisease contains 5571 manually curated associations between 4759 tsRNAs and 166 diseases with experimental evidence from 346 studies.In addition,it also contains 5013 predicted associations between 1297 tsRNAs and 111 diseases.tsRNADisease provides a user-friendly interface to browse,retrieve,and download data conveniently.This database can improve our understanding of tsRNA deregulation in diseases and serve as a valuable resource for investigating the mechanism of disease-related tsRNAs.tsRNADisease is freely available at http://gffzz9c504e06f78b4edahf0w9ko5vw66w6bvc.ffgz.tsg.suse.edu.cn.展开更多
Intelligent Connected Vehicles(ICVs)generate massive heterogeneous multi-modal data during operation,and due to the limited computing resources on board,graded data encryption protection is of great significance for b...Intelligent Connected Vehicles(ICVs)generate massive heterogeneous multi-modal data during operation,and due to the limited computing resources on board,graded data encryption protection is of great significance for balancing data security and efficient utilization.However,the current data grading processes struggle to address the evolving inference attacks and dynamic operational environments,and traditional grading approaches relying on static expert judgment or information-theoretic metrics.To bridge this gap,this paper proposes a novel inference strength-driven data grading framework,where inference strength quantifies the susceptibility of one dataset to infer another through adversarial reasoning.The framework employs a systematic methodology combining graph theory,optimization,and Large Language Models to construct an inference library and calculate inference strength.The framework also provides a PageRank-based algorithm to generate interpretable data grading lists for both static policy and vehicle-end application,prioritizing core data protection while respecting computational constraints.Validated on the Audi A2D2 dataset and real vehicle controller,our approach demonstrates improved protection utility compared to default grading baselines.The results highlight its potential to enhance data security in ICVs through prioritized protection of core data under computational constraints.展开更多
The full waveform inversion(FWI)utilizes full wave field data to invert subsurface parameters and is considered one of the most promising data-driven tools for obtaining high precision velocity models.However,the succ...The full waveform inversion(FWI)utilizes full wave field data to invert subsurface parameters and is considered one of the most promising data-driven tools for obtaining high precision velocity models.However,the successful application of FWI in geophysical explo ration remains limited,primarily due to the cycle-skipping issue caused by the absence of low-frequency data,which is one of the main reasons for FWI failures.Incorporating prior regularization constraints FWI can effectively compensate for the lacking low-frequency components and constrain the iterative updates of FWI toward the desired direction,offering a natural advantage in addressing this challenge.However,the weights of the prior information terms are still determined empirically,which introduces significant subjectivity and randomness to the inversion results.To solve this issue,we propose an adaptive method to determine the weight factor based on posterior probability distribution within the Bayesian theoretical framework.This factor adaptively adjusts during each iteration to balance the contributions of the data error term and the prior information term in FWI,which can effectively mitigate the cycle-skipping problem and alleviating the nonlinearity of the inversion process.Numerical examples from the Overthrust model and the Marmousi model show that our method not only enhance the accuracy of FWI,but also demonstrate strong noise resistance.展开更多
This study investigates the pivotal role of data clustering in both data science and management,focusing on core methodologies,tools,and diverse applications.It examines traditional clustering techniques such as parti...This study investigates the pivotal role of data clustering in both data science and management,focusing on core methodologies,tools,and diverse applications.It examines traditional clustering techniques such as partitional and hierarchical methods,alongside more advanced approaches,including data stream,density-based,graphbased,and model-based clustering,which are essential for processing complex and structured datasets.The study highlights fundamental principles,presents commonly adapted tools and frameworks,outlines the clustering workflow within data science,and discusses major implementation challenges.Beyond technical applications,this study emphasizes how clustering supports managerial tasks and decision-making through a comprehensive survey of recent literature.By bridging analytical techniques with real-world business needs,clustering remains an essential tool in both data science and management.The study concludes by outlining future research directions,underscoring the role of clustering in driving innovation and enabling informed strategic and operational decisions.展开更多
Effective fault diagnosis is crucial for the reliable running of Electromechanical coupling Systems(EMS),yet hampered by insufficient entity fault data.Digital Twin(DT)technology offers the potential for virtual fault...Effective fault diagnosis is crucial for the reliable running of Electromechanical coupling Systems(EMS),yet hampered by insufficient entity fault data.Digital Twin(DT)technology offers the potential for virtual fault data augmentation and fault diagnosis improvement.However,there is still a lack of an effective and systematic methodology,to decouple complicated EMS entities and construct their full-system DT.To address this,a hierarchical collaborative DT construction framework is proposed for fault data augmentation of EMS.Specially,we decouple EMS entity into the triplet representations of element,data,and relationship,which establish the profound understanding of coupling characteristics from multiple modalities.Furthermore,we develop a hierarchical DT modeling method to mirror these complicated couplings as four-level sub-DTs of space,behavior,process,and status.Each level of sub-DT utilizes the data-mechanism combined technique to balance modeling adaptability and precision.Finally,these heterogeneous sub-DTs are integrated as full-system DT driven by collaborative orchestration algorithm,which achieves the global consistency mirror with the real fault manifestation under diverse fault modes.Experiments on a multi-coupled electromechanical fault test bench validate our framework.Results exhibit the average improvements of 17.29%and 9.97%in accuracy of data augmentation fault classification,confirming its superiority and effectiveness.展开更多
Real-world studies(RWSs)have emerged as a transformative force in oncology research,complementing traditional randomized controlled trials(RCTs)by providing comprehensive insights into cancer care within routine clini...Real-world studies(RWSs)have emerged as a transformative force in oncology research,complementing traditional randomized controlled trials(RCTs)by providing comprehensive insights into cancer care within routine clinical settings.This review examines the evolving landscape of RWSs in oncology,focusing on their implementation,methodological considerations,and impact on precision medicine.We systematically analyze how RWSs leverage diverse data sources,including electronic health records(EHRs),insurance claims,and patient registries,to generate evidence that bridges the gap between controlled clinical trials and real-world clinical practice.The review underscores the key contributions of RWSs,including capturing therapeutic outcomes in traditionally underrepresented populations,expanding drug indications,and evaluating long-term safety and effectiveness in routine clinical settings.While acknowledging significant challenges,including data quality variability and privacy concerns,we discuss how emerging technologies like artificial intelligence are helping to address these limitations.The integration of RWSs with traditional clinical research is revolutionizing the paradigm of precision oncology and enabling more personalized treatment approaches based on real-world evidence.展开更多
To overcome the limitations of ground-based telescopes in spatial resolution and imaging quality, recent researchhas concentrated on image super-resolution (SR) reconstruction. This approach aims to enhance ground-bas...To overcome the limitations of ground-based telescopes in spatial resolution and imaging quality, recent researchhas concentrated on image super-resolution (SR) reconstruction. This approach aims to enhance ground-basedimages by recovering “space-based-like” images with higher spatial resolution and more detailed structuralinformation, without the need for expensive space-based observations. The training of SR models typically relieson large-scale, high-quality paired datasets of ground-based and space-based images. However, the acquisition ofsuch data is highly costly, which significantly limits the widespread application of these models. Based on this,this paper proposes using realistic synthetic galaxy images generated by IllustrisTNG cosmological simulation toreplace real space-based images and constructs a training dataset for image SR tasks, TNG-RealSR. To verify thevalidity of this dataset, we selected five mainstream lightweight SR models for evaluation. The results show thatTNG-RealSR exhibits good applicability in the task of galaxy image SR, providing reliable data support forrelated research. To the best of our knowledge, this work is the first to apply cosmological simulation-generatedsynthetic galaxy images to the field of galaxy image SR, demonstrating the feasibility of deep learning methodsbased on synthetic data for astronomical image enhancement tasks. This provides a low-cost and efficient solutionfor improving data quality in future large-scale surveys. The dataset is available online at http://gffzz188fe103f8f1460asf0w9ko5vw66w6bvc.ffgz.tsg.suse.edu.cn/jiaweimmiao/TNG-RealSR.展开更多
This study presents a method to correct the lithology of mud-logging profile with logging data based on neural network,which aims to solve the problems of time-consuming,high labor intensity and great infl uence of hu...This study presents a method to correct the lithology of mud-logging profile with logging data based on neural network,which aims to solve the problems of time-consuming,high labor intensity and great infl uence of human factors in the process of traditional lithology correction of mud-logging profi le.Firstly,the lithology of mud-logging profi le is processed by digital technology and converted into digital curve which is consistent with the logging sampling interval,and the logging lithology curve is calculated by using the optimal logging method.Then,combining automatic depth-correction technology with manual correction methods,the lithology of mud-logging profi le is corrected for depth.On the basis of lithology depth-correction of mudlogging profile,the multi-layer perceptron(MLP)neural network is used to learn logging data and realize accurate identification of multiple lithologies,so as to construct a high-precision logging profile lithology curve and provide accurate basis for lithology correction of mud-logging profile.The effectiveness and accuracy of the proposed method are verifi ed by practical application cases.The corrected lithology of mudlogging profi le is highly consistent with the lithology of logging profi le,which provides a solid foundation for subsequent geological interpretation,reservoir evaluation and oil and gas resource assessment.This study not only improves the effi ciency of mud-logging data processing,but also ensures that the needs of exploration and exploitation work are met in a timely manner,which has important theoretical signifi cance and application value.展开更多
In this study,buoy measurements of sea surface wind fields(10 m above the sea surface)were collected at 12 stations in the Bohai and Yellow Seas over a period of 11 years(2011-2021).An evaluation was conducted of the ...In this study,buoy measurements of sea surface wind fields(10 m above the sea surface)were collected at 12 stations in the Bohai and Yellow Seas over a period of 11 years(2011-2021).An evaluation was conducted of the 10 m wind performance of the Climate Forecast System Reanalysis version 2(CFSv2),Modern-Era Retrospective Analysis for Research and Applications version 2(MERRA2),and the fifth-generation European Centre for Medium-Range Weather Forecasts reanalysis products(ERA5)in reproducing wind fields over the Bohai and Yellow Seas.The mean BIAS,root mean square error(RMSE),and correlation coefficient(R)were calculated and used as performance indicators.All three reanalysis products demonstrate commendable performance.The correlation coefficients for wind speed and wind direction exceed 0.83 and 0.90,respectively.Of the three products,the fifth-generation European Centre for Medium-Range Weather Forecasts reanalysis demonstrates the highest level of performance.The model demonstrated the lowest average BIAS and RMSE for wind speed,and the highest R for wind direction.The findings indicate that product performance is subject to variation in both season and wind speed.展开更多
Speech Emotion Recognition(SER)is a critical component of affective computing with broad applications in human–computer interaction,mental health monitoring,and intelligent multimedia systems.However,SER remains chal...Speech Emotion Recognition(SER)is a critical component of affective computing with broad applications in human–computer interaction,mental health monitoring,and intelligent multimedia systems.However,SER remains challenging due to the emotional ambiguity,lack of labeled data,class imbalance,and speaker variability.This study presents an effective SER framework that integrates contrastive representation learning,optimized spectrogram-based data augmentation,and selective synthetic data generation by using TimeGAN to enhance emotion classification performance.Contrastive learning enables the model to better discriminate acoustically similar emotions while Optuna automatically tunes augmentation strategies such as noise injection,time shifting,and time-frequency masking.Unlike existing approaches that apply synthetic generation uniformly across all classes,the proposed method targets only confusing or under-represented emotion classes to preserve the inter-class separability.A CNN-BiLSTM architecture is used to extract spectral and temporal information of the speech.The framework is evaluated with benchmark SER datasets—EMO-DB and RAVDESS—under speaker independent protocols.Experimental results demonstrate improved accuracy,robustness,and generalization under limited and imbalanced data conditions,supported by confusion matrices,UMAP,and t-SNE visualizations.展开更多
Acupuncture research increasingly involves heterogeneous and multimodal data that are difficult to analyze using conventional methods.This review summarizes data-driven approaches in acupuncture research within a fram...Acupuncture research increasingly involves heterogeneous and multimodal data that are difficult to analyze using conventional methods.This review summarizes data-driven approaches in acupuncture research within a framework encompassing intervention,response,and contextual data.We discuss causal inference,artificial intelligence,text mining,and integrative analysis,along with their applications in efficacy evaluation,outcome prediction,mechanistic investigation,and clinical decision support.These approaches shift the focus of acupuncture research from population-level average effects toward individualized clinical decision-making by enabling the analysis of treatment heterogeneity and underlying mechanisms.However,current research remains limited by inadequate data standardization,insufficient external validation,and limited model interpretability.Despite these challenges,data-driven approaches offer substantial promise for advancing more rigorous and personalized acupuncture research.展开更多
Urban traffic generates massive and diverse data,yet most systems remain fragmented.Current approaches to congestion management suffer from weak data consistency and poor scalability.This study addresses this gap by p...Urban traffic generates massive and diverse data,yet most systems remain fragmented.Current approaches to congestion management suffer from weak data consistency and poor scalability.This study addresses this gap by proposing the Urban Traffic Congestion Unified Metadata Model(UTC-UMM).The goal is to provide a standardized and extensible framework for describing,extracting,and storing multisource traffic data in smart cities.The model defines a two-tier specification that organizes nine core traffic resource classes.It employs an eXtensible Markup Language(XML)Schema that connects general elements with resource-specific elements.This design ensures both syntactic and semantic interoperability across siloed datasets.Extension principles allow new elements or constraints to be introducedwithout breaking backward compatibility.Adistributed pipeline is implemented usingHadoop Distributed File System(HDFS)and HBase.It integrates computer vision for video and natural language processing for text to automate metadata extraction.Optimized row-key designs enable low-latency queries.Performance is tested with the Yahoo!Cloud Serving Benchmark(YCSB),which shows linear scalability and high throughput.The results demonstrate that UTC-UMM can unify heterogeneous traffic data while supporting real-time analytics.The discussion highlights its potential to improve data reuse,portability,and scalability in urban congestion studies.Future research will explore integration with association rulemining and advanced knowledge representation to capture richer spatiotemporal traffic patterns.展开更多
The publisher regrets the CRediT authorship contribution statement was inserted incorrectly and the correct statement should be updated as below:Zengji Liu:Writing-review&editing,Writing-original draft,Visualizati...The publisher regrets the CRediT authorship contribution statement was inserted incorrectly and the correct statement should be updated as below:Zengji Liu:Writing-review&editing,Writing-original draft,Visualization,Validation,Supervision,Software,Resources,Project administration,Methodology,Investigation,Funding acquisition,Formal analysis,Data curation,Conceptualization.Mengge Liu:Writing-review&editing,Writing-original draft,Investigation.Qi Wang:Writing-review&editing,Writing-original draft.Yi Tang:Writing-review&editing,Writing-original draft.展开更多
摘要The National Data Administration announced the establishment of the“data infrastructure technology community”and launched the“data standard semantic service platform”on April 28,which aims to accelerate the national data infrastructure and data standardization process during the 15th Five-Year Plan period(2026-2030).
摘要Earthquakes are highly destructive spatio-temporal phenomena whose analysis is essential for disaster preparedness and risk mitigation.Modern seismological research produces vast volumes of heterogeneous data from seismic networks,satellite observations,and geospatial repositories,creating the need for scalable infrastructures capable of integrating and analyzing such data to support intelligent decision-making.Data warehousing technologies provide a robust foundation for this purpose;however,existing earthquake-oriented data warehouses remain limited,often relying on simplified schemas,domain-specific analytics,or cataloguing efforts.This paper presents the design and implementation of a spatio-temporal data warehouse for seismic activity.The framework integrates spatial and temporal dimensions in a unified schema and introduces a novel array-based approach for managing many-to-many relationships between facts and dimensions without intermediate bridge tables.A comparative evaluation against a conventional bridge-table schema demonstrates that the array-based design improves fact-centric query performance,while the bridge-table schema remains advantageous for dimension-centric queries.To reconcile these trade-offs,a hybrid schema is proposed that retains both representations,ensuring balanced efficiency across heterogeneous workloads.The proposed framework demonstrates how spatio-temporal data warehousing can address schema complexity,improve query performance,and support multidimensional visualization.In doing so,it provides a foundation for integrating seismic analysis into broader big data-driven intelligent decision systems for disaster resilience,risk mitigation,and emergency management.
基金supported by a project of the China Geological Survey(DD20240126).
摘要With continuous advancement in geological studies,three-dimensional(3D)geological modeling technology based on big data and artificial intelligence(AI)has become a prominent focus in the interdisciplinary field of earth and information sciences.Through the analysis and comparison of existing 3D modeling methods,this study introduces a high-precision 3D geological modeling approach.The new method leverages advanced computing technologies,including multisource heterogeneous data processing,integrated model databases,seamless splicing of local models,and cluster analysis.Furthermore,it enables unified 3D visualization of subsurface and surface conditions,providing new insights into disaster prevention,mitigation,and intelligent mineral exploration.To validate its practicality,this study conducts 3D geological modeling of the X area within the Sichuan Basin,China.The data sources include remote sensing images,geological maps,geophysical data,borehole data,and X-ray fluorescence(XRF)spectroscopy data.Preliminary exploration in the X area has successfully identified new mineralization belts,verifying the feasibility and effectiveness of the new 3D geological modeling method that integrates big data processing and intelligent techniques.
基金supported in part by the National Natural Science Foundation of China(Grant Nos.12588101,12535002,12175184,12433001,and 12205015)。
摘要The Ok null test can not only assess whether the cosmic curvature is zero—thereby,if true,reducing degeneracies between cosmic curvature and other cosmological parameters—but also provide a model-independent check of compatibility between different data sets.However,traditional implementations often require absolute distance data from Type Ia supernovae(SNe Ia)or baryon acoustic oscillation(BAO)measurements,limiting their applicability because such absolute distance data are usually not accessible.The BAO Alcock-Paczynski(AP)parameter FAP is a measurement of a distance ratio,making the Dark Energy Spectroscopic Instrument(DESI)AP measurements particularly well suited for the Oknull test,as no absolute distance measurements are required.We propose a novel null test of cosmic curvature tailored to DESI BAO data that combines FAPwith ratios such as D′V/DVor D′M/DM.Crucially,this construction eliminates the need for absolute distance measurements.We further develop multi-task Gaussian processes to perform the null test.This approach can also be applied to a joint DESI BAO and SNe Ia dataset,and we find that DESI BAO and SNe Ia data are compatible.Although there is~2σ evidence of nonzero curvature at low redshift z■0.5,this result is not conclusive,largely due to the lack of observational data in the corresponding redshift range.
基金supported by the National Science Foundation of China(Nos.62027826,62502066)the Fundamental Research Funds for the Central Universities of China(No.DUT25RC(3)044)the Natural Science Foundation Project of Liaoning Province of China(No.2025-BS-0003)。
摘要Airborne Mobile Networks(AMNs)require high-precision situational awareness in enclosed spaces for critical tasks like border surveillance and maritime monitoring.Compared to vision and wearable technologies,WiFi signals in AMN environments are a promising sensing medium due to their non-intrusiveness,cost-effectiveness,and privacy benefits.However,the development of WiFi-based sensing in AMNs is hindered by the scarcity of high-quality large-scale data.While Data Augmentation(DA)can address this scarcity,traditional time-series DA may distort WiFi Channel State Information(CSI)'s time–frequency properties,and deep generative models suffer from mode collapse,spectral distortion,and high computational costs.To overcome these limitations,we propose OT-ADG,an adaptive data generation algorithm based on Optimal Transport(OT)theory to synthesize high-fidelity WiFi sensing samples.It dynamically models subcarrierspecific energy distributions and computes optimal transport plans using Sinkhorn's algorithm with entropy regularization.Furthermore,we present ViFi,the first large-scale and scenario-rich WiFi sensing dataset to fill the gap in AMN applications,encompassing 26 real-world scenarios with 20640 CSI samples and synchronized videos.Experimental results show OT-ADG's superior performance,with maximum improvements of up to 14.88%over baseline recognition methods,while outperforming existing DA approaches.Moreover,the robust performance of multiple WiFi-based HAR models validates the effectiveness of the ViFi dataset.
基金The work described in this paper was fully supported by a grant from Hong Kong Metropolitan University(RIF/2021/05).
摘要Parkinson’s disease(PD)is a debilitating neurological disorder affecting over 10 million people worldwide.PD classification models using voice signals as input are common in the literature.It is believed that using deep learning algorithms further enhances performance;nevertheless,it is challenging due to the nature of small-scale and imbalanced PD datasets.This paper proposed a convolutional neural network-based deep support vector machine(CNN-DSVM)to automate the feature extraction process using CNN and extend the conventional SVM to a DSVM for better classification performance in small-scale PD datasets.A customized kernel function reduces the impact of biased classification towards the majority class(healthy candidates in our consideration).An improved generative adversarial network(IGAN)was designed to generate additional training data to enhance the model’s performance.For performance evaluation,the proposed algorithm achieves a sensitivity of 97.6%and a specificity of 97.3%.The performance comparison is evaluated from five perspectives,including comparisons with different data generation algorithms,feature extraction techniques,kernel functions,and existing works.Results reveal the effectiveness of the IGAN algorithm,which improves the sensitivity and specificity by 4.05%–4.72%and 4.96%–5.86%,respectively;and the effectiveness of the CNN-DSVM algorithm,which improves the sensitivity by 1.24%–57.4%and specificity by 1.04%–163%and reduces biased detection towards the majority class.The ablation experiments confirm the effectiveness of individual components.Two future research directions have also been suggested.
摘要Biomedical data is surging due to technological innovations and integration of multidisciplinary data,posing challenges to data management.This article summarizes the policies,data collection efforts,platform construction,and applications of biomedical data in China,aiming to identify key issues and needs,enhance the capacity-building of platform construction,unleash the value of data,and leverage the advantages of China's vast amount of data.
基金supported by the National Natural Science Foundation of China(91959106)the Foundation of the Shanghai Municipal Education Commission(24RGZNC02)+4 种基金Shanghai Key Laboratory of Intelligent Information Processing,Fudan University(IIPL-2025-RD3-02)Key University Science Research Project of Anhui Province(2023AH030108)Climbing Peak Training Program for Innovative Technology team of Yijishan Hospital,Wannan Medical College(PF201904)Peak Training Program for Scientific Research of Yijishan Hospital,Wannan Medical College(GF2019G15)the talent project of the First Affiliated Hospital of Wannan Medical College(Yijishan Hospital of Wannan Medical College)(YR202422).
摘要tRNA-derived small RNAs(tsRNAs),as a class of regulatory small noncoding RNA,have been implicated in a wide variety of human diseases.Large amounts of tsRNA–disease associations have been identified in recent years from accumulating studies.However,repositories for cataloging the detailed information on tsRNA–disease associations are scarce.In this study,we provide a tsRNADisease database by integrating experimentally and computationally supported tsRNA–disease associations from manual curation of literatures and other related resources.tsRNADisease contains 5571 manually curated associations between 4759 tsRNAs and 166 diseases with experimental evidence from 346 studies.In addition,it also contains 5013 predicted associations between 1297 tsRNAs and 111 diseases.tsRNADisease provides a user-friendly interface to browse,retrieve,and download data conveniently.This database can improve our understanding of tsRNA deregulation in diseases and serve as a valuable resource for investigating the mechanism of disease-related tsRNAs.tsRNADisease is freely available at http://gffzz9c504e06f78b4edahf0w9ko5vw66w6bvc.ffgz.tsg.suse.edu.cn.
基金supported by National Key R&D Program of China(No.2022YFB3104900)the National Natural Science Foundation of China(No.U22A202101)State Key Laboratory of Intelligent Vehicle Safety Technology program(No.IVSTSKL-202439)。
摘要Intelligent Connected Vehicles(ICVs)generate massive heterogeneous multi-modal data during operation,and due to the limited computing resources on board,graded data encryption protection is of great significance for balancing data security and efficient utilization.However,the current data grading processes struggle to address the evolving inference attacks and dynamic operational environments,and traditional grading approaches relying on static expert judgment or information-theoretic metrics.To bridge this gap,this paper proposes a novel inference strength-driven data grading framework,where inference strength quantifies the susceptibility of one dataset to infer another through adversarial reasoning.The framework employs a systematic methodology combining graph theory,optimization,and Large Language Models to construct an inference library and calculate inference strength.The framework also provides a PageRank-based algorithm to generate interpretable data grading lists for both static policy and vehicle-end application,prioritizing core data protection while respecting computational constraints.Validated on the Audi A2D2 dataset and real vehicle controller,our approach demonstrates improved protection utility compared to default grading baselines.The results highlight its potential to enhance data security in ICVs through prioritized protection of core data under computational constraints.
基金partly supported by National Natural Science Foundation of China(42274154)the Fund of State Key Laboratory of Deep Oil and Gas,China University of Petroleum(East China),China(SKLDOG2024-ZYTS-03)。
摘要The full waveform inversion(FWI)utilizes full wave field data to invert subsurface parameters and is considered one of the most promising data-driven tools for obtaining high precision velocity models.However,the successful application of FWI in geophysical explo ration remains limited,primarily due to the cycle-skipping issue caused by the absence of low-frequency data,which is one of the main reasons for FWI failures.Incorporating prior regularization constraints FWI can effectively compensate for the lacking low-frequency components and constrain the iterative updates of FWI toward the desired direction,offering a natural advantage in addressing this challenge.However,the weights of the prior information terms are still determined empirically,which introduces significant subjectivity and randomness to the inversion results.To solve this issue,we propose an adaptive method to determine the weight factor based on posterior probability distribution within the Bayesian theoretical framework.This factor adaptively adjusts during each iteration to balance the contributions of the data error term and the prior information term in FWI,which can effectively mitigate the cycle-skipping problem and alleviating the nonlinearity of the inversion process.Numerical examples from the Overthrust model and the Marmousi model show that our method not only enhance the accuracy of FWI,but also demonstrate strong noise resistance.
摘要This study investigates the pivotal role of data clustering in both data science and management,focusing on core methodologies,tools,and diverse applications.It examines traditional clustering techniques such as partitional and hierarchical methods,alongside more advanced approaches,including data stream,density-based,graphbased,and model-based clustering,which are essential for processing complex and structured datasets.The study highlights fundamental principles,presents commonly adapted tools and frameworks,outlines the clustering workflow within data science,and discusses major implementation challenges.Beyond technical applications,this study emphasizes how clustering supports managerial tasks and decision-making through a comprehensive survey of recent literature.By bridging analytical techniques with real-world business needs,clustering remains an essential tool in both data science and management.The study concludes by outlining future research directions,underscoring the role of clustering in driving innovation and enabling informed strategic and operational decisions.
基金supported by the National Key R&D Programof China under Grant STI 2030—Major Projects(No.2021ZD0201300)the National Natural Science Foundation of China(Nos.52472442,72471013)+1 种基金the Zhejiang Provincial Natural Science Foundation(Grant No.LMS26E050036)the Research Start-up Funds of Hangzhou International Innovation Institute of Beihang University,China(Nos.2024KQ069,2024KQ036,2024KQ035 and 2025BKZ055)。
摘要Effective fault diagnosis is crucial for the reliable running of Electromechanical coupling Systems(EMS),yet hampered by insufficient entity fault data.Digital Twin(DT)technology offers the potential for virtual fault data augmentation and fault diagnosis improvement.However,there is still a lack of an effective and systematic methodology,to decouple complicated EMS entities and construct their full-system DT.To address this,a hierarchical collaborative DT construction framework is proposed for fault data augmentation of EMS.Specially,we decouple EMS entity into the triplet representations of element,data,and relationship,which establish the profound understanding of coupling characteristics from multiple modalities.Furthermore,we develop a hierarchical DT modeling method to mirror these complicated couplings as four-level sub-DTs of space,behavior,process,and status.Each level of sub-DT utilizes the data-mechanism combined technique to balance modeling adaptability and precision.Finally,these heterogeneous sub-DTs are integrated as full-system DT driven by collaborative orchestration algorithm,which achieves the global consistency mirror with the real fault manifestation under diverse fault modes.Experiments on a multi-coupled electromechanical fault test bench validate our framework.Results exhibit the average improvements of 17.29%and 9.97%in accuracy of data augmentation fault classification,confirming its superiority and effectiveness.
基金supported by the Zhejiang Provincial Natural Science Foundation(No.ZCLY24H1601)the National Natural Science Foundation of China(No.82403697)+1 种基金the Medical and Health Science and Technology Project of Zhejiang Province(No.2025KY411)the National Key R&D Program of China(No.2022YFC2505100).
摘要Real-world studies(RWSs)have emerged as a transformative force in oncology research,complementing traditional randomized controlled trials(RCTs)by providing comprehensive insights into cancer care within routine clinical settings.This review examines the evolving landscape of RWSs in oncology,focusing on their implementation,methodological considerations,and impact on precision medicine.We systematically analyze how RWSs leverage diverse data sources,including electronic health records(EHRs),insurance claims,and patient registries,to generate evidence that bridges the gap between controlled clinical trials and real-world clinical practice.The review underscores the key contributions of RWSs,including capturing therapeutic outcomes in traditionally underrepresented populations,expanding drug indications,and evaluating long-term safety and effectiveness in routine clinical settings.While acknowledging significant challenges,including data quality variability and privacy concerns,we discuss how emerging technologies like artificial intelligence are helping to address these limitations.The integration of RWSs with traditional clinical research is revolutionizing the paradigm of precision oncology and enabling more personalized treatment approaches based on real-world evidence.
基金supported by the National Natural Science Foundation of China(NSFC,grant No.U1731128).
摘要To overcome the limitations of ground-based telescopes in spatial resolution and imaging quality, recent researchhas concentrated on image super-resolution (SR) reconstruction. This approach aims to enhance ground-basedimages by recovering “space-based-like” images with higher spatial resolution and more detailed structuralinformation, without the need for expensive space-based observations. The training of SR models typically relieson large-scale, high-quality paired datasets of ground-based and space-based images. However, the acquisition ofsuch data is highly costly, which significantly limits the widespread application of these models. Based on this,this paper proposes using realistic synthetic galaxy images generated by IllustrisTNG cosmological simulation toreplace real space-based images and constructs a training dataset for image SR tasks, TNG-RealSR. To verify thevalidity of this dataset, we selected five mainstream lightweight SR models for evaluation. The results show thatTNG-RealSR exhibits good applicability in the task of galaxy image SR, providing reliable data support forrelated research. To the best of our knowledge, this work is the first to apply cosmological simulation-generatedsynthetic galaxy images to the field of galaxy image SR, demonstrating the feasibility of deep learning methodsbased on synthetic data for astronomical image enhancement tasks. This provides a low-cost and efficient solutionfor improving data quality in future large-scale surveys. The dataset is available online at http://gffzz188fe103f8f1460asf0w9ko5vw66w6bvc.ffgz.tsg.suse.edu.cn/jiaweimmiao/TNG-RealSR.
摘要This study presents a method to correct the lithology of mud-logging profile with logging data based on neural network,which aims to solve the problems of time-consuming,high labor intensity and great infl uence of human factors in the process of traditional lithology correction of mud-logging profi le.Firstly,the lithology of mud-logging profi le is processed by digital technology and converted into digital curve which is consistent with the logging sampling interval,and the logging lithology curve is calculated by using the optimal logging method.Then,combining automatic depth-correction technology with manual correction methods,the lithology of mud-logging profi le is corrected for depth.On the basis of lithology depth-correction of mudlogging profile,the multi-layer perceptron(MLP)neural network is used to learn logging data and realize accurate identification of multiple lithologies,so as to construct a high-precision logging profile lithology curve and provide accurate basis for lithology correction of mud-logging profile.The effectiveness and accuracy of the proposed method are verifi ed by practical application cases.The corrected lithology of mudlogging profi le is highly consistent with the lithology of logging profi le,which provides a solid foundation for subsequent geological interpretation,reservoir evaluation and oil and gas resource assessment.This study not only improves the effi ciency of mud-logging data processing,but also ensures that the needs of exploration and exploitation work are met in a timely manner,which has important theoretical signifi cance and application value.
基金supported by the Three Year Action Project for Science and Technology Innovation of Beihai Bureau(Nos.2023B06-YJC and 2023B17-YJC)。
摘要In this study,buoy measurements of sea surface wind fields(10 m above the sea surface)were collected at 12 stations in the Bohai and Yellow Seas over a period of 11 years(2011-2021).An evaluation was conducted of the 10 m wind performance of the Climate Forecast System Reanalysis version 2(CFSv2),Modern-Era Retrospective Analysis for Research and Applications version 2(MERRA2),and the fifth-generation European Centre for Medium-Range Weather Forecasts reanalysis products(ERA5)in reproducing wind fields over the Bohai and Yellow Seas.The mean BIAS,root mean square error(RMSE),and correlation coefficient(R)were calculated and used as performance indicators.All three reanalysis products demonstrate commendable performance.The correlation coefficients for wind speed and wind direction exceed 0.83 and 0.90,respectively.Of the three products,the fifth-generation European Centre for Medium-Range Weather Forecasts reanalysis demonstrates the highest level of performance.The model demonstrated the lowest average BIAS and RMSE for wind speed,and the highest R for wind direction.The findings indicate that product performance is subject to variation in both season and wind speed.
基金supported by Princess Nourah bint Abdulrahman University,Riyadh,Saudi Arabia through the Researchers Supporting Project PNURSP2026R760.
摘要Speech Emotion Recognition(SER)is a critical component of affective computing with broad applications in human–computer interaction,mental health monitoring,and intelligent multimedia systems.However,SER remains challenging due to the emotional ambiguity,lack of labeled data,class imbalance,and speaker variability.This study presents an effective SER framework that integrates contrastive representation learning,optimized spectrogram-based data augmentation,and selective synthetic data generation by using TimeGAN to enhance emotion classification performance.Contrastive learning enables the model to better discriminate acoustically similar emotions while Optuna automatically tunes augmentation strategies such as noise injection,time shifting,and time-frequency masking.Unlike existing approaches that apply synthetic generation uniformly across all classes,the proposed method targets only confusing or under-represented emotion classes to preserve the inter-class separability.A CNN-BiLSTM architecture is used to extract spectral and temporal information of the speech.The framework is evaluated with benchmark SER datasets—EMO-DB and RAVDESS—under speaker independent protocols.Experimental results demonstrate improved accuracy,robustness,and generalization under limited and imbalanced data conditions,supported by confusion matrices,UMAP,and t-SNE visualizations.
基金supported by the Zhongshan TCM Heritage and Innovation Research Program(No.2024B3006)the Peak-Shaping Project under Guangzhou University of Chinese Medicine's Action Plan for Double First-Class and High-Level Disciplinary Development(No.GZY2025ZJ18)+1 种基金the Sanming Project of Medicine in Shenzhen(No.SZZYSM202311015)the Shenzhen Medical Research Fund(No.C2501027).
摘要Acupuncture research increasingly involves heterogeneous and multimodal data that are difficult to analyze using conventional methods.This review summarizes data-driven approaches in acupuncture research within a framework encompassing intervention,response,and contextual data.We discuss causal inference,artificial intelligence,text mining,and integrative analysis,along with their applications in efficacy evaluation,outcome prediction,mechanistic investigation,and clinical decision support.These approaches shift the focus of acupuncture research from population-level average effects toward individualized clinical decision-making by enabling the analysis of treatment heterogeneity and underlying mechanisms.However,current research remains limited by inadequate data standardization,insufficient external validation,and limited model interpretability.Despite these challenges,data-driven approaches offer substantial promise for advancing more rigorous and personalized acupuncture research.
基金supported by the National Natural Science Foundation of China(Grant No.62172033).
摘要Urban traffic generates massive and diverse data,yet most systems remain fragmented.Current approaches to congestion management suffer from weak data consistency and poor scalability.This study addresses this gap by proposing the Urban Traffic Congestion Unified Metadata Model(UTC-UMM).The goal is to provide a standardized and extensible framework for describing,extracting,and storing multisource traffic data in smart cities.The model defines a two-tier specification that organizes nine core traffic resource classes.It employs an eXtensible Markup Language(XML)Schema that connects general elements with resource-specific elements.This design ensures both syntactic and semantic interoperability across siloed datasets.Extension principles allow new elements or constraints to be introducedwithout breaking backward compatibility.Adistributed pipeline is implemented usingHadoop Distributed File System(HDFS)and HBase.It integrates computer vision for video and natural language processing for text to automate metadata extraction.Optimized row-key designs enable low-latency queries.Performance is tested with the Yahoo!Cloud Serving Benchmark(YCSB),which shows linear scalability and high throughput.The results demonstrate that UTC-UMM can unify heterogeneous traffic data while supporting real-time analytics.The discussion highlights its potential to improve data reuse,portability,and scalability in urban congestion studies.Future research will explore integration with association rulemining and advanced knowledge representation to capture richer spatiotemporal traffic patterns.
摘要The publisher regrets the CRediT authorship contribution statement was inserted incorrectly and the correct statement should be updated as below:Zengji Liu:Writing-review&editing,Writing-original draft,Visualization,Validation,Supervision,Software,Resources,Project administration,Methodology,Investigation,Funding acquisition,Formal analysis,Data curation,Conceptualization.Mengge Liu:Writing-review&editing,Writing-original draft,Investigation.Qi Wang:Writing-review&editing,Writing-original draft.Yi Tang:Writing-review&editing,Writing-original draft.