Randomness and nonlinearity are essential properties of the real world,and their interaction gives rise to highly complex phenomena.With the advancement of technology,merely observing data of the current system state ...Randomness and nonlinearity are essential properties of the real world,and their interaction gives rise to highly complex phenomena.With the advancement of technology,merely observing data of the current system state is no longer sufficient for prediction and application in various fields.Consequently,extracting the nonlinear evolution nature of the system from noisy data has become a prominent and challenging issue.To address this,we propose an integrated approach that combines data-driven stochastic model identification with a knowledge-based model predictive control strategy.By leveraging high-precision model identification,our data-driven control design is particularly effective for continuous target tracking problems that are difficult to address using traditional precise-model-based control theory.Furthermore,the central challenge in data science lies in maximizing the informational value of datasets while minimizing the effects of observation noise.In this study,we propose and rigorously demonstrate the stochastic Occam’s razor principle,a stochastic error estimation theory that evaluates and enhances the design of data-driven schemes to mitigate the effect of observation noise.Notably,our approach offers valuable insights for contemporary data-driven,end-to-end control challenges,particularly those involving uncertain governing equations and substantial non-Gaussian observation noise.展开更多
With the rapid advancement of machine learning technology and its growing adoption in research and engineering applications,an increasing number of studies have embraced data-driven approaches for modeling wind turbin...With the rapid advancement of machine learning technology and its growing adoption in research and engineering applications,an increasing number of studies have embraced data-driven approaches for modeling wind turbine wakes.These models leverage the ability to capture complex,high-dimensional characteristics of wind turbine wakes while offering significantly greater efficiency in the prediction process than physics-driven models.As a result,data-driven wind turbine wake models are regarded as powerful and effective tools for predicting wake behavior and turbine power output.This paper aims to provide a concise yet comprehensive review of existing studies on wind turbine wake modeling that employ data-driven approaches.It begins by defining and classifying machine learning methods to facilitate a clearer understanding of the reviewed literature.Subsequently,the related studies are categorized into four key areas:wind turbine power prediction,data-driven analytic wake models,wake field reconstruction,and the incorporation of explicit physical constraints.The accuracy of data-driven models is influenced by two primary factors:the quality of the training data and the performance of the model itself.Accordingly,both data accuracy and model structure are discussed in detail within the review.展开更多
The distillation process is an important chemical process,and the application of data-driven modelling approach has the potential to reduce model complexity compared to mechanistic modelling,thus improving the efficie...The distillation process is an important chemical process,and the application of data-driven modelling approach has the potential to reduce model complexity compared to mechanistic modelling,thus improving the efficiency of process optimization or monitoring studies.However,the distillation process is highly nonlinear and has multiple uncertainty perturbation intervals,which brings challenges to accurate data-driven modelling of distillation processes.This paper proposes a systematic data-driven modelling framework to solve these problems.Firstly,data segment variance was introduced into the K-means algorithm to form K-means data interval(KMDI)clustering in order to cluster the data into perturbed and steady state intervals for steady-state data extraction.Secondly,maximal information coefficient(MIC)was employed to calculate the nonlinear correlation between variables for removing redundant features.Finally,extreme gradient boosting(XGBoost)was integrated as the basic learner into adaptive boosting(AdaBoost)with the error threshold(ET)set to improve weights update strategy to construct the new integrated learning algorithm,XGBoost-AdaBoost-ET.The superiority of the proposed framework is verified by applying this data-driven modelling framework to a real industrial process of propylene distillation.展开更多
In industrial production,the acquisition of critical quality variables often faces significant challenges due to high costs and data scarcity,which not only limit the improvement of production efficiency but also incr...In industrial production,the acquisition of critical quality variables often faces significant challenges due to high costs and data scarcity,which not only limit the improvement of production efficiency but also increase the difficulty of quality control.With the advent of the industrial big data era,the availability and diversity of data have greatly increased,offering opportunities to address these issues.To address the problem of data scarcity,this paper proposes a novel data augmentation method for soft sensing—FVAE-WGAN,which generates high-quality synthetic data to expand the training dataset of soft sensors,thereby enhancing their prediction accuracy and generalization capability.This method integrates two stacked variational autoencoder(VAE)models with a Wasserstein generative adversarial network(WGAN),constructing a generator capable of learning from a broader data distribution.Additionally,an encoder is embedded in the discriminator,enhancing the model's ability to utilize late nt features of the data.By freezing specific layers of the discriminato r,the pro posed method reduces computational resource consumption during training and effectively mitigates overfitting.Experiments conducted on industrial process datasets show that the FVAE-WGAN model outperforms comparative models in terms of accuracy and robustness.This approach not only alleviates the impact of data scarcity,but also optimizes the efficiency and reliability of industrial processes,thereby bringing substantial economic benefits to industrial production.展开更多
We propose an integrated method of data-driven and mechanism models for well logging formation evaluation,explicitly focusing on predicting reservoir parameters,such as porosity and water saturation.Accurately interpr...We propose an integrated method of data-driven and mechanism models for well logging formation evaluation,explicitly focusing on predicting reservoir parameters,such as porosity and water saturation.Accurately interpreting these parameters is crucial for effectively exploring and developing oil and gas.However,with the increasing complexity of geological conditions in this industry,there is a growing demand for improved accuracy in reservoir parameter prediction,leading to higher costs associated with manual interpretation.The conventional logging interpretation methods rely on empirical relationships between logging data and reservoir parameters,which suffer from low interpretation efficiency,intense subjectivity,and suitability for ideal conditions.The application of artificial intelligence in the interpretation of logging data provides a new solution to the problems existing in traditional methods.It is expected to improve the accuracy and efficiency of the interpretation.If large and high-quality datasets exist,data-driven models can reveal relationships of arbitrary complexity.Nevertheless,constructing sufficiently large logging datasets with reliable labels remains challenging,making it difficult to apply data-driven models effectively in logging data interpretation.Furthermore,data-driven models often act as“black boxes”without explaining their predictions or ensuring compliance with primary physical constraints.This paper proposes a machine learning method with strong physical constraints by integrating mechanism and data-driven models.Prior knowledge of logging data interpretation is embedded into machine learning regarding network structure,loss function,and optimization algorithm.We employ the Physically Informed Auto-Encoder(PIAE)to predict porosity and water saturation,which can be trained without labeled reservoir parameters using self-supervised learning techniques.This approach effectively achieves automated interpretation and facilitates generalization across diverse datasets.展开更多
Identifying an interpretable and tractable model is a crucial step for the analysis and control of dynamical systems.In this work,we employ the recently introduced Kolmogorov–Arnold Networks(KANs),a novel neural netw...Identifying an interpretable and tractable model is a crucial step for the analysis and control of dynamical systems.In this work,we employ the recently introduced Kolmogorov–Arnold Networks(KANs),a novel neural network architecture tailored for symbolic regression and interpretability,to learn symbolic models directly from data without any a priori knowledge of the observed dynamics.We then extend our result to the distributed case through federated learning introducing the FedKANs algorithm.FedKANs allows agents observing similar,but non-identical,systems to cooperate to learn more efficiently a symbolic model without the need to exchange any process data.To our knowledge,this represents the first distributed deep learning framework for symbolic regression in dynamical systems.Numerical simulations validate the proposed solutions in various settings,involving linear,nonlinear,discrete-time and continuous-time dynamics.展开更多
The constitutive models of shape memory alloys(SMAs)play an important role in facilitating the widespread application of such types of alloys in various engineering fields.However,to accurately describe the deformatio...The constitutive models of shape memory alloys(SMAs)play an important role in facilitating the widespread application of such types of alloys in various engineering fields.However,to accurately describe the deformation behaviors of SMAs,the concepts in classical plasticity are employed in the existing constitutive models,and a series of complex mathematical equations are involved.Such complexity brings inconvenience for the construction,implementation,and application of the constitutive models.To overcome these shortcomings,a data-driven constitutive model of SMAs is developed in this work based on the artificial neural network(ANN).In the proposed model,the components of the strain tensor in principal space,ambient temperature,and the maximum equivalent strain in the deformation history from the initial state to the current loading state are chosen as the input features,and the components of the stress tensor in principal space are set as the output.The proposed ANN-based constitutive model is implemented into the finite element program ABAQUS by deriving its consistent tangent modulus and writing a user-defined material subroutine.The stress-strain responses of SMA material under various loading paths and at different ambient temperatures are used to train the ANN model,which is generated from the existing constitutive model(numerical experiments).To validate the capability of the proposed model,the predicted stress-strain responses of SMA material,and the global and local responses of two typical SMA structures are compared with the corresponding numerical experiments.This work demonstrates a good potential to obtain the constitutive model of SMAs by pure data and avoid the need for vast stores of knowledge for the construction of constitutive models.展开更多
To ensure the safe operation of batteries,accurately obtaining key internal state parameters is essential.However,traditional parameter measurement methods either require opening the battery or long-term measurements,...To ensure the safe operation of batteries,accurately obtaining key internal state parameters is essential.However,traditional parameter measurement methods either require opening the battery or long-term measurements,which are impractical.Therefore,the fixed values are commonly used for these parameters in electrochemical models and have significant limitations.To overcome these limitations,this paper proposes a deep neural network(DNN)based data-driven evaluation method to determine model parameters.By coupling an improved one-dimensional isothermal pseudo-twodimensional(P2D)model with DNN,this study identified concentration-dependent parameters through detailed discharge curve analysis.The results show that the data-driven method can effectively obtain the change trend of concentration-dependent parameters through the charge and discharge curve,and the method can be extended to different battery systems in different discharge rates and aging applications.This work is expected to provide new parameter selection insights for data-driven battery prediction and monitoring models.展开更多
Permanent magnet synchronous motor(PMSM)is widely used in alternating current servo systems as it provides high eficiency,high power density,and a wide speed regulation range.The servo system is placing higher demands...Permanent magnet synchronous motor(PMSM)is widely used in alternating current servo systems as it provides high eficiency,high power density,and a wide speed regulation range.The servo system is placing higher demands on its control performance.The model predictive control(MPC)algorithm is emerging as a potential high-performance motor control algorithm due to its capability of handling multiple-input and multipleoutput variables and imposed constraints.For the MPC used in the PMSM control process,there is a nonlinear disturbance caused by the change of electromagnetic parameters or load disturbance that may lead to a mismatch between the nominal model and the controlled object,which causes the prediction error and thus affects the dynamic stability of the control system.This paper proposes a data-driven MPC strategy in which the historical data in an appropriate range are utilized to eliminate the impact of parameter mismatch and further improve the control performance.The stability of the proposed algorithm is proved as the simulation demonstrates the feasibility.Compared with the classical MPC strategy,the superiority of the algorithm has also been verified.展开更多
A data-driven optimization framework that integrates machine learning surrogate models,finite element analysis(FEA),and a multi-objective optimization algorithm is used in this study for developing thermoplastic elast...A data-driven optimization framework that integrates machine learning surrogate models,finite element analysis(FEA),and a multi-objective optimization algorithm is used in this study for developing thermoplastic elastomer(TPE)parts for aerospace applications.By using FEA simulations and experiments,a database of input design parameters(e.g.,geometry and structural shape modifier)is generated.Afterwards,we train surrogate models(e.g.,Gaussian Process Regression,neural networks)to approximate mappings from design space to performance space.Finally,we propose Pareto-optimal TPE designs using the surrogate embedded in a multi-objective optimization loop(such as NSGA-Ⅱ or gradient-based methods).The novelty of this approach is demonstrated by employing highly simplified surrogate models,including an artificial neural network(ANN)with 10 hidden neurons trained on analytically generated synthetic data.The proposed methodology has been validated using an aerospace-related case study:a vibration-damping plate.Compared with the baseline configuration,Pareto-optimal designs identified by the proposed framework achieved a reduction in maximum deflection of 23%-28%and a reduction in von Mises stress of 18%-24%,depending on the selected trade-off solution,as the number of full FEA simulations required for optimization was reduced from 500 to 50.This framework enables faster design of TPE components for aerospace systems.Validation against high-fidelity ANSYS simulations showed a mean error of~1.18%and a maximum deviation of~2.6%.展开更多
A data-driven model predictive control(MPC)algorithm based on the input-mapping method is proposed for piecewise affine(PWA)systems.These systems are characterized by unknown but constant parameters and are subject to...A data-driven model predictive control(MPC)algorithm based on the input-mapping method is proposed for piecewise affine(PWA)systems.These systems are characterized by unknown but constant parameters and are subject to disturbances,as well as state and input constraints.To support the control strategy,an offline algorithm is developed to compute a non-convex robust positively invariant set that serves as the terminal set within the MPC framework tailored for PWA systems.The online MPC algorithm directly maps the future control input and predicted state to the historical input-state data associated with the corresponding state subregion.This mapping process leverages the more accurate relationships contained in the historical input-state data to enhance the prediction accuracy of future states.A state-dependent weight embedded in the cost function enables the controller to balance prediction accuracy against convergence speed,enhancing overall performance.Moreover,conditions ensuring the recursive feasibility of the optimization problem and stability of the closed-loop system are established.The effectiveness of the proposed algorithm is demonstrated through a numerical example,which highlights its ability to handle complex system dynamics and constraints while maintaining robust performance.展开更多
With the rising water cut in mature oil fields,polymer flooding has emerged as a critical Enhanced Oil Recovery(EOR)technique.However,high-fidelity numerical simulations for history matching and polymer flooding optim...With the rising water cut in mature oil fields,polymer flooding has emerged as a critical Enhanced Oil Recovery(EOR)technique.However,high-fidelity numerical simulations for history matching and polymer flooding optimization remain computationally intensive,limiting their practicality for ClosedLoop Reservoir Management(CLRM),which is inherently dependent on rapid iterative simulations for real-time model updating and operational decision-making.Although physics-based data-driven flownetwork models,such as General-Purpose Simulator-powered Network model(GPSNet),can accelerate simulations,their lack of geological constraints compromises predictive reliability.To address this limitation,we propose a novel facies-constrained flow-network model(GPSNet-FC)within the GPSNet framework.This model simplifies reservoir geometry into a 1D discretized grid between wells while incorporating sedimentary facies boundaries identified through edge detection and level-set methods.Grid properties are assigned and calibrated based on facies-specific attributes to ensure geological consistency.GPSNet-FC is applied to history matching using the Ensemble Smoother with Multiple Data Assimilation(ESMDA)and to polymer flooding optimization via the Differential Evolution(DE)algorithm.Numerical case studies validate the method,demonstrating that GPSNet-FC outperforms the original GPSNet in both reliability and accuracy.By integrating facies-based geological constraints,this approach reduces non-uniqueness in history matching and enables rapid and accurate decision-making fo r polymer flooding strategies.This work advances the integration of geological data into physics-based data-driven models,offering a robust and efficient tool for the CLRM of polymer flooding reservoirs.展开更多
This paper focuses on the numerical solution of a tumor growth model under a data-driven approach.Based on the inherent laws of the data and reasonable assumptions,an ordinary differential equation model for tumor gro...This paper focuses on the numerical solution of a tumor growth model under a data-driven approach.Based on the inherent laws of the data and reasonable assumptions,an ordinary differential equation model for tumor growth is established.Nonlinear fitting is employed to obtain the optimal parameter estimation of the mathematical model,and the numerical solution is carried out using the Matlab software.By comparing the clinical data with the simulation results,a good agreement is achieved,which verifies the rationality and feasibility of the model.展开更多
This paper focuses on the development of smart construction sites, providing a detailed exploration of how IoT technology can drive innovation and improvement in management practices. It first clarifies the fundamenta...This paper focuses on the development of smart construction sites, providing a detailed exploration of how IoT technology can drive innovation and improvement in management practices. It first clarifies the fundamental concepts and historical context of smart construction sites, emphasizing the critical role of IoT data in enhancing the precision and intelligence of site management. The study further highlights that such management enhancements have become an inevitable trend. Addressing prominent challenges in current smart construction management—including decentralized data collection, severe information silos, low collaboration efficiency between systems, and traditional methods' inadequacy in meeting dynamic construction demands—the paper conducts thorough analysis and research. To tackle these issues, researchers have developed a data-driven management improvement framework supported by IoT technologies. This system encompasses comprehensive implementation strategies for data collection and transmission, establishment of a unified data center integrating multi-source information, and advanced data applications throughout the entire construction process. The paper elaborates on leveraging data to drive management innovation, proposing concrete implementation approaches such as real-time monitoring of worker conditions, machinery operations, and material usage patterns with proactive risk alerts. Finally, it advocates for data sharing to facilitate efficient collaboration among project stakeholders, optimize resource allocation, and ensure successful project execution and achievement of objectives. This paper conducts an in-depth and comprehensive analysis of the practical effectiveness of innovative management models, focusing on specific measures to ensure data security and effectively promote standardization.展开更多
Increasing the production and utilization of shale gas is of great significance for building a clean and low-carbon energy system.Sharp decline of gas production has been widely observed in shale gas reservoirs.How to...Increasing the production and utilization of shale gas is of great significance for building a clean and low-carbon energy system.Sharp decline of gas production has been widely observed in shale gas reservoirs.How to forecast shale gas production is still challenging due to complex fracture networks,dynamic fracture properties,frac hits,complicated multiphase flow,and multi-scale flow as well as data quality and uncertainty.This work develops an integrated framework for evaluating shale gas well production based on data-driven models.Firstly,a comprehensive dominated-factor system has been established,including geological,drilling,fracturing,and production factors.Data processing and visualization are required to ensure data quality and determine final data set.A shale gas production evaluation model is developed to evaluate shale gas production levels.Finally,the random forest algorithm is used to forecast shale gas production.The prediction accuracy of shale gas production level is higher than 95%based on the shale gas reservoirs in China.Forty-one wells are randomly selected to predict cumulative gas production using the optimal regression model.The proposed shale gas production evaluation frame-work overcomes too many assumptions of analytical or semi-analytical models and avoids huge computation cost and poor generalization for numerical modelling.展开更多
The world’s increasing population requires the process industry to produce food,fuels,chemicals,and consumer products in a more efficient and sustainable way.Functional process materials lie at the heart of this chal...The world’s increasing population requires the process industry to produce food,fuels,chemicals,and consumer products in a more efficient and sustainable way.Functional process materials lie at the heart of this challenge.Traditionally,new advanced materials are found empirically or through trial-and-error approaches.As theoretical methods and associated tools are being continuously improved and computer power has reached a high level,it is now efficient and popular to use computational methods to guide material selection and design.Due to the strong interaction between material selection and the operation of the process in which the material is used,it is essential to perform material and process design simultaneously.Despite this significant connection,the solution of the integrated material and process design problem is not easy because multiple models at different scales are usually required.Hybrid modeling provides a promising option to tackle such complex design problems.In hybrid modeling,the material properties,which are computationally expensive to obtain,are described by data-driven models,while the well-known process-related principles are represented by mechanistic models.This article highlights the significance of hybrid modeling in multiscale material and process design.The generic design methodology is first introduced.Six important application areas are then selected:four from the chemical engineering field and two from the energy systems engineering domain.For each selected area,state-ofthe-art work using hybrid modeling for multiscale material and process design is discussed.Concluding remarks are provided at the end,and current limitations and future opportunities are pointed out.展开更多
Aerodynamic surrogate modeling mostly relies only on integrated loads data obtained from simulation or experiment,while neglecting and wasting the valuable distributed physical information on the surface.To make full ...Aerodynamic surrogate modeling mostly relies only on integrated loads data obtained from simulation or experiment,while neglecting and wasting the valuable distributed physical information on the surface.To make full use of both integrated and distributed loads,a modeling paradigm,called the heterogeneous data-driven aerodynamic modeling,is presented.The essential concept is to incorporate the physical information of distributed loads as additional constraints within the end-to-end aerodynamic modeling.Towards heterogenous data,a novel and easily applicable physical feature embedding modeling framework is designed.This framework extracts lowdimensional physical features from pressure distribution and then effectively enhances the modeling of the integrated loads via feature embedding.The proposed framework can be coupled with multiple feature extraction methods,and the well-performed generalization capabilities over different airfoils are verified through a transonic case.Compared with traditional direct modeling,the proposed framework can reduce testing errors by almost 50%.Given the same prediction accuracy,it can save more than half of the training samples.Furthermore,the visualization analysis has revealed a significant correlation between the discovered low-dimensional physical features and the heterogeneous aerodynamic loads,which shows the interpretability and credibility of the superior performance offered by the proposed deep learning framework.展开更多
The complex sand-casting process combined with the interactions between process parameters makes it difficult to control the casting quality,resulting in a high scrap rate.A strategy based on a data-driven model was p...The complex sand-casting process combined with the interactions between process parameters makes it difficult to control the casting quality,resulting in a high scrap rate.A strategy based on a data-driven model was proposed to reduce casting defects and improve production efficiency,which includes the random forest(RF)classification model,the feature importance analysis,and the process parameters optimization with Monte Carlo simulation.The collected data includes four types of defects and corresponding process parameters were used to construct the RF model.Classification results show a recall rate above 90% for all categories.The Gini Index was used to assess the importance of the process parameters in the formation of various defects in the RF model.Finally,the classification model was applied to different production conditions for quality prediction.In the case of process parameters optimization for gas porosity defects,this model serves as an experimental process in the Monte Carlo method to estimate a better temperature distribution.The prediction model,when applied to the factory,greatly improved the efficiency of defect detection.Results show that the scrap rate decreased from 10.16% to 6.68%.展开更多
Vortex induced vibration(VIV)is a challenge in ocean engineering.Several devices including fairings have been designed to suppress VIV.However,how to optimize the design of suppression devices is still a problem to be...Vortex induced vibration(VIV)is a challenge in ocean engineering.Several devices including fairings have been designed to suppress VIV.However,how to optimize the design of suppression devices is still a problem to be solved.In this paper,an optimization design methodology is presented based on data-driven models and genetic algorithm(GA).Data-driven models are introduced to substitute complex physics-based equations.GA is used to rapidly search for the optimal suppression device from all possible solutions.Taking fairings as example,VIV response database for different fairings is established based on parameterized models in which model sections of fairings are controlled by several control points and Bezier curves.Then a data-driven model,which can predict the VIV response of fairings with different sections accurately and efficiently,is trained through BP neural network.Finally,a comprehensive optimization method and process is proposed based on GA and the data-driven model.The proposed method is demonstrated by its application to a case.It turns out that the proposed method can perform the optimization design of fairings effectively.VIV can be reduced obviously through the optimization design.展开更多
This study explores the effectiveness of machine learning models in predicting the air-side performance of microchannel heat exchangers.The data were generated by experimentally validated Computational Fluid Dynam-ics...This study explores the effectiveness of machine learning models in predicting the air-side performance of microchannel heat exchangers.The data were generated by experimentally validated Computational Fluid Dynam-ics(CFD)simulations of air-to-water microchannel heat exchangers.A distinctive aspect of this research is the comparative analysis of four diverse machine learning algorithms:Artificial Neural Networks(ANN),Support Vector Machines(SVM),Random Forest(RF),and Gaussian Process Regression(GPR).These models are adeptly applied to predict air-side heat transfer performance with high precision,with ANN and GPR exhibiting notably superior accuracy.Additionally,this research further delves into the influence of both geometric and operational parameters—including louvered angle,fin height,fin spacing,air inlet temperature,velocity,and tube temperature—on model performance.Moreover,it innovatively incorporates dimensionless numbers such as aspect ratio,fin height-to-spacing ratio,Reynolds number,Nusselt number,normalized air inlet temperature,temperature difference,and louvered angle into the input variables.This strategic inclusion significantly refines the predictive capabilities of the models by establishing a robust analytical framework supported by the CFD-generated database.The results show the enhanced prediction accuracy achieved by integrating dimensionless numbers,highlighting the effectiveness of data-driven approaches in precisely forecasting heat exchanger performance.This advancement is pivotal for the geometric optimization of heat exchangers,illustrating the considerable potential of integrating sophisticated modeling techniques with traditional engineering metrics.展开更多
基金supported by the National Natural Science Foundation of China(Grant No.12172167).
摘要Randomness and nonlinearity are essential properties of the real world,and their interaction gives rise to highly complex phenomena.With the advancement of technology,merely observing data of the current system state is no longer sufficient for prediction and application in various fields.Consequently,extracting the nonlinear evolution nature of the system from noisy data has become a prominent and challenging issue.To address this,we propose an integrated approach that combines data-driven stochastic model identification with a knowledge-based model predictive control strategy.By leveraging high-precision model identification,our data-driven control design is particularly effective for continuous target tracking problems that are difficult to address using traditional precise-model-based control theory.Furthermore,the central challenge in data science lies in maximizing the informational value of datasets while minimizing the effects of observation noise.In this study,we propose and rigorously demonstrate the stochastic Occam’s razor principle,a stochastic error estimation theory that evaluates and enhances the design of data-driven schemes to mitigate the effect of observation noise.Notably,our approach offers valuable insights for contemporary data-driven,end-to-end control challenges,particularly those involving uncertain governing equations and substantial non-Gaussian observation noise.
基金Supported by the National Natural Science Foundation of China under Grant No.52131102.
摘要With the rapid advancement of machine learning technology and its growing adoption in research and engineering applications,an increasing number of studies have embraced data-driven approaches for modeling wind turbine wakes.These models leverage the ability to capture complex,high-dimensional characteristics of wind turbine wakes while offering significantly greater efficiency in the prediction process than physics-driven models.As a result,data-driven wind turbine wake models are regarded as powerful and effective tools for predicting wake behavior and turbine power output.This paper aims to provide a concise yet comprehensive review of existing studies on wind turbine wake modeling that employ data-driven approaches.It begins by defining and classifying machine learning methods to facilitate a clearer understanding of the reviewed literature.Subsequently,the related studies are categorized into four key areas:wind turbine power prediction,data-driven analytic wake models,wake field reconstruction,and the incorporation of explicit physical constraints.The accuracy of data-driven models is influenced by two primary factors:the quality of the training data and the performance of the model itself.Accordingly,both data accuracy and model structure are discussed in detail within the review.
基金supported by the National Key Research and Development Program of China(2023YFB3307801)the National Natural Science Foundation of China(62394343,62373155,62073142)+3 种基金Major Science and Technology Project of Xinjiang(No.2022A01006-4)the Programme of Introducing Talents of Discipline to Universities(the 111 Project)under Grant B17017the Fundamental Research Funds for the Central Universities,Science Foundation of China University of Petroleum,Beijing(No.2462024YJRC011)the Open Research Project of the State Key Laboratory of Industrial Control Technology,China(Grant No.ICT2024B70).
摘要The distillation process is an important chemical process,and the application of data-driven modelling approach has the potential to reduce model complexity compared to mechanistic modelling,thus improving the efficiency of process optimization or monitoring studies.However,the distillation process is highly nonlinear and has multiple uncertainty perturbation intervals,which brings challenges to accurate data-driven modelling of distillation processes.This paper proposes a systematic data-driven modelling framework to solve these problems.Firstly,data segment variance was introduced into the K-means algorithm to form K-means data interval(KMDI)clustering in order to cluster the data into perturbed and steady state intervals for steady-state data extraction.Secondly,maximal information coefficient(MIC)was employed to calculate the nonlinear correlation between variables for removing redundant features.Finally,extreme gradient boosting(XGBoost)was integrated as the basic learner into adaptive boosting(AdaBoost)with the error threshold(ET)set to improve weights update strategy to construct the new integrated learning algorithm,XGBoost-AdaBoost-ET.The superiority of the proposed framework is verified by applying this data-driven modelling framework to a real industrial process of propylene distillation.
基金supported by the National Natural Science Foundation of China(62341314)。
摘要In industrial production,the acquisition of critical quality variables often faces significant challenges due to high costs and data scarcity,which not only limit the improvement of production efficiency but also increase the difficulty of quality control.With the advent of the industrial big data era,the availability and diversity of data have greatly increased,offering opportunities to address these issues.To address the problem of data scarcity,this paper proposes a novel data augmentation method for soft sensing—FVAE-WGAN,which generates high-quality synthetic data to expand the training dataset of soft sensors,thereby enhancing their prediction accuracy and generalization capability.This method integrates two stacked variational autoencoder(VAE)models with a Wasserstein generative adversarial network(WGAN),constructing a generator capable of learning from a broader data distribution.Additionally,an encoder is embedded in the discriminator,enhancing the model's ability to utilize late nt features of the data.By freezing specific layers of the discriminato r,the pro posed method reduces computational resource consumption during training and effectively mitigates overfitting.Experiments conducted on industrial process datasets show that the FVAE-WGAN model outperforms comparative models in terms of accuracy and robustness.This approach not only alleviates the impact of data scarcity,but also optimizes the efficiency and reliability of industrial processes,thereby bringing substantial economic benefits to industrial production.
基金supported by National Key Research and Development Program (2019YFA0708301)National Natural Science Foundation of China (51974337)+2 种基金the Strategic Cooperation Projects of CNPC and CUPB (ZLZX2020-03)Science and Technology Innovation Fund of CNPC (2021DQ02-0403)Open Fund of Petroleum Exploration and Development Research Institute of CNPC (2022-KFKT-09)
摘要We propose an integrated method of data-driven and mechanism models for well logging formation evaluation,explicitly focusing on predicting reservoir parameters,such as porosity and water saturation.Accurately interpreting these parameters is crucial for effectively exploring and developing oil and gas.However,with the increasing complexity of geological conditions in this industry,there is a growing demand for improved accuracy in reservoir parameter prediction,leading to higher costs associated with manual interpretation.The conventional logging interpretation methods rely on empirical relationships between logging data and reservoir parameters,which suffer from low interpretation efficiency,intense subjectivity,and suitability for ideal conditions.The application of artificial intelligence in the interpretation of logging data provides a new solution to the problems existing in traditional methods.It is expected to improve the accuracy and efficiency of the interpretation.If large and high-quality datasets exist,data-driven models can reveal relationships of arbitrary complexity.Nevertheless,constructing sufficiently large logging datasets with reliable labels remains challenging,making it difficult to apply data-driven models effectively in logging data interpretation.Furthermore,data-driven models often act as“black boxes”without explaining their predictions or ensuring compliance with primary physical constraints.This paper proposes a machine learning method with strong physical constraints by integrating mechanism and data-driven models.Prior knowledge of logging data interpretation is embedded into machine learning regarding network structure,loss function,and optimization algorithm.We employ the Physically Informed Auto-Encoder(PIAE)to predict porosity and water saturation,which can be trained without labeled reservoir parameters using self-supervised learning techniques.This approach effectively achieves automated interpretation and facilitates generalization across diverse datasets.
摘要Identifying an interpretable and tractable model is a crucial step for the analysis and control of dynamical systems.In this work,we employ the recently introduced Kolmogorov–Arnold Networks(KANs),a novel neural network architecture tailored for symbolic regression and interpretability,to learn symbolic models directly from data without any a priori knowledge of the observed dynamics.We then extend our result to the distributed case through federated learning introducing the FedKANs algorithm.FedKANs allows agents observing similar,but non-identical,systems to cooperate to learn more efficiently a symbolic model without the need to exchange any process data.To our knowledge,this represents the first distributed deep learning framework for symbolic regression in dynamical systems.Numerical simulations validate the proposed solutions in various settings,involving linear,nonlinear,discrete-time and continuous-time dynamics.
基金supported by the National Natural Science Foundation of China(NSFC)(Grant No.12322203).
摘要The constitutive models of shape memory alloys(SMAs)play an important role in facilitating the widespread application of such types of alloys in various engineering fields.However,to accurately describe the deformation behaviors of SMAs,the concepts in classical plasticity are employed in the existing constitutive models,and a series of complex mathematical equations are involved.Such complexity brings inconvenience for the construction,implementation,and application of the constitutive models.To overcome these shortcomings,a data-driven constitutive model of SMAs is developed in this work based on the artificial neural network(ANN).In the proposed model,the components of the strain tensor in principal space,ambient temperature,and the maximum equivalent strain in the deformation history from the initial state to the current loading state are chosen as the input features,and the components of the stress tensor in principal space are set as the output.The proposed ANN-based constitutive model is implemented into the finite element program ABAQUS by deriving its consistent tangent modulus and writing a user-defined material subroutine.The stress-strain responses of SMA material under various loading paths and at different ambient temperatures are used to train the ANN model,which is generated from the existing constitutive model(numerical experiments).To validate the capability of the proposed model,the predicted stress-strain responses of SMA material,and the global and local responses of two typical SMA structures are compared with the corresponding numerical experiments.This work demonstrates a good potential to obtain the constitutive model of SMAs by pure data and avoid the need for vast stores of knowledge for the construction of constitutive models.
基金supported by National Natural Science Foundation of China(22478239)Science and Technology Commission of Shanghai Municipality(19DZ2271100)National Natural Science Foundation of China(22208208)。
摘要To ensure the safe operation of batteries,accurately obtaining key internal state parameters is essential.However,traditional parameter measurement methods either require opening the battery or long-term measurements,which are impractical.Therefore,the fixed values are commonly used for these parameters in electrochemical models and have significant limitations.To overcome these limitations,this paper proposes a deep neural network(DNN)based data-driven evaluation method to determine model parameters.By coupling an improved one-dimensional isothermal pseudo-twodimensional(P2D)model with DNN,this study identified concentration-dependent parameters through detailed discharge curve analysis.The results show that the data-driven method can effectively obtain the change trend of concentration-dependent parameters through the charge and discharge curve,and the method can be extended to different battery systems in different discharge rates and aging applications.This work is expected to provide new parameter selection insights for data-driven battery prediction and monitoring models.
摘要Permanent magnet synchronous motor(PMSM)is widely used in alternating current servo systems as it provides high eficiency,high power density,and a wide speed regulation range.The servo system is placing higher demands on its control performance.The model predictive control(MPC)algorithm is emerging as a potential high-performance motor control algorithm due to its capability of handling multiple-input and multipleoutput variables and imposed constraints.For the MPC used in the PMSM control process,there is a nonlinear disturbance caused by the change of electromagnetic parameters or load disturbance that may lead to a mismatch between the nominal model and the controlled object,which causes the prediction error and thus affects the dynamic stability of the control system.This paper proposes a data-driven MPC strategy in which the historical data in an appropriate range are utilized to eliminate the impact of parameter mismatch and further improve the control performance.The stability of the proposed algorithm is proved as the simulation demonstrates the feasibility.Compared with the classical MPC strategy,the superiority of the algorithm has also been verified.
摘要A data-driven optimization framework that integrates machine learning surrogate models,finite element analysis(FEA),and a multi-objective optimization algorithm is used in this study for developing thermoplastic elastomer(TPE)parts for aerospace applications.By using FEA simulations and experiments,a database of input design parameters(e.g.,geometry and structural shape modifier)is generated.Afterwards,we train surrogate models(e.g.,Gaussian Process Regression,neural networks)to approximate mappings from design space to performance space.Finally,we propose Pareto-optimal TPE designs using the surrogate embedded in a multi-objective optimization loop(such as NSGA-Ⅱ or gradient-based methods).The novelty of this approach is demonstrated by employing highly simplified surrogate models,including an artificial neural network(ANN)with 10 hidden neurons trained on analytically generated synthetic data.The proposed methodology has been validated using an aerospace-related case study:a vibration-damping plate.Compared with the baseline configuration,Pareto-optimal designs identified by the proposed framework achieved a reduction in maximum deflection of 23%-28%and a reduction in von Mises stress of 18%-24%,depending on the selected trade-off solution,as the number of full FEA simulations required for optimization was reduced from 500 to 50.This framework enables faster design of TPE components for aerospace systems.Validation against high-fidelity ANSYS simulations showed a mean error of~1.18%and a maximum deviation of~2.6%.
基金supported by the National Key Research and Development Project(No.2024YFB4105200)the National Science Foundation of China(Nos.62573284,62333015,62261160385)+1 种基金the Science Foundation of Shanghai(No.24ZR1438800)the China Postdoctoral Science Foundation(No.2025M771696).
摘要A data-driven model predictive control(MPC)algorithm based on the input-mapping method is proposed for piecewise affine(PWA)systems.These systems are characterized by unknown but constant parameters and are subject to disturbances,as well as state and input constraints.To support the control strategy,an offline algorithm is developed to compute a non-convex robust positively invariant set that serves as the terminal set within the MPC framework tailored for PWA systems.The online MPC algorithm directly maps the future control input and predicted state to the historical input-state data associated with the corresponding state subregion.This mapping process leverages the more accurate relationships contained in the historical input-state data to enhance the prediction accuracy of future states.A state-dependent weight embedded in the cost function enables the controller to balance prediction accuracy against convergence speed,enhancing overall performance.Moreover,conditions ensuring the recursive feasibility of the optimization problem and stability of the closed-loop system are established.The effectiveness of the proposed algorithm is demonstrated through a numerical example,which highlights its ability to handle complex system dynamics and constraints while maintaining robust performance.
基金supported by the National Natural Science Foundation of China(Nos.52474067,52441411,52325402,52034010,12131014)Natural Science Foundation of Shandong Province,China(No.ZR2024ME005)+1 种基金Fundamental Research Funds for the Central Universities(Nos.25CX02025A and 21CX06031A)Youth Innovation and Technology Support Program for Higher Education Institutions of Shandong Province,China(No.2022KJ070)。
摘要With the rising water cut in mature oil fields,polymer flooding has emerged as a critical Enhanced Oil Recovery(EOR)technique.However,high-fidelity numerical simulations for history matching and polymer flooding optimization remain computationally intensive,limiting their practicality for ClosedLoop Reservoir Management(CLRM),which is inherently dependent on rapid iterative simulations for real-time model updating and operational decision-making.Although physics-based data-driven flownetwork models,such as General-Purpose Simulator-powered Network model(GPSNet),can accelerate simulations,their lack of geological constraints compromises predictive reliability.To address this limitation,we propose a novel facies-constrained flow-network model(GPSNet-FC)within the GPSNet framework.This model simplifies reservoir geometry into a 1D discretized grid between wells while incorporating sedimentary facies boundaries identified through edge detection and level-set methods.Grid properties are assigned and calibrated based on facies-specific attributes to ensure geological consistency.GPSNet-FC is applied to history matching using the Ensemble Smoother with Multiple Data Assimilation(ESMDA)and to polymer flooding optimization via the Differential Evolution(DE)algorithm.Numerical case studies validate the method,demonstrating that GPSNet-FC outperforms the original GPSNet in both reliability and accuracy.By integrating facies-based geological constraints,this approach reduces non-uniqueness in history matching and enables rapid and accurate decision-making fo r polymer flooding strategies.This work advances the integration of geological data into physics-based data-driven models,offering a robust and efficient tool for the CLRM of polymer flooding reservoirs.
基金National Natural Science Foundation of China(Project No.:12371428)Projects of the Provincial College Students’Innovation and Training Program in 2024(Project No.:S202413023106,S202413023110)。
摘要This paper focuses on the numerical solution of a tumor growth model under a data-driven approach.Based on the inherent laws of the data and reasonable assumptions,an ordinary differential equation model for tumor growth is established.Nonlinear fitting is employed to obtain the optimal parameter estimation of the mathematical model,and the numerical solution is carried out using the Matlab software.By comparing the clinical data with the simulation results,a good agreement is achieved,which verifies the rationality and feasibility of the model.
摘要This paper focuses on the development of smart construction sites, providing a detailed exploration of how IoT technology can drive innovation and improvement in management practices. It first clarifies the fundamental concepts and historical context of smart construction sites, emphasizing the critical role of IoT data in enhancing the precision and intelligence of site management. The study further highlights that such management enhancements have become an inevitable trend. Addressing prominent challenges in current smart construction management—including decentralized data collection, severe information silos, low collaboration efficiency between systems, and traditional methods' inadequacy in meeting dynamic construction demands—the paper conducts thorough analysis and research. To tackle these issues, researchers have developed a data-driven management improvement framework supported by IoT technologies. This system encompasses comprehensive implementation strategies for data collection and transmission, establishment of a unified data center integrating multi-source information, and advanced data applications throughout the entire construction process. The paper elaborates on leveraging data to drive management innovation, proposing concrete implementation approaches such as real-time monitoring of worker conditions, machinery operations, and material usage patterns with proactive risk alerts. Finally, it advocates for data sharing to facilitate efficient collaboration among project stakeholders, optimize resource allocation, and ensure successful project execution and achievement of objectives. This paper conducts an in-depth and comprehensive analysis of the practical effectiveness of innovative management models, focusing on specific measures to ensure data security and effectively promote standardization.
基金funded by National Natural Science Foundation of China(52004238)China Postdoctoral Science Foundation(2019M663561).
摘要Increasing the production and utilization of shale gas is of great significance for building a clean and low-carbon energy system.Sharp decline of gas production has been widely observed in shale gas reservoirs.How to forecast shale gas production is still challenging due to complex fracture networks,dynamic fracture properties,frac hits,complicated multiphase flow,and multi-scale flow as well as data quality and uncertainty.This work develops an integrated framework for evaluating shale gas well production based on data-driven models.Firstly,a comprehensive dominated-factor system has been established,including geological,drilling,fracturing,and production factors.Data processing and visualization are required to ensure data quality and determine final data set.A shale gas production evaluation model is developed to evaluate shale gas production levels.Finally,the random forest algorithm is used to forecast shale gas production.The prediction accuracy of shale gas production level is higher than 95%based on the shale gas reservoirs in China.Forty-one wells are randomly selected to predict cumulative gas production using the optimal regression model.The proposed shale gas production evaluation frame-work overcomes too many assumptions of analytical or semi-analytical models and avoids huge computation cost and poor generalization for numerical modelling.
摘要The world’s increasing population requires the process industry to produce food,fuels,chemicals,and consumer products in a more efficient and sustainable way.Functional process materials lie at the heart of this challenge.Traditionally,new advanced materials are found empirically or through trial-and-error approaches.As theoretical methods and associated tools are being continuously improved and computer power has reached a high level,it is now efficient and popular to use computational methods to guide material selection and design.Due to the strong interaction between material selection and the operation of the process in which the material is used,it is essential to perform material and process design simultaneously.Despite this significant connection,the solution of the integrated material and process design problem is not easy because multiple models at different scales are usually required.Hybrid modeling provides a promising option to tackle such complex design problems.In hybrid modeling,the material properties,which are computationally expensive to obtain,are described by data-driven models,while the well-known process-related principles are represented by mechanistic models.This article highlights the significance of hybrid modeling in multiscale material and process design.The generic design methodology is first introduced.Six important application areas are then selected:four from the chemical engineering field and two from the energy systems engineering domain.For each selected area,state-ofthe-art work using hybrid modeling for multiscale material and process design is discussed.Concluding remarks are provided at the end,and current limitations and future opportunities are pointed out.
基金supported by the National Natural Science Foundation of China(Nos.92152301,12072282)。
摘要Aerodynamic surrogate modeling mostly relies only on integrated loads data obtained from simulation or experiment,while neglecting and wasting the valuable distributed physical information on the surface.To make full use of both integrated and distributed loads,a modeling paradigm,called the heterogeneous data-driven aerodynamic modeling,is presented.The essential concept is to incorporate the physical information of distributed loads as additional constraints within the end-to-end aerodynamic modeling.Towards heterogenous data,a novel and easily applicable physical feature embedding modeling framework is designed.This framework extracts lowdimensional physical features from pressure distribution and then effectively enhances the modeling of the integrated loads via feature embedding.The proposed framework can be coupled with multiple feature extraction methods,and the well-performed generalization capabilities over different airfoils are verified through a transonic case.Compared with traditional direct modeling,the proposed framework can reduce testing errors by almost 50%.Given the same prediction accuracy,it can save more than half of the training samples.Furthermore,the visualization analysis has revealed a significant correlation between the discovered low-dimensional physical features and the heterogeneous aerodynamic loads,which shows the interpretability and credibility of the superior performance offered by the proposed deep learning framework.
基金financially supported by the National Key Research and Development Program of China(2022YFB3706800,2020YFB1710100)the National Natural Science Foundation of China(51821001,52090042,52074183)。
摘要The complex sand-casting process combined with the interactions between process parameters makes it difficult to control the casting quality,resulting in a high scrap rate.A strategy based on a data-driven model was proposed to reduce casting defects and improve production efficiency,which includes the random forest(RF)classification model,the feature importance analysis,and the process parameters optimization with Monte Carlo simulation.The collected data includes four types of defects and corresponding process parameters were used to construct the RF model.Classification results show a recall rate above 90% for all categories.The Gini Index was used to assess the importance of the process parameters in the formation of various defects in the RF model.Finally,the classification model was applied to different production conditions for quality prediction.In the case of process parameters optimization for gas porosity defects,this model serves as an experimental process in the Monte Carlo method to estimate a better temperature distribution.The prediction model,when applied to the factory,greatly improved the efficiency of defect detection.Results show that the scrap rate decreased from 10.16% to 6.68%.
基金supported by the National Natural Science Foundation of China(Grant No.51809279)the Major National Science and Technology Program(Grant No.2016ZX05028-001-05)+1 种基金Program for Changjiang Scholars and Innovative Research Team in University(Grant No.IRT14R58)the Fundamental Research Funds for the Central Universities,that is,the Opening Fund of National Engineering Laboratory of Offshore Geophysical and Exploration Equipment(Grant No.20CX02302A).
摘要Vortex induced vibration(VIV)is a challenge in ocean engineering.Several devices including fairings have been designed to suppress VIV.However,how to optimize the design of suppression devices is still a problem to be solved.In this paper,an optimization design methodology is presented based on data-driven models and genetic algorithm(GA).Data-driven models are introduced to substitute complex physics-based equations.GA is used to rapidly search for the optimal suppression device from all possible solutions.Taking fairings as example,VIV response database for different fairings is established based on parameterized models in which model sections of fairings are controlled by several control points and Bezier curves.Then a data-driven model,which can predict the VIV response of fairings with different sections accurately and efficiently,is trained through BP neural network.Finally,a comprehensive optimization method and process is proposed based on GA and the data-driven model.The proposed method is demonstrated by its application to a case.It turns out that the proposed method can perform the optimization design of fairings effectively.VIV can be reduced obviously through the optimization design.
基金supported by the National Natural Science Foundation of China(Grant No.52306026)the Wenzhou Municipal Science and Technology Research Program(Grant No.G20220012)+2 种基金the Special Innovation Project Fund of the Institute of Wenzhou,Zhejiang University(XMGL-KJZX202205)the State Key Laboratory of Air-Conditioning Equipment and System Energy Conservation Open Project(Project No.ACSKL2021KT01)the Special Innovation Project Fund of the Institute of Wenzhou,Zhejiang University(XMGL-KJZX-202205).
摘要This study explores the effectiveness of machine learning models in predicting the air-side performance of microchannel heat exchangers.The data were generated by experimentally validated Computational Fluid Dynam-ics(CFD)simulations of air-to-water microchannel heat exchangers.A distinctive aspect of this research is the comparative analysis of four diverse machine learning algorithms:Artificial Neural Networks(ANN),Support Vector Machines(SVM),Random Forest(RF),and Gaussian Process Regression(GPR).These models are adeptly applied to predict air-side heat transfer performance with high precision,with ANN and GPR exhibiting notably superior accuracy.Additionally,this research further delves into the influence of both geometric and operational parameters—including louvered angle,fin height,fin spacing,air inlet temperature,velocity,and tube temperature—on model performance.Moreover,it innovatively incorporates dimensionless numbers such as aspect ratio,fin height-to-spacing ratio,Reynolds number,Nusselt number,normalized air inlet temperature,temperature difference,and louvered angle into the input variables.This strategic inclusion significantly refines the predictive capabilities of the models by establishing a robust analytical framework supported by the CFD-generated database.The results show the enhanced prediction accuracy achieved by integrating dimensionless numbers,highlighting the effectiveness of data-driven approaches in precisely forecasting heat exchanger performance.This advancement is pivotal for the geometric optimization of heat exchangers,illustrating the considerable potential of integrating sophisticated modeling techniques with traditional engineering metrics.