As an important subtask of fine-grained sentiment analysis,aspect term extraction(ATE)aims to identify aspect terms within user-generated comments.ATE supervised learning approaches are heavily based on the availabili...As an important subtask of fine-grained sentiment analysis,aspect term extraction(ATE)aims to identify aspect terms within user-generated comments.ATE supervised learning approaches are heavily based on the availability of annotated data with token-level labels.However,obtaining these annotations for each domain in sufficient quantity is often a costly process that limits the applicability of supervised methods.The cross-domain ATE(CDATE)problem involves transferring knowledge from a well-annotated source domain to a less-annotated target domain.We propose an approach combining the pre-trained language model with the pre-training and fine-tuning strategy for solving CDATE(named LM-PF).The first stage adapts bidirectional encoder representations from transformers(BERT)using unlabeled data from both domains to capture domain-specific characteristics,the second stage fine-tunes a downstream sequence labeling model with minimal labeled data to improve target domain performance,the third stage pre-trains on source-domain data and fine-tunes the bidirectional long short-term memory with conditional random fields(Bi-LSTM+CRF)with limited target labels.Specifically,we first adapt a pre-trained BERT model using unlabeled data from both the source and target domains to better capture domain-specific language patterns.Then,we build an ATE model initialized with this adapted BERT and pre-train it using labeled source domain data.Finally,we fine-tune the model on labeled target domain data to enhance its generalization to new domains.We conducted six CDATE tasks on three benchmark datasets,restaurants,laptops,and digital devices.The results show that our model achieves the highest average Micro-F1 score(60.09%)across all tasks and outperforms strong baselines such as generative cross-domain data augmentation by an average margin of 2.7%,confirming the effectiveness of combining domain-adaptive pretraining with task-specific fine-tuning in enhancing cross-domain generalization for ATE tasks.展开更多
Sentence classification is the process of categorizing a sentence based on the context of the sentence.Sentence categorization requires more semantic highlights than other tasks,such as dependence parsing,which requir...Sentence classification is the process of categorizing a sentence based on the context of the sentence.Sentence categorization requires more semantic highlights than other tasks,such as dependence parsing,which requires more syntactic elements.Most existing strategies focus on the general semantics of a conversation without involving the context of the sentence,recognizing the progress and comparing impacts.An ensemble pre-trained language model was taken up here to classify the conversation sentences from the conversation corpus.The conversational sentences are classified into four categories:information,question,directive,and commission.These classification label sequences are for analyzing the conversation progress and predicting the pecking order of the conversation.Ensemble of Bidirectional Encoder for Representation of Transformer(BERT),Robustly Optimized BERT pretraining Approach(RoBERTa),Generative Pre-Trained Transformer(GPT),DistilBERT and Generalized Autoregressive Pretraining for Language Understanding(XLNet)models are trained on conversation corpus with hyperparameters.Hyperparameter tuning approach is carried out for better performance on sentence classification.This Ensemble of Pre-trained Language Models with a Hyperparameter Tuning(EPLM-HT)system is trained on an annotated conversation dataset.The proposed approach outperformed compared to the base BERT,GPT,DistilBERT and XLNet transformer models.The proposed ensemble model with the fine-tuned parameters achieved an F1_score of 0.88.展开更多
To improve the accuracy and generalization of well logging curve reconstruction,this paper proposes an artificial intelligence large language model“Gaia”and conducts model evaluation experiments.By fine-tuning the p...To improve the accuracy and generalization of well logging curve reconstruction,this paper proposes an artificial intelligence large language model“Gaia”and conducts model evaluation experiments.By fine-tuning the pre-trained large language model,the Gaia significantly improved its ability in extracting sequential patterns and spatial features from well-log curves.Leveraging the adapter method for fine-tuning,this model required training only about 1/70 of its original parameters,greatly improving training efficiency.Comparative experiments,ablation experiments,and generalization experiments were designed and conducted using well-log data from 250 wells.In the comparative experiment,the Gaia model was benchmarked against cutting-edge small deep learning models and conventional large language models,demonstrating that the Gaia model reduced the mean absolute error(MAE)by at least 20%.In the ablation experiments,the synergistic effect of the Gaia model's multiple components was validated,with its MAE being at least 30%lower than that of single-component models.In the generalization experiments,the superior performance of the Gaia model in blind-well predictions was further confirmed.Compared to traditional models,the Gaia model is significantly superior in accuracy and generalization for logging curve reconstruction,fully showcasing the potential of large language models in the field of well-logging.This provides a new approach for future intelligent logging data processing.展开更多
To overcome the challenges associated with predicting gas extraction performance and mitigating the gradual decline in extraction volume,which adversely impacts gas utilization efficiency in mines,a gas extraction pur...To overcome the challenges associated with predicting gas extraction performance and mitigating the gradual decline in extraction volume,which adversely impacts gas utilization efficiency in mines,a gas extraction pure volume prediction model was developed using Support Vector Regression(SVR)and Random Forest(RF),with hyperparameters fine-tuned via the Genetic Algorithm(GA).Building upon this,an adaptive control model for gas extraction negative pressure was formulated to maximize the extracted gas volume within the pipeline network,followed by field validation experiments.Experimental results indicate that the GA-SVR model surpasses comparable models in terms of mean absolute error,root mean square error,and mean absolute percentage error.In the extraction process of bedding boreholes,the influence of negative pressure on gas extraction concentration diminishes over time,yet it remains a critical factor in determining the extracted pure volume.In contrast,throughout the entire extraction period of cross-layer boreholes,both extracted pure volume and concentration exhibit pronounced sensitivity to fluctuations in extraction negative pressure.Field experiments demonstrated that the adaptive controlmodel enhanced the average extracted gas volume by 5.08% in the experimental borehole group compared to the control group during the later extraction stage,with a more pronounced increase of 7.15% in the first 15 days.The research findings offer essential technical support for the efficient utilization and long-term sustainable development of mine gas resources.The research findings offer essential technical support for gas disaster mitigation and the sustained,efficient utilization of mine gas.展开更多
Using gas and rock samples from major petroliferous basins in the world,the helium content,composition,isotopic compositions and the U and Th contents in rocks are analyzed to clarify the helium enrichment mechanism a...Using gas and rock samples from major petroliferous basins in the world,the helium content,composition,isotopic compositions and the U and Th contents in rocks are analyzed to clarify the helium enrichment mechanism and distribution pattern and the exploration ideas for helium-rich gas reservoirs.It is believed that the formation of helium-rich gas reservoirs depends on the amount of helium supplied to the reservoir and the degree of helium dilution by natural gas,and that the reservoir-forming process can be summarized as"multi-source helium supply,main-source helium enrichment,helium-nitrogen coupling,and homogeneous symbiosis".Helium mainly comes from the radioactive decay of U and Th in rocks.All rocks contain trace amounts of U and Th,so they are effective helium sources.Especially,large-scale ancient basement dominated by granite or metamorphic rocks is the main helium source.The helium generated by the decay of U and Th in the ancient basement in a long geologic history,together with the nitrogen generated by the cracking of the inorganic nitrogenous compounds in the basement rocks,is dissolved in the water and preserved.With the tectonic uplift,the ground water is transported upward along the fracture to the gas reservoirs,with helium and nitrogen released.Thus,the reservoirs are enriched with both helium and nitrogen,which present a clear concomitant and coupling relationship.In tensional basins in eastern China,where tectonic activities are strong,a certain proportion of mantle-derived helium is mixed in the natural gas.The helium-rich gas reservoirs are mostly located in normal or low-pressure zones above ancient basement with fracture communication,which later experience substantial tectonic uplift and present relatively weak seal,low intensity of natural gas charging,and active groundwater.Helium exploration should focus on gas reservoirs with fractures connecting ancient basement,large tectonic uplift,relatively weak sealing capacity,insufficient natural gas charging intensity,and rich ancient formation water,depending on the characteristics of helium enrichment,beyond the traditional idea of searching for natural gas sweetspots and high-yield giant gas fields simultaneously.展开更多
This article elucidates the concept of large model technology,summarizes the research status of large model technology both domestically and internationally,provides an overview of the application status of large mode...This article elucidates the concept of large model technology,summarizes the research status of large model technology both domestically and internationally,provides an overview of the application status of large models in vertical industries,outlines the challenges and issues confronted in applying large models in the oil and gas sector,and offers prospects for the application of large models in the oil and gas industry.The existing large models can be briefly divided into three categories:large language models,visual large models,and multimodal large models.The application of large models in the oil and gas industry is still in its infancy.Based on open-source large language models,some oil and gas enterprises have released large language model products using methods like fine-tuning and retrieval augmented generation.Scholars have attempted to develop scenario-specific models for oil and gas operations by using visual/multimodal foundation models.A few researchers have constructed pre-trained foundation models for seismic data processing and interpretation,as well as core analysis.The application of large models in the oil and gas industry faces challenges such as current data quantity and quality being difficult to support the training of large models,high research and development costs,and poor algorithm autonomy and control.The application of large models should be guided by the needs of oil and gas business,taking the application of large models as an opportunity to improve data lifecycle management,enhance data governance capabilities,promote the construction of computing power,strengthen the construction of“artificial intelligence+energy”composite teams,and boost the autonomy and control of large model technology.展开更多
A large language model(LLM)is constructed to address the sophisticated demands of data retrieval and analysis,detailed well profiling,computation of key technical indicators,and the solutions to complex problems in re...A large language model(LLM)is constructed to address the sophisticated demands of data retrieval and analysis,detailed well profiling,computation of key technical indicators,and the solutions to complex problems in reservoir performance analysis(RPA).The LLM is constructed for RPA scenarios with incremental pre-training,fine-tuning,and functional subsystems coupling.Functional subsystem and efficient coupling methods are proposed based on named entity recognition(NER),tool invocation,and Text-to-SQL construction,all aimed at resolving pivotal challenges in developing the specific application of LLMs for RDA.This study conducted a detailed accuracy test on feature extraction models,tool classification models,data retrieval models and analysis recommendation models.The results indicate that these models have demonstrated good performance in various key aspects of reservoir dynamic analysis.The research takes some injection and production well groups in the PK3 Block of the Daqing Oilfield as an example for testing.Testing results show that our model has significant potential and practical value in assisting reservoir engineers with RDA.The research results provide a powerful support to the application of LLM in reservoir performance analysis.展开更多
The pursuit of optimal neural network architectures is foundational to the progression of Neural Architecture Search (NAS). However, the existing NAS methods suffer from the following problem using traditional search ...The pursuit of optimal neural network architectures is foundational to the progression of Neural Architecture Search (NAS). However, the existing NAS methods suffer from the following problem using traditional search strategies, i.e., when facing a large and complex search space, it is difficult to mine more effective architectures within a reasonable time, resulting in inferior search results. This research introduces the Generative Pre-trained Transformer NAS (GPT-NAS), an innovative approach designed to overcome the limitations which are inherent in traditional NAS strategies. This approach improves search efficiency and obtains better architectures by integrating GPT model into the search process. Specifically, we design a reconstruction strategy that utilizes the trained GPT to reorganize the architectures obtained from the search. In addition, to equip the GPT model with the design capabilities of neural architecture, we propose the use of the GPT model for training on a neural architecture dataset. For each architecture, the structural information of its previous layers is utilized to predict the next layer of structure, iteratively traversing the entire architecture. In this way, the GPT model can efficiently learn the key features required for neural architectures. Extensive experimental validation shows that our GPT-NAS approach beats both manually constructed neural architectures and automatically generated architectures by NAS. In addition, we validate the superiority of introducing the GPT model in several ways, and find that the accuracy of the neural architecture on the image dataset obtained from the search after introducing the GPT model is improved by up to about 9%.展开更多
Fine-tuning pre-trained language models like BERT have become an effective way in natural language processing(NLP)and yield state-of-the-art results on many downstream tasks.Recent studies on adapting BERT to new task...Fine-tuning pre-trained language models like BERT have become an effective way in natural language processing(NLP)and yield state-of-the-art results on many downstream tasks.Recent studies on adapting BERT to new tasks mainly focus on modifying the model structure,re-designing the pre-training tasks,and leveraging external data and knowledge.The fine-tuning strategy itself has yet to be fully explored.In this paper,we improve the fine-tuning of BERT with two effective mechanisms:self-ensemble and self-distillation.The self-ensemble mechanism utilizes the checkpoints from an experience pool to integrate the teacher model.In order to transfer knowledge from the teacher model to the student model efficiently,we further use knowledge distillation,which is called self-distillation because the distillation comes from the model itself through the time dimension.Experiments on the GLUE benchmark and the Text Classification benchmark show that our proposed approach can significantly improve the adaption of BERT without any external data or knowledge.We conduct exhaustive experiments to investigate the efficiency of the self-ensemble and self-distillation mechanisms,and our proposed approach achieves a new state-of-the-art result on the SNLI dataset.展开更多
With current success of large-scale pre-trained models(PTMs),how efficiently adapting PTMs to downstream tasks has attracted tremendous attention,especially for PTMs with billions of parameters.Previous work focuses o...With current success of large-scale pre-trained models(PTMs),how efficiently adapting PTMs to downstream tasks has attracted tremendous attention,especially for PTMs with billions of parameters.Previous work focuses on designing parameter-efficient tuning paradigms but needs to save and compute the gradient of the whole computational graph.In this paper,we propose y-Tuning,an efficient yet effective paradigm to adapt frozen large-scale PTMs to specific downstream tasks.y-Tuning learns dense representations for labels y defined in a given task and aligns them to fixed feature representation.Without computing the gradients of text encoder at training phrase,y-Tuning is not only parameterefficient but also training-efficient.Experimental results show that for DeBERTaxxL with 1.6 billion parameters,y-Tuning achieves performance more than 96%of full fine-tuning on GLUE Benchmark with only 2%tunable parameters and much fewer training costs.展开更多
Large models have accelerated the development of intelligent interpretation in remote sensing.Many remote sensing foundation models(RSFM)have emerged in recent years,sparking a new wave of deep learning in this field....Large models have accelerated the development of intelligent interpretation in remote sensing.Many remote sensing foundation models(RSFM)have emerged in recent years,sparking a new wave of deep learning in this field.Fine-tuning techniques serve as a bridge between remote sensing downstream tasks and advanced foundation models.As RSFMs become more powerful,fine-tuning techniques are expected to lead the next research frontier in numerous critical remote sensing applications.Advanced fine-tuning techniques can reduce the data and computational resource requirements during the downstream adaptation process.Current fine-tuning techniques for remote sensing are still in their early stages,leaving a large space for optimization and application.To elucidate the current development and future trends of remote sensing fine-tuning techniques,this survey offers a comprehensive overview of recent research.Specifically,this survey summarizes the applications and innovations of each work and categorizes recent remote sensing fine-tuning techniques into six types:adapter-based,prompt-based,reparameterization-based,hybrid methods,partial tuning,and improved tuning.展开更多
Continuous Chinese sign language recognition(CCSLR)methods have shown their strong ability to learn excellent model architectures from datasets.However,due to data insufficiency,it is difficult to complete the CCSLR t...Continuous Chinese sign language recognition(CCSLR)methods have shown their strong ability to learn excellent model architectures from datasets.However,due to data insufficiency,it is difficult to complete the CCSLR task.In this work,we focus on a simple but important solution to alleviate data insufficiency:how to refine the model architecture of a CCSLR network to improve the robustness of feature processing by using some better-quality non-Chinese sign language datasets.To this end,a simple empirical study wasfirst conducted to verify the feasibility of knowledge transfer in the CCSLR task.Surprisingly,just by pre-training of our recognition model on a foreign sign language dataset,we can refine the model architecture and improve its robustness significantly.To make it more practical,the key issue of how tofine-tune the existing feature processing models for effective guidance should be carefully investigated.Then,we propose a novel scheme forfine-tuning of pre-trained models named FTP,which updates the spatial feature extractor initialized by a pre-trained backbone and freezes the temporal feature extractor implemented by a better shareable transformer encoder.Compared with the baseline method,our FTP method can achieve significant performance improvement on the public dataset USTC-CCSL.展开更多
基金supported by Program for the Innovative Talents of Higher Learning Institutions of Shanxi(Grant No.2024Q018)Shanxi Provincial Philosophy and Social Sciences Planning Project(Grant No.2024YB095)+3 种基金Fundamental Research Program of Shanxi Province(Grant No.202303021211139)Shanxi Province Science and Technology Cooperation and Exchange Special Project(Grant No.202404041101003)Shanxi Province Science and Technology Strategic Research Project(Grant No.202404030401078)Humanity and Social Science Youth Foundation of Ministry of Education in China(Grant No.24YCZH246).
摘要As an important subtask of fine-grained sentiment analysis,aspect term extraction(ATE)aims to identify aspect terms within user-generated comments.ATE supervised learning approaches are heavily based on the availability of annotated data with token-level labels.However,obtaining these annotations for each domain in sufficient quantity is often a costly process that limits the applicability of supervised methods.The cross-domain ATE(CDATE)problem involves transferring knowledge from a well-annotated source domain to a less-annotated target domain.We propose an approach combining the pre-trained language model with the pre-training and fine-tuning strategy for solving CDATE(named LM-PF).The first stage adapts bidirectional encoder representations from transformers(BERT)using unlabeled data from both domains to capture domain-specific characteristics,the second stage fine-tunes a downstream sequence labeling model with minimal labeled data to improve target domain performance,the third stage pre-trains on source-domain data and fine-tunes the bidirectional long short-term memory with conditional random fields(Bi-LSTM+CRF)with limited target labels.Specifically,we first adapt a pre-trained BERT model using unlabeled data from both the source and target domains to better capture domain-specific language patterns.Then,we build an ATE model initialized with this adapted BERT and pre-train it using labeled source domain data.Finally,we fine-tune the model on labeled target domain data to enhance its generalization to new domains.We conducted six CDATE tasks on three benchmark datasets,restaurants,laptops,and digital devices.The results show that our model achieves the highest average Micro-F1 score(60.09%)across all tasks and outperforms strong baselines such as generative cross-domain data augmentation by an average margin of 2.7%,confirming the effectiveness of combining domain-adaptive pretraining with task-specific fine-tuning in enhancing cross-domain generalization for ATE tasks.
摘要Sentence classification is the process of categorizing a sentence based on the context of the sentence.Sentence categorization requires more semantic highlights than other tasks,such as dependence parsing,which requires more syntactic elements.Most existing strategies focus on the general semantics of a conversation without involving the context of the sentence,recognizing the progress and comparing impacts.An ensemble pre-trained language model was taken up here to classify the conversation sentences from the conversation corpus.The conversational sentences are classified into four categories:information,question,directive,and commission.These classification label sequences are for analyzing the conversation progress and predicting the pecking order of the conversation.Ensemble of Bidirectional Encoder for Representation of Transformer(BERT),Robustly Optimized BERT pretraining Approach(RoBERTa),Generative Pre-Trained Transformer(GPT),DistilBERT and Generalized Autoregressive Pretraining for Language Understanding(XLNet)models are trained on conversation corpus with hyperparameters.Hyperparameter tuning approach is carried out for better performance on sentence classification.This Ensemble of Pre-trained Language Models with a Hyperparameter Tuning(EPLM-HT)system is trained on an annotated conversation dataset.The proposed approach outperformed compared to the base BERT,GPT,DistilBERT and XLNet transformer models.The proposed ensemble model with the fine-tuned parameters achieved an F1_score of 0.88.
基金Supported by the National Natural Science Foundation of China(52288101)National Key R&D Program of China(2024YFF1500600)。
摘要To improve the accuracy and generalization of well logging curve reconstruction,this paper proposes an artificial intelligence large language model“Gaia”and conducts model evaluation experiments.By fine-tuning the pre-trained large language model,the Gaia significantly improved its ability in extracting sequential patterns and spatial features from well-log curves.Leveraging the adapter method for fine-tuning,this model required training only about 1/70 of its original parameters,greatly improving training efficiency.Comparative experiments,ablation experiments,and generalization experiments were designed and conducted using well-log data from 250 wells.In the comparative experiment,the Gaia model was benchmarked against cutting-edge small deep learning models and conventional large language models,demonstrating that the Gaia model reduced the mean absolute error(MAE)by at least 20%.In the ablation experiments,the synergistic effect of the Gaia model's multiple components was validated,with its MAE being at least 30%lower than that of single-component models.In the generalization experiments,the superior performance of the Gaia model in blind-well predictions was further confirmed.Compared to traditional models,the Gaia model is significantly superior in accuracy and generalization for logging curve reconstruction,fully showcasing the potential of large language models in the field of well-logging.This provides a new approach for future intelligent logging data processing.
基金funded by the National Key Research and Development Program of China,grant number:2023YFF0615404.
摘要To overcome the challenges associated with predicting gas extraction performance and mitigating the gradual decline in extraction volume,which adversely impacts gas utilization efficiency in mines,a gas extraction pure volume prediction model was developed using Support Vector Regression(SVR)and Random Forest(RF),with hyperparameters fine-tuned via the Genetic Algorithm(GA).Building upon this,an adaptive control model for gas extraction negative pressure was formulated to maximize the extracted gas volume within the pipeline network,followed by field validation experiments.Experimental results indicate that the GA-SVR model surpasses comparable models in terms of mean absolute error,root mean square error,and mean absolute percentage error.In the extraction process of bedding boreholes,the influence of negative pressure on gas extraction concentration diminishes over time,yet it remains a critical factor in determining the extracted pure volume.In contrast,throughout the entire extraction period of cross-layer boreholes,both extracted pure volume and concentration exhibit pronounced sensitivity to fluctuations in extraction negative pressure.Field experiments demonstrated that the adaptive controlmodel enhanced the average extracted gas volume by 5.08% in the experimental borehole group compared to the control group during the later extraction stage,with a more pronounced increase of 7.15% in the first 15 days.The research findings offer essential technical support for the efficient utilization and long-term sustainable development of mine gas resources.The research findings offer essential technical support for gas disaster mitigation and the sustained,efficient utilization of mine gas.
基金Supported by the National Natural Science Foundation of China(42141022,42272189)Project of Ministry of Natural Resources of China(QGYQZYPJ2022-1)CNPC Core Project(2021ZG12)。
摘要Using gas and rock samples from major petroliferous basins in the world,the helium content,composition,isotopic compositions and the U and Th contents in rocks are analyzed to clarify the helium enrichment mechanism and distribution pattern and the exploration ideas for helium-rich gas reservoirs.It is believed that the formation of helium-rich gas reservoirs depends on the amount of helium supplied to the reservoir and the degree of helium dilution by natural gas,and that the reservoir-forming process can be summarized as"multi-source helium supply,main-source helium enrichment,helium-nitrogen coupling,and homogeneous symbiosis".Helium mainly comes from the radioactive decay of U and Th in rocks.All rocks contain trace amounts of U and Th,so they are effective helium sources.Especially,large-scale ancient basement dominated by granite or metamorphic rocks is the main helium source.The helium generated by the decay of U and Th in the ancient basement in a long geologic history,together with the nitrogen generated by the cracking of the inorganic nitrogenous compounds in the basement rocks,is dissolved in the water and preserved.With the tectonic uplift,the ground water is transported upward along the fracture to the gas reservoirs,with helium and nitrogen released.Thus,the reservoirs are enriched with both helium and nitrogen,which present a clear concomitant and coupling relationship.In tensional basins in eastern China,where tectonic activities are strong,a certain proportion of mantle-derived helium is mixed in the natural gas.The helium-rich gas reservoirs are mostly located in normal or low-pressure zones above ancient basement with fracture communication,which later experience substantial tectonic uplift and present relatively weak seal,low intensity of natural gas charging,and active groundwater.Helium exploration should focus on gas reservoirs with fractures connecting ancient basement,large tectonic uplift,relatively weak sealing capacity,insufficient natural gas charging intensity,and rich ancient formation water,depending on the characteristics of helium enrichment,beyond the traditional idea of searching for natural gas sweetspots and high-yield giant gas fields simultaneously.
基金Supported by the National Natural Science Foundation of China(72088101,42372175)PetroChina Science and Technology Innovation Fund Program(2021DQ02-0904)。
摘要This article elucidates the concept of large model technology,summarizes the research status of large model technology both domestically and internationally,provides an overview of the application status of large models in vertical industries,outlines the challenges and issues confronted in applying large models in the oil and gas sector,and offers prospects for the application of large models in the oil and gas industry.The existing large models can be briefly divided into three categories:large language models,visual large models,and multimodal large models.The application of large models in the oil and gas industry is still in its infancy.Based on open-source large language models,some oil and gas enterprises have released large language model products using methods like fine-tuning and retrieval augmented generation.Scholars have attempted to develop scenario-specific models for oil and gas operations by using visual/multimodal foundation models.A few researchers have constructed pre-trained foundation models for seismic data processing and interpretation,as well as core analysis.The application of large models in the oil and gas industry faces challenges such as current data quantity and quality being difficult to support the training of large models,high research and development costs,and poor algorithm autonomy and control.The application of large models should be guided by the needs of oil and gas business,taking the application of large models as an opportunity to improve data lifecycle management,enhance data governance capabilities,promote the construction of computing power,strengthen the construction of“artificial intelligence+energy”composite teams,and boost the autonomy and control of large model technology.
基金Supported by the National Talent Fund of the Ministry of Science and Technology of China(20230240011)China University of Geosciences(Wuhan)Research Fund(162301192687)。
摘要A large language model(LLM)is constructed to address the sophisticated demands of data retrieval and analysis,detailed well profiling,computation of key technical indicators,and the solutions to complex problems in reservoir performance analysis(RPA).The LLM is constructed for RPA scenarios with incremental pre-training,fine-tuning,and functional subsystems coupling.Functional subsystem and efficient coupling methods are proposed based on named entity recognition(NER),tool invocation,and Text-to-SQL construction,all aimed at resolving pivotal challenges in developing the specific application of LLMs for RDA.This study conducted a detailed accuracy test on feature extraction models,tool classification models,data retrieval models and analysis recommendation models.The results indicate that these models have demonstrated good performance in various key aspects of reservoir dynamic analysis.The research takes some injection and production well groups in the PK3 Block of the Daqing Oilfield as an example for testing.Testing results show that our model has significant potential and practical value in assisting reservoir engineers with RDA.The research results provide a powerful support to the application of LLM in reservoir performance analysis.
基金supported by the National Nature Science Foundation of China(No.62106161)the Fundamental Research Funds for the Central Universities(No.1082204112364)+4 种基金the Sichuan University Luzhou Municipal Government Strategic Cooperation Project(No.2022CDLZ-8)the Key R&D Program of Sichuan Province(Nos.2022YFN0017 and 2023YFG0019)the Natural Science Foundation of Sichuan(No.2023NSFSC0474)the Tianfiu Yongxing Laboratory Organized Research Project Funding(No.2023CXXM14)the Digital Media Art,Key Laboratory of Sichuan Province,Sichuan Conservatory of Music(No.22DMAKL04).
摘要The pursuit of optimal neural network architectures is foundational to the progression of Neural Architecture Search (NAS). However, the existing NAS methods suffer from the following problem using traditional search strategies, i.e., when facing a large and complex search space, it is difficult to mine more effective architectures within a reasonable time, resulting in inferior search results. This research introduces the Generative Pre-trained Transformer NAS (GPT-NAS), an innovative approach designed to overcome the limitations which are inherent in traditional NAS strategies. This approach improves search efficiency and obtains better architectures by integrating GPT model into the search process. Specifically, we design a reconstruction strategy that utilizes the trained GPT to reorganize the architectures obtained from the search. In addition, to equip the GPT model with the design capabilities of neural architecture, we propose the use of the GPT model for training on a neural architecture dataset. For each architecture, the structural information of its previous layers is utilized to predict the next layer of structure, iteratively traversing the entire architecture. In this way, the GPT model can efficiently learn the key features required for neural architectures. Extensive experimental validation shows that our GPT-NAS approach beats both manually constructed neural architectures and automatically generated architectures by NAS. In addition, we validate the superiority of introducing the GPT model in several ways, and find that the accuracy of the neural architecture on the image dataset obtained from the search after introducing the GPT model is improved by up to about 9%.
基金supported by the National Key Research and Development Program of China under Grant No.2020AAA0106700the National Natural Science Foundation of China under Grant No.62022027.
摘要Fine-tuning pre-trained language models like BERT have become an effective way in natural language processing(NLP)and yield state-of-the-art results on many downstream tasks.Recent studies on adapting BERT to new tasks mainly focus on modifying the model structure,re-designing the pre-training tasks,and leveraging external data and knowledge.The fine-tuning strategy itself has yet to be fully explored.In this paper,we improve the fine-tuning of BERT with two effective mechanisms:self-ensemble and self-distillation.The self-ensemble mechanism utilizes the checkpoints from an experience pool to integrate the teacher model.In order to transfer knowledge from the teacher model to the student model efficiently,we further use knowledge distillation,which is called self-distillation because the distillation comes from the model itself through the time dimension.Experiments on the GLUE benchmark and the Text Classification benchmark show that our proposed approach can significantly improve the adaption of BERT without any external data or knowledge.We conduct exhaustive experiments to investigate the efficiency of the self-ensemble and self-distillation mechanisms,and our proposed approach achieves a new state-of-the-art result on the SNLI dataset.
基金National Key R&D Program of China(No.2020AAA0108702)National Natural Science Foundation of China(Grant No.62022027).
摘要With current success of large-scale pre-trained models(PTMs),how efficiently adapting PTMs to downstream tasks has attracted tremendous attention,especially for PTMs with billions of parameters.Previous work focuses on designing parameter-efficient tuning paradigms but needs to save and compute the gradient of the whole computational graph.In this paper,we propose y-Tuning,an efficient yet effective paradigm to adapt frozen large-scale PTMs to specific downstream tasks.y-Tuning learns dense representations for labels y defined in a given task and aligns them to fixed feature representation.Without computing the gradients of text encoder at training phrase,y-Tuning is not only parameterefficient but also training-efficient.Experimental results show that for DeBERTaxxL with 1.6 billion parameters,y-Tuning achieves performance more than 96%of full fine-tuning on GLUE Benchmark with only 2%tunable parameters and much fewer training costs.
基金supported by the National Natural Science Foundation of China(62495061,62495064,and 62476143)the Tsinghua-Tencent Joint Laboratory for Internet Innovation Technology,and the Shuimu Tsinghua Scholar Program.
摘要Large models have accelerated the development of intelligent interpretation in remote sensing.Many remote sensing foundation models(RSFM)have emerged in recent years,sparking a new wave of deep learning in this field.Fine-tuning techniques serve as a bridge between remote sensing downstream tasks and advanced foundation models.As RSFMs become more powerful,fine-tuning techniques are expected to lead the next research frontier in numerous critical remote sensing applications.Advanced fine-tuning techniques can reduce the data and computational resource requirements during the downstream adaptation process.Current fine-tuning techniques for remote sensing are still in their early stages,leaving a large space for optimization and application.To elucidate the current development and future trends of remote sensing fine-tuning techniques,this survey offers a comprehensive overview of recent research.Specifically,this survey summarizes the applications and innovations of each work and categorizes recent remote sensing fine-tuning techniques into six types:adapter-based,prompt-based,reparameterization-based,hybrid methods,partial tuning,and improved tuning.
基金supported by the National Natural Science Foundation of China(Nos.62376197,62020106004,92048301,and 62202332).
摘要Continuous Chinese sign language recognition(CCSLR)methods have shown their strong ability to learn excellent model architectures from datasets.However,due to data insufficiency,it is difficult to complete the CCSLR task.In this work,we focus on a simple but important solution to alleviate data insufficiency:how to refine the model architecture of a CCSLR network to improve the robustness of feature processing by using some better-quality non-Chinese sign language datasets.To this end,a simple empirical study wasfirst conducted to verify the feasibility of knowledge transfer in the CCSLR task.Surprisingly,just by pre-training of our recognition model on a foreign sign language dataset,we can refine the model architecture and improve its robustness significantly.To make it more practical,the key issue of how tofine-tune the existing feature processing models for effective guidance should be carefully investigated.Then,we propose a novel scheme forfine-tuning of pre-trained models named FTP,which updates the spatial feature extractor initialized by a pre-trained backbone and freezes the temporal feature extractor implemented by a better shareable transformer encoder.Compared with the baseline method,our FTP method can achieve significant performance improvement on the public dataset USTC-CCSL.