This study focuses on meeting the challenges of big data visualization by using of data reduction methods based the feature selection methods.To reduce the volume of big data and minimize model training time(Tt)while ...This study focuses on meeting the challenges of big data visualization by using of data reduction methods based the feature selection methods.To reduce the volume of big data and minimize model training time(Tt)while maintaining data quality.We contributed to meeting the challenges of big data visualization using the embedded method based“Select from model(SFM)”method by using“Random forest Importance algorithm(RFI)”and comparing it with the filter method by using“Select percentile(SP)”method based chi square“Chi2”tool for selecting the most important features,which are then fed into a classification process using the logistic regression(LR)algorithm and the k-nearest neighbor(KNN)algorithm.Thus,the classification accuracy(AC)performance of LRis also compared to theKNN approach in python on eight data sets to see which method produces the best rating when feature selection methods are applied.Consequently,the study concluded that the feature selection methods have a significant impact on the analysis and visualization of the data after removing the repetitive data and the data that do not affect the goal.After making several comparisons,the study suggests(SFMLR)using SFM based on RFI algorithm for feature selection,with LR algorithm for data classify.The proposal proved its efficacy by comparing its results with recent literature.展开更多
The study of marine data visualization is of great value. Marine data, due to its large scale, random variation and multiresolution in nature, are hard to be visualized and analyzed. Nowadays, constructing an ocean mo...The study of marine data visualization is of great value. Marine data, due to its large scale, random variation and multiresolution in nature, are hard to be visualized and analyzed. Nowadays, constructing an ocean model and visualizing model results have become some of the most important research topics of ‘Digital Ocean'. In this paper, a spherical ray casting method is developed to improve the traditional ray-casting algorithm and to make efficient use of GPUs. Aiming at the ocean current data, a 3D view-dependent line integral convolution method is used, in which the spatial frequency is adapted according to the distance from a camera. The study is based on a 3D virtual reality and visualization engine, namely the VV-Ocean. Some interactive operations are also provided to highlight the interesting structures and the characteristics of volumetric data. Finally, the marine data gathered in the East China Sea are displayed and analyzed. The results show that the method meets the requirements of real-time and interactive rendering.展开更多
Data visualization blends art and science to convey stories from data via graphical representations.Considering different problems,applications,requirements,and design goals,it is challenging to combine these two comp...Data visualization blends art and science to convey stories from data via graphical representations.Considering different problems,applications,requirements,and design goals,it is challenging to combine these two components at their full force.While the art component involves creating visually appealing and easily interpreted graphics for users,the science component requires accurate representations of a large amount of input data.With a lack of the science component,visualization cannot serve its role of creating correct representations of the actual data,thus leading to wrong perception,interpretation,and decision.It might be even worse if incorrect visual representations were intentionally produced to deceive the viewers.To address common pitfalls in graphical representations,this paper focuses on identifying and understanding the root causes of misinformation in graphical representations.We reviewed the misleading data visualization examples in the scientific publications collected from indexing databases and then projected them onto the fundamental units of visual communication such as color,shape,size,and spatial orientation.Moreover,a text mining technique was applied to extract practical insights from common visualization pitfalls.Cochran’s Q test and McNemar’s test were conducted to examine if there is any difference in the proportions of common errors among color,shape,size,and spatial orientation.The findings showed that the pie chart is the most misused graphical representation,and size is the most critical issue.It was also observed that there were statistically significant differences in the proportion of errors among color,shape,size,and spatial orientation.展开更多
The Growth Value Model(GVM)proposed theoretical closed form formulas consist-ing of Return on Equity(ROE)and the Price-to-Book value ratio(P/B)for fair stock prices and expected rates of return.Although regression ana...The Growth Value Model(GVM)proposed theoretical closed form formulas consist-ing of Return on Equity(ROE)and the Price-to-Book value ratio(P/B)for fair stock prices and expected rates of return.Although regression analysis can be employed to verify these theoretical closed form formulas,they cannot be explored by classical quintile or decile sorting approaches with intuition due to the essence of multi-factors and dynamical processes.This article uses visualization techniques to help intuitively explore GVM.The discerning findings and contributions of this paper is that we put forward the concept of the smart frontier,which can be regarded as the reasonable lower limit of P/B at a specific ROE by exploring fair P/B with ROE-P/B 2D dynamical process visualization.The coefficients in the formula can be determined by the quantile regression analysis with market data.The moving paths of the ROE and P/B in the cur-rent quarter and the subsequent quarters show that the portfolios at the lower right of the curve approaches this curve and stagnates here after the portfolios are formed.Furthermore,exploring expected rates of return with ROE-P/B-Return 3D dynamical process visualization,the results show that the data outside of the lower right edge of the“smart frontier”has positive quarterly return rates not only in the t+1 quarter but also in the t+2 quarter.The farther away the data in the t quarter is from the“smart frontier”,the larger the return rates in the t+1 and t+2 quarter.展开更多
Exploration of artworks is enjoyable but often time consuming.For example,it is not always easy to discover the favorite types of unknown painting works.It is not also always easy to explore unpopular painting works w...Exploration of artworks is enjoyable but often time consuming.For example,it is not always easy to discover the favorite types of unknown painting works.It is not also always easy to explore unpopular painting works which looks similar to painting works created by famous artists.This paper presents a painting image browser which assists the explorative discovery of user-interested painting works.The presented browser applies a new multidimensional data visualization technique that highlights particular ranges of particular numeric values based on association rules to suggest cues to find favorite painting images.This study assumes a large number of painting images are provided where categorical information(e.g.,names of artists,created year)is assigned to the images.The presented system firstly calculates the feature values of the images as a preprocessing step.Then the browser visualizes the multidimensional feature values as a heatmap and highlights association rules discovered from the relationships between the feature values and categorical information.This mechanism enables users to explore favorite painting images or painting images that look similar to famous painting works.Our case study and user evaluation demonstrates the effectiveness of the presented image browser.展开更多
Background:Reproductive,maternal,newborn,child health,and nutrition(RMNCH&N)data is an indispensable tool for program and policy decisions in low-and middle-income countries.However,being equipped with evidence do...Background:Reproductive,maternal,newborn,child health,and nutrition(RMNCH&N)data is an indispensable tool for program and policy decisions in low-and middle-income countries.However,being equipped with evidence doesn’t necessarily translate to program and policy changes.This study aimed to characterize data visualization interpretation capacity and preferences among RMNCH&N Tanzanian program implementers and policymakers(“decision-makers”)to design more effective approaches towards promoting evidence-based RMNCH&N decisions in Tanzania.Methods:We conducted 25 semi-structured interviews in Kiswahili with junior,mid-level,and senior RMNCH&N decision-makers working in Tanzanian government institutions.We used snowball sampling to recruit participants with different rank and roles in RMNCH&N decision-making.Using semi-structured interviews,we probed participants on their statistical skills and data use,and asked participants to identify key messages and rank prepared RMNCH&N visualizations.We used a grounded theory approach to organize themes and identify findings.Results:The findings suggest that data literacy and statistical skills among RMNCH&N decision-makers in Tanzania varies.Most participants demonstrated awareness of many critical factors that should influence a visualization choice—audience,key message,simplicity—but assessments of data interpretation and preferences suggest that there may be weak knowledge of basic statistics.A majority of decision-makers have not had any statistical training since attending university.There appeared to be some discomfort with interpreting and using visualizations that are not bar charts,pie charts,and maps.Conclusions:Decision-makers must be able to understand and interpret RMNCH&N data they receive to be empowered to act.Addressing inadequate data literacy and presentation skills among decision-makers is vital to bridging gaps between evidence and policymaking.It would be beneficial to host basic data literacy and visualization training for RMNCH&N decision-makers at all levels in Tanzania,and to expand skills on developing key messages from visualizations.展开更多
Many countries are paying more and more attention to the protection of water resources at present,and how to protect water resources has received extensive attention from society.Water quality monitoring is the key wo...Many countries are paying more and more attention to the protection of water resources at present,and how to protect water resources has received extensive attention from society.Water quality monitoring is the key work to water resources protection.How to efficiently collect and analyze water quality monitoring data is an important aspect of water resources protection.In this paper,python programming tools and regular expressions were used to design a web crawler for the acquisition of water quality monitoring data from Global Freshwater Quality Database(GEMStat)sites,and the multi-thread parallelism was added to improve the efficiency in the process of downloading and parsing.In order to analyze and process the crawled water quality data,Pandas and Pyecharts are used to visualize the water quality data to show the intrinsic correlation and spatiotemporal relationship of the data.展开更多
The advent of the big data era has made data visualization a crucial tool for enhancing the efficiency and insights of data analysis. This theoretical research delves into the current applications and potential future...The advent of the big data era has made data visualization a crucial tool for enhancing the efficiency and insights of data analysis. This theoretical research delves into the current applications and potential future trends of data visualization in big data analysis. The article first systematically reviews the theoretical foundations and technological evolution of data visualization, and thoroughly analyzes the challenges faced by visualization in the big data environment, such as massive data processing, real-time visualization requirements, and multi-dimensional data display. Through extensive literature research, it explores innovative application cases and theoretical models of data visualization in multiple fields including business intelligence, scientific research, and public decision-making. The study reveals that interactive visualization, real-time visualization, and immersive visualization technologies may become the main directions for future development and analyzes the potential of these technologies in enhancing user experience and data comprehension. The paper also delves into the theoretical potential of artificial intelligence technology in enhancing data visualization capabilities, such as automated chart generation, intelligent recommendation of visualization schemes, and adaptive visualization interfaces. The research also focuses on the role of data visualization in promoting interdisciplinary collaboration and data democratization. Finally, the paper proposes theoretical suggestions for promoting data visualization technology innovation and application popularization, including strengthening visualization literacy education, developing standardized visualization frameworks, and promoting open-source sharing of visualization tools. This study provides a comprehensive theoretical perspective for understanding the importance of data visualization in the big data era and its future development directions.展开更多
This study aims to explore the application of Bayesian analysis based on neural networks and deep learning in data visualization.The research background is that with the increasing amount and complexity of data,tradit...This study aims to explore the application of Bayesian analysis based on neural networks and deep learning in data visualization.The research background is that with the increasing amount and complexity of data,traditional data analysis methods have been unable to meet the needs.Research methods include building neural networks and deep learning models,optimizing and improving them through Bayesian analysis,and applying them to the visualization of large-scale data sets.The results show that the neural network combined with Bayesian analysis and deep learning method can effectively improve the accuracy and efficiency of data visualization,and enhance the intuitiveness and depth of data interpretation.The significance of the research is that it provides a new solution for data visualization in the big data environment and helps to further promote the development and application of data science.展开更多
Financial literacy empowers individuals to make informed and effective financial decisions,improving their overall financial well-being and security.However,for many people understanding financial concepts can be daun...Financial literacy empowers individuals to make informed and effective financial decisions,improving their overall financial well-being and security.However,for many people understanding financial concepts can be daunting and only half of US adults are considered financially literate.Data visualization simplifies these concepts,making them accessible and engaging for learners of all ages.This systematic review analyzes 37 research papers exploring the use of data visualization and visual analytics in financial education and literacy enhancement.We classify these studies into five key areas:(1)the evolution of visualization use across time and space,(2)motivations for using visualization tools,(3)the financial topics addressed and instructional approaches used,(4)the types of tools and technologies applied,and(5)how the effectiveness of teaching interventions was evaluated.Furthermore,we identify research gaps and highlight opportunities for advancing financial literacy.Our findings offer practical insights for educators and professionals to effectively utilize or design visual tools for financial literacy.展开更多
Translation is a crucial step in gene expression.Over the past decade,the development and application of ribosome profiling(Ribo-seq)have significantly advanced our understanding of translational regulation in vivo.Ho...Translation is a crucial step in gene expression.Over the past decade,the development and application of ribosome profiling(Ribo-seq)have significantly advanced our understanding of translational regulation in vivo.However,the analysis and visualization of Ribo-seq data remain challenging.Despite the availability of various analytical pipelines,improvements in comprehensiveness,accuracy,and user-friendliness are still necessary.In this study,we develop RiboParser/RiboShiny,a robust framework for analyzing and visualizing Ribo-seq data.Building on published methods,we optimize ribosome structure-based and start/stopbased models to improve the accuracy and stability of P-site detection,even in species with a high proportion of leaderless transcripts.Leveraging these improvements,RiboParser offers comprehensive analyses,including quality control,gene-level analysis,codon-level analysis,and the analysis of Ribo-seq variants.Meanwhile,RiboShiny provides a user-friendly and adaptable platform for data visualization,facilitating deeper insights into the translational landscape.Furthermore,the integration of standardized genome annotation renders our platform universally applicable to various organisms with sequenced genomes.This framework has the potential to significantly improve the precision and efficiency of Ribo-seq data interpretation,thereby deepening our understanding of translational regulation.展开更多
In the digital age,where the amount of data is growing exponentially,traditional big data visualization methods often fall short,especially when it comes to processing high-dimensional data or matching visualization r...In the digital age,where the amount of data is growing exponentially,traditional big data visualization methods often fall short,especially when it comes to processing high-dimensional data or matching visualization results to user needs.This highlights the need for better integration between Artificial Intelligence(AI)and big data visualization.AI has brought about significant changes in this field.This paper reviews research from the past five years to explore the integration of AI with big data visualization.It looks at the core technologies,common use cases,and challenges of this combination.The review finds that AI can address many issues in traditional visualization techniques,such as handling complex data and improving interaction and evaluation.However,challenges like model interpretability,data privacy concerns,and inconsistent standards remain.In the future,this integration is expected to move toward more transparent models,stronger privacy protection,and standardized systems.This will help unlock the full value of data in real-world applications.展开更多
With the further advancement of China’s strategy of building a maritime power, marine surveying and mapping, as the core fundamental work for marine development and protection, has witnessed an exponential growth in ...With the further advancement of China’s strategy of building a maritime power, marine surveying and mapping, as the core fundamental work for marine development and protection, has witnessed an exponential growth in data volume. The traditional file-based management architecture and static visualization mode can hardly meet the strategic demands of modern marine informatization construction. With its outstanding capabilities in spatial data processing, analysis and representation, GIS technology provides systematic technical support for the refined management and efficient visualization of marine surveying and mapping data. Combined with the inherent characteristics and management pain points of marine surveying and mapping data, this paper systematically discusses the application mechanisms of GIS technology in core technical links, including multi-source heterogeneous data integration and standardization, spatiotemporal database construction and efficient retrieval, data quality control and dynamic update, as well as 2D-3D integrated visualization, dynamic spatiotemporal process simulation, and interactive decision support. It aims to provide theoretical references and practical experience for improving the resource utilization efficiency and professional service capacity of marine surveying and mapping data.展开更多
Earthquakes are highly destructive spatio-temporal phenomena whose analysis is essential for disaster preparedness and risk mitigation.Modern seismological research produces vast volumes of heterogeneous data from sei...Earthquakes are highly destructive spatio-temporal phenomena whose analysis is essential for disaster preparedness and risk mitigation.Modern seismological research produces vast volumes of heterogeneous data from seismic networks,satellite observations,and geospatial repositories,creating the need for scalable infrastructures capable of integrating and analyzing such data to support intelligent decision-making.Data warehousing technologies provide a robust foundation for this purpose;however,existing earthquake-oriented data warehouses remain limited,often relying on simplified schemas,domain-specific analytics,or cataloguing efforts.This paper presents the design and implementation of a spatio-temporal data warehouse for seismic activity.The framework integrates spatial and temporal dimensions in a unified schema and introduces a novel array-based approach for managing many-to-many relationships between facts and dimensions without intermediate bridge tables.A comparative evaluation against a conventional bridge-table schema demonstrates that the array-based design improves fact-centric query performance,while the bridge-table schema remains advantageous for dimension-centric queries.To reconcile these trade-offs,a hybrid schema is proposed that retains both representations,ensuring balanced efficiency across heterogeneous workloads.The proposed framework demonstrates how spatio-temporal data warehousing can address schema complexity,improve query performance,and support multidimensional visualization.In doing so,it provides a foundation for integrating seismic analysis into broader big data-driven intelligent decision systems for disaster resilience,risk mitigation,and emergency management.展开更多
The widespread use of numerical simulations in different scientific domains provides a variety of research opportunities.They often output a great deal of spatio-temporal simulation data,which are traditionally charac...The widespread use of numerical simulations in different scientific domains provides a variety of research opportunities.They often output a great deal of spatio-temporal simulation data,which are traditionally characterized as single-run,multi-run,multi-variate,multi-modal and multi-dimensional.From the perspective of data exploration and analysis,we noticed that many works focusing on spatiotemporal simulation data often share similar exploration techniques,for example,the exploration schemes designed in simulation space,parameter space,feature space and combinations of them.However,it lacks a survey to have a systematic overview of the essential commonalities shared by those works.In this survey,we take a novel multi-space perspective to categorize the state-ofthe-art works into three major categories.Specifically,the works are characterized as using similar techniques such as visual designs in simulation space(e.g,visual mapping,boxplot-based visual summarization,etc.),parameter space analysis(e.g,visual steering,parameter space projection,etc.)and data processing in feature space(e.g,feature definition and extraction,sampling,reduction and clustering of simulation data,etc.).展开更多
A visualization tool was developed through a web browser based on Java applets embedded into HTML pages, in order to provide a world access to the EAST experimental data. It can display data from various trees in diff...A visualization tool was developed through a web browser based on Java applets embedded into HTML pages, in order to provide a world access to the EAST experimental data. It can display data from various trees in different servers in a single panel. With WebScope, it is easier to make a comparison between different data sources and perform a simple calculation over different data sources.展开更多
Cyber security has been thrust into the limelight in the modern technological era because of an array of attacks often bypassing tmtrained intrusion detection systems (IDSs). Therefore, greater attention has been di...Cyber security has been thrust into the limelight in the modern technological era because of an array of attacks often bypassing tmtrained intrusion detection systems (IDSs). Therefore, greater attention has been directed on being able deciphering better methods for identifying attack types to train IDSs more effectively. Keycyber-attack insights exist in big data; however, an efficient approach is required to determine strong attack types to train IDSs to become more effective in key areas. Despite the rising growth in IDS research, there is a lack of studies involving big data visualization, which is key. The KDD99 data set has served as a strong benchmark since 1999; therefore, we utilized this data set in our experiment. In this study, we utilized hash algorithm, a weight table, and sampling method to deal with the inherent problems caused by analyzing big data; volume, variety, and velocity. By utilizing a visualization algorithm, we were able to gain insights into the KDD99 data set with a clear iden- tification of "normal" clusters and described distinct clusters of effective attacks.展开更多
Fine art authentication plays a significant role in protecting cultural heritage and ensuring the integrity of artworks.Traditional authentication methods require professionals to collect many reference materials and ...Fine art authentication plays a significant role in protecting cultural heritage and ensuring the integrity of artworks.Traditional authentication methods require professionals to collect many reference materials and conduct detailed analyses.To ease the difficulty,we collaborate with domain experts to develop a GPT-based agent,namely ArtEyer,that offers accurate attributions,determines the origin and authorship,and executes visual analytics.Despite the convenience of the conversational user interface,novice users may still face challenges due to the hallucination issue and the steep learning curve associated with prompting.To face these obstacles,we propose a novel solution that places interactive data visualizations into the conversations.We create contextual visualizations from an external domain-dependent database to ensure data trustworthiness and allow users to provide precise instructions to the agent by interacting directly with these visualizations,thus overcoming the vagueness inherent in natural language-based prompting.We evaluate ArtEyer through an in-lab user study and demonstrate its usage with a real-world case.展开更多
In recent years,with the wide application of image data visual extraction technology in the field of industrial engineering,the development of industrial economy has reached a new situation.To explore the interaction ...In recent years,with the wide application of image data visual extraction technology in the field of industrial engineering,the development of industrial economy has reached a new situation.To explore the interaction between the pellet microstructure and compressive strength,firstly,the pellet microstructure needed for the experiment was obtained using a Leica DM4500P microscope.The area proportions of hematite,calcium ferrite,magnetite,calcium silicate and pore in pellet microstructure were extracted by visual extraction technology of image data.Moreover,the relationship between the area proportions of mineral components and compressive strength was established by backpropagation neural network(BPNN),generalized regression neural network(GRNN)and beetle antennae search-generalized regression neural network(BAS-GRNN)algorithms,which proves that the pellet microstructure can be used as the prediction standard of compressive strength.The errors of BPNN and BAS-GRNN are 5.13%and 3.37%,respectively,both of which are less than 5.5%.Therefore,through data visualization,we are able to discuss the connection between various components of pellet microstructure and compressive strength and provide new research ideas for improving the compressive strength and metallurgical performance of pellet.展开更多
Water resources are one of the basic resources for human survival,and water protection has been becoming a major problem for countries around the world.However,most of the traditional water quality monitoring research...Water resources are one of the basic resources for human survival,and water protection has been becoming a major problem for countries around the world.However,most of the traditional water quality monitoring research work is still concerned with the collection of water quality indicators,and ignored the analysis of water quality monitoring data and its value.In this paper,by adopting Laravel and AdminTE framework,we introduced how to design and implement a water quality data visualization platform based on Baidu ECharts.Through the deployed water quality sensor,the collected water quality indicator data is transmitted to the big data processing platform that deployed on Tencent Cloud in real time through the 4G network.The collected monitoring data is analyzed,and the processing result is visualized by Baidu ECharts.The test results showed that the designed system could run well and will provide decision support for water resource protection.展开更多
摘要This study focuses on meeting the challenges of big data visualization by using of data reduction methods based the feature selection methods.To reduce the volume of big data and minimize model training time(Tt)while maintaining data quality.We contributed to meeting the challenges of big data visualization using the embedded method based“Select from model(SFM)”method by using“Random forest Importance algorithm(RFI)”and comparing it with the filter method by using“Select percentile(SP)”method based chi square“Chi2”tool for selecting the most important features,which are then fed into a classification process using the logistic regression(LR)algorithm and the k-nearest neighbor(KNN)algorithm.Thus,the classification accuracy(AC)performance of LRis also compared to theKNN approach in python on eight data sets to see which method produces the best rating when feature selection methods are applied.Consequently,the study concluded that the feature selection methods have a significant impact on the analysis and visualization of the data after removing the repetitive data and the data that do not affect the goal.After making several comparisons,the study suggests(SFMLR)using SFM based on RFI algorithm for feature selection,with LR algorithm for data classify.The proposal proved its efficacy by comparing its results with recent literature.
基金supported by the Natural Science Foundation of China under Project 41076115the Global Change Research Program of China under project 2012CB955603the Public Science and Technology Research Funds of the Ocean under project 201005019
摘要The study of marine data visualization is of great value. Marine data, due to its large scale, random variation and multiresolution in nature, are hard to be visualized and analyzed. Nowadays, constructing an ocean model and visualizing model results have become some of the most important research topics of ‘Digital Ocean'. In this paper, a spherical ray casting method is developed to improve the traditional ray-casting algorithm and to make efficient use of GPUs. Aiming at the ocean current data, a 3D view-dependent line integral convolution method is used, in which the spatial frequency is adapted according to the distance from a camera. The study is based on a 3D virtual reality and visualization engine, namely the VV-Ocean. Some interactive operations are also provided to highlight the interesting structures and the characteristics of volumetric data. Finally, the marine data gathered in the East China Sea are displayed and analyzed. The results show that the method meets the requirements of real-time and interactive rendering.
摘要Data visualization blends art and science to convey stories from data via graphical representations.Considering different problems,applications,requirements,and design goals,it is challenging to combine these two components at their full force.While the art component involves creating visually appealing and easily interpreted graphics for users,the science component requires accurate representations of a large amount of input data.With a lack of the science component,visualization cannot serve its role of creating correct representations of the actual data,thus leading to wrong perception,interpretation,and decision.It might be even worse if incorrect visual representations were intentionally produced to deceive the viewers.To address common pitfalls in graphical representations,this paper focuses on identifying and understanding the root causes of misinformation in graphical representations.We reviewed the misleading data visualization examples in the scientific publications collected from indexing databases and then projected them onto the fundamental units of visual communication such as color,shape,size,and spatial orientation.Moreover,a text mining technique was applied to extract practical insights from common visualization pitfalls.Cochran’s Q test and McNemar’s test were conducted to examine if there is any difference in the proportions of common errors among color,shape,size,and spatial orientation.The findings showed that the pie chart is the most misused graphical representation,and size is the most critical issue.It was also observed that there were statistically significant differences in the proportion of errors among color,shape,size,and spatial orientation.
摘要The Growth Value Model(GVM)proposed theoretical closed form formulas consist-ing of Return on Equity(ROE)and the Price-to-Book value ratio(P/B)for fair stock prices and expected rates of return.Although regression analysis can be employed to verify these theoretical closed form formulas,they cannot be explored by classical quintile or decile sorting approaches with intuition due to the essence of multi-factors and dynamical processes.This article uses visualization techniques to help intuitively explore GVM.The discerning findings and contributions of this paper is that we put forward the concept of the smart frontier,which can be regarded as the reasonable lower limit of P/B at a specific ROE by exploring fair P/B with ROE-P/B 2D dynamical process visualization.The coefficients in the formula can be determined by the quantile regression analysis with market data.The moving paths of the ROE and P/B in the cur-rent quarter and the subsequent quarters show that the portfolios at the lower right of the curve approaches this curve and stagnates here after the portfolios are formed.Furthermore,exploring expected rates of return with ROE-P/B-Return 3D dynamical process visualization,the results show that the data outside of the lower right edge of the“smart frontier”has positive quarterly return rates not only in the t+1 quarter but also in the t+2 quarter.The farther away the data in the t quarter is from the“smart frontier”,the larger the return rates in the t+1 and t+2 quarter.
摘要Exploration of artworks is enjoyable but often time consuming.For example,it is not always easy to discover the favorite types of unknown painting works.It is not also always easy to explore unpopular painting works which looks similar to painting works created by famous artists.This paper presents a painting image browser which assists the explorative discovery of user-interested painting works.The presented browser applies a new multidimensional data visualization technique that highlights particular ranges of particular numeric values based on association rules to suggest cues to find favorite painting images.This study assumes a large number of painting images are provided where categorical information(e.g.,names of artists,created year)is assigned to the images.The presented system firstly calculates the feature values of the images as a preprocessing step.Then the browser visualizes the multidimensional feature values as a heatmap and highlights association rules discovered from the relationships between the feature values and categorical information.This mechanism enables users to explore favorite painting images or painting images that look similar to famous painting works.Our case study and user evaluation demonstrates the effectiveness of the presented image browser.
基金Grant Number 7059904 on the“National Evaluation Platform Approach for Accountability in Women’s and Children’s Health”from the Department of Global Affairs Canada to the Institute for International Programs at the Johns Hopkins Bloomberg School of Public Health.
摘要Background:Reproductive,maternal,newborn,child health,and nutrition(RMNCH&N)data is an indispensable tool for program and policy decisions in low-and middle-income countries.However,being equipped with evidence doesn’t necessarily translate to program and policy changes.This study aimed to characterize data visualization interpretation capacity and preferences among RMNCH&N Tanzanian program implementers and policymakers(“decision-makers”)to design more effective approaches towards promoting evidence-based RMNCH&N decisions in Tanzania.Methods:We conducted 25 semi-structured interviews in Kiswahili with junior,mid-level,and senior RMNCH&N decision-makers working in Tanzanian government institutions.We used snowball sampling to recruit participants with different rank and roles in RMNCH&N decision-making.Using semi-structured interviews,we probed participants on their statistical skills and data use,and asked participants to identify key messages and rank prepared RMNCH&N visualizations.We used a grounded theory approach to organize themes and identify findings.Results:The findings suggest that data literacy and statistical skills among RMNCH&N decision-makers in Tanzania varies.Most participants demonstrated awareness of many critical factors that should influence a visualization choice—audience,key message,simplicity—but assessments of data interpretation and preferences suggest that there may be weak knowledge of basic statistics.A majority of decision-makers have not had any statistical training since attending university.There appeared to be some discomfort with interpreting and using visualizations that are not bar charts,pie charts,and maps.Conclusions:Decision-makers must be able to understand and interpret RMNCH&N data they receive to be empowered to act.Addressing inadequate data literacy and presentation skills among decision-makers is vital to bridging gaps between evidence and policymaking.It would be beneficial to host basic data literacy and visualization training for RMNCH&N decision-makers at all levels in Tanzania,and to expand skills on developing key messages from visualizations.
基金This research was funded by the National Natural Science Foundation of China(No.51775185)Scientific Research Fund of Hunan Province Education Department(18C0003)+2 种基金Research project on teaching reform in colleges and universities of Hunan Province Education Department(20190147)Innovation and Entrepreneurship Training Program for College Students in Hunan Province(2021-1980)Hunan Normal University University-Industry Cooperation.This work is implemented at the 2011 Collaborative Innovation Center for Development and Utilization of Finance and Economics Big Data Property,Universities of Hunan Province,Open project,Grant Number 20181901CRP04.
摘要Many countries are paying more and more attention to the protection of water resources at present,and how to protect water resources has received extensive attention from society.Water quality monitoring is the key work to water resources protection.How to efficiently collect and analyze water quality monitoring data is an important aspect of water resources protection.In this paper,python programming tools and regular expressions were used to design a web crawler for the acquisition of water quality monitoring data from Global Freshwater Quality Database(GEMStat)sites,and the multi-thread parallelism was added to improve the efficiency in the process of downloading and parsing.In order to analyze and process the crawled water quality data,Pandas and Pyecharts are used to visualize the water quality data to show the intrinsic correlation and spatiotemporal relationship of the data.
摘要The advent of the big data era has made data visualization a crucial tool for enhancing the efficiency and insights of data analysis. This theoretical research delves into the current applications and potential future trends of data visualization in big data analysis. The article first systematically reviews the theoretical foundations and technological evolution of data visualization, and thoroughly analyzes the challenges faced by visualization in the big data environment, such as massive data processing, real-time visualization requirements, and multi-dimensional data display. Through extensive literature research, it explores innovative application cases and theoretical models of data visualization in multiple fields including business intelligence, scientific research, and public decision-making. The study reveals that interactive visualization, real-time visualization, and immersive visualization technologies may become the main directions for future development and analyzes the potential of these technologies in enhancing user experience and data comprehension. The paper also delves into the theoretical potential of artificial intelligence technology in enhancing data visualization capabilities, such as automated chart generation, intelligent recommendation of visualization schemes, and adaptive visualization interfaces. The research also focuses on the role of data visualization in promoting interdisciplinary collaboration and data democratization. Finally, the paper proposes theoretical suggestions for promoting data visualization technology innovation and application popularization, including strengthening visualization literacy education, developing standardized visualization frameworks, and promoting open-source sharing of visualization tools. This study provides a comprehensive theoretical perspective for understanding the importance of data visualization in the big data era and its future development directions.
摘要This study aims to explore the application of Bayesian analysis based on neural networks and deep learning in data visualization.The research background is that with the increasing amount and complexity of data,traditional data analysis methods have been unable to meet the needs.Research methods include building neural networks and deep learning models,optimizing and improving them through Bayesian analysis,and applying them to the visualization of large-scale data sets.The results show that the neural network combined with Bayesian analysis and deep learning method can effectively improve the accuracy and efficiency of data visualization,and enhance the intuitiveness and depth of data interpretation.The significance of the research is that it provides a new solution for data visualization in the big data environment and helps to further promote the development and application of data science.
摘要Financial literacy empowers individuals to make informed and effective financial decisions,improving their overall financial well-being and security.However,for many people understanding financial concepts can be daunting and only half of US adults are considered financially literate.Data visualization simplifies these concepts,making them accessible and engaging for learners of all ages.This systematic review analyzes 37 research papers exploring the use of data visualization and visual analytics in financial education and literacy enhancement.We classify these studies into five key areas:(1)the evolution of visualization use across time and space,(2)motivations for using visualization tools,(3)the financial topics addressed and instructional approaches used,(4)the types of tools and technologies applied,and(5)how the effectiveness of teaching interventions was evaluated.Furthermore,we identify research gaps and highlight opportunities for advancing financial literacy.Our findings offer practical insights for educators and professionals to effectively utilize or design visual tools for financial literacy.
基金supported by the National Key Research and Development Program of China(2022YFA0912100)the National Natural Science Foundation of China(32270098 and 32470073)+1 种基金the Fundamental Research Funds for the Central Universities(2662024JC015)the National Key Laboratory of Agricultural Microbiology(AML2024D02)to Z.Z.
摘要Translation is a crucial step in gene expression.Over the past decade,the development and application of ribosome profiling(Ribo-seq)have significantly advanced our understanding of translational regulation in vivo.However,the analysis and visualization of Ribo-seq data remain challenging.Despite the availability of various analytical pipelines,improvements in comprehensiveness,accuracy,and user-friendliness are still necessary.In this study,we develop RiboParser/RiboShiny,a robust framework for analyzing and visualizing Ribo-seq data.Building on published methods,we optimize ribosome structure-based and start/stopbased models to improve the accuracy and stability of P-site detection,even in species with a high proportion of leaderless transcripts.Leveraging these improvements,RiboParser offers comprehensive analyses,including quality control,gene-level analysis,codon-level analysis,and the analysis of Ribo-seq variants.Meanwhile,RiboShiny provides a user-friendly and adaptable platform for data visualization,facilitating deeper insights into the translational landscape.Furthermore,the integration of standardized genome annotation renders our platform universally applicable to various organisms with sequenced genomes.This framework has the potential to significantly improve the precision and efficiency of Ribo-seq data interpretation,thereby deepening our understanding of translational regulation.
摘要In the digital age,where the amount of data is growing exponentially,traditional big data visualization methods often fall short,especially when it comes to processing high-dimensional data or matching visualization results to user needs.This highlights the need for better integration between Artificial Intelligence(AI)and big data visualization.AI has brought about significant changes in this field.This paper reviews research from the past five years to explore the integration of AI with big data visualization.It looks at the core technologies,common use cases,and challenges of this combination.The review finds that AI can address many issues in traditional visualization techniques,such as handling complex data and improving interaction and evaluation.However,challenges like model interpretability,data privacy concerns,and inconsistent standards remain.In the future,this integration is expected to move toward more transparent models,stronger privacy protection,and standardized systems.This will help unlock the full value of data in real-world applications.
摘要With the further advancement of China’s strategy of building a maritime power, marine surveying and mapping, as the core fundamental work for marine development and protection, has witnessed an exponential growth in data volume. The traditional file-based management architecture and static visualization mode can hardly meet the strategic demands of modern marine informatization construction. With its outstanding capabilities in spatial data processing, analysis and representation, GIS technology provides systematic technical support for the refined management and efficient visualization of marine surveying and mapping data. Combined with the inherent characteristics and management pain points of marine surveying and mapping data, this paper systematically discusses the application mechanisms of GIS technology in core technical links, including multi-source heterogeneous data integration and standardization, spatiotemporal database construction and efficient retrieval, data quality control and dynamic update, as well as 2D-3D integrated visualization, dynamic spatiotemporal process simulation, and interactive decision support. It aims to provide theoretical references and practical experience for improving the resource utilization efficiency and professional service capacity of marine surveying and mapping data.
摘要Earthquakes are highly destructive spatio-temporal phenomena whose analysis is essential for disaster preparedness and risk mitigation.Modern seismological research produces vast volumes of heterogeneous data from seismic networks,satellite observations,and geospatial repositories,creating the need for scalable infrastructures capable of integrating and analyzing such data to support intelligent decision-making.Data warehousing technologies provide a robust foundation for this purpose;however,existing earthquake-oriented data warehouses remain limited,often relying on simplified schemas,domain-specific analytics,or cataloguing efforts.This paper presents the design and implementation of a spatio-temporal data warehouse for seismic activity.The framework integrates spatial and temporal dimensions in a unified schema and introduces a novel array-based approach for managing many-to-many relationships between facts and dimensions without intermediate bridge tables.A comparative evaluation against a conventional bridge-table schema demonstrates that the array-based design improves fact-centric query performance,while the bridge-table schema remains advantageous for dimension-centric queries.To reconcile these trade-offs,a hybrid schema is proposed that retains both representations,ensuring balanced efficiency across heterogeneous workloads.The proposed framework demonstrates how spatio-temporal data warehousing can address schema complexity,improve query performance,and support multidimensional visualization.In doing so,it provides a foundation for integrating seismic analysis into broader big data-driven intelligent decision systems for disaster resilience,risk mitigation,and emergency management.
基金supported by the National Natural Science Foundation of China(NSFC)Grant Nos.61702271,61702270.
摘要The widespread use of numerical simulations in different scientific domains provides a variety of research opportunities.They often output a great deal of spatio-temporal simulation data,which are traditionally characterized as single-run,multi-run,multi-variate,multi-modal and multi-dimensional.From the perspective of data exploration and analysis,we noticed that many works focusing on spatiotemporal simulation data often share similar exploration techniques,for example,the exploration schemes designed in simulation space,parameter space,feature space and combinations of them.However,it lacks a survey to have a systematic overview of the essential commonalities shared by those works.In this survey,we take a novel multi-space perspective to categorize the state-ofthe-art works into three major categories.Specifically,the works are characterized as using similar techniques such as visual designs in simulation space(e.g,visual mapping,boxplot-based visual summarization,etc.),parameter space analysis(e.g,visual steering,parameter space projection,etc.)and data processing in feature space(e.g,feature definition and extraction,sampling,reduction and clustering of simulation data,etc.).
基金supported by National Natural Science Foundation of China (No.10835009)Chinese Academy of Sciences for the Key Project of Knowledge Innovation Program (No.KJCX3.SYW.N4)Chinese Ministry of Sciences for the 973 project (No.2009GB103000)
摘要A visualization tool was developed through a web browser based on Java applets embedded into HTML pages, in order to provide a world access to the EAST experimental data. It can display data from various trees in different servers in a single panel. With WebScope, it is easier to make a comparison between different data sources and perform a simple calculation over different data sources.
摘要Cyber security has been thrust into the limelight in the modern technological era because of an array of attacks often bypassing tmtrained intrusion detection systems (IDSs). Therefore, greater attention has been directed on being able deciphering better methods for identifying attack types to train IDSs more effectively. Keycyber-attack insights exist in big data; however, an efficient approach is required to determine strong attack types to train IDSs to become more effective in key areas. Despite the rising growth in IDS research, there is a lack of studies involving big data visualization, which is key. The KDD99 data set has served as a strong benchmark since 1999; therefore, we utilized this data set in our experiment. In this study, we utilized hash algorithm, a weight table, and sampling method to deal with the inherent problems caused by analyzing big data; volume, variety, and velocity. By utilizing a visualization algorithm, we were able to gain insights into the KDD99 data set with a clear iden- tification of "normal" clusters and described distinct clusters of effective attacks.
基金This document contains the results of the research project funded by the National Social Science Fund of China (19ZDA046)NSF of China (62302440,U22A2032)+1 种基金China Postdoctoral Science Foundation (2023TQ0288)the Fundamental Research Funds for the Central Universities,China.
摘要Fine art authentication plays a significant role in protecting cultural heritage and ensuring the integrity of artworks.Traditional authentication methods require professionals to collect many reference materials and conduct detailed analyses.To ease the difficulty,we collaborate with domain experts to develop a GPT-based agent,namely ArtEyer,that offers accurate attributions,determines the origin and authorship,and executes visual analytics.Despite the convenience of the conversational user interface,novice users may still face challenges due to the hallucination issue and the steep learning curve associated with prompting.To face these obstacles,we propose a novel solution that places interactive data visualizations into the conversations.We create contextual visualizations from an external domain-dependent database to ensure data trustworthiness and allow users to provide precise instructions to the agent by interacting directly with these visualizations,thus overcoming the vagueness inherent in natural language-based prompting.We evaluate ArtEyer through an in-lab user study and demonstrate its usage with a real-world case.
基金supported by the National Natural Science Foundation of China(51674121)Fund for Distinguished Youth Scholars in North China University of Science and Technology(JQ201705).
摘要In recent years,with the wide application of image data visual extraction technology in the field of industrial engineering,the development of industrial economy has reached a new situation.To explore the interaction between the pellet microstructure and compressive strength,firstly,the pellet microstructure needed for the experiment was obtained using a Leica DM4500P microscope.The area proportions of hematite,calcium ferrite,magnetite,calcium silicate and pore in pellet microstructure were extracted by visual extraction technology of image data.Moreover,the relationship between the area proportions of mineral components and compressive strength was established by backpropagation neural network(BPNN),generalized regression neural network(GRNN)and beetle antennae search-generalized regression neural network(BAS-GRNN)algorithms,which proves that the pellet microstructure can be used as the prediction standard of compressive strength.The errors of BPNN and BAS-GRNN are 5.13%and 3.37%,respectively,both of which are less than 5.5%.Therefore,through data visualization,we are able to discuss the connection between various components of pellet microstructure and compressive strength and provide new research ideas for improving the compressive strength and metallurgical performance of pellet.
基金This work is supported by National Natural Science Foundation of China 61304208by the 2011 Collaborative Innovation Center for Development and Utilization of Finance and Economics Big Data Property Open Fund Project 20181901CRP04+2 种基金by the Scientific Research Fund of Hunan Province Education Department 18C0003by the Research Project on Teaching Reform in General Colleges and Universities,Hunan Provincial Education Department 20190147by the Hunan Normal University Ungraduated Innovation and Entrepreneurship Training Plan Project 2019127.
摘要Water resources are one of the basic resources for human survival,and water protection has been becoming a major problem for countries around the world.However,most of the traditional water quality monitoring research work is still concerned with the collection of water quality indicators,and ignored the analysis of water quality monitoring data and its value.In this paper,by adopting Laravel and AdminTE framework,we introduced how to design and implement a water quality data visualization platform based on Baidu ECharts.Through the deployed water quality sensor,the collected water quality indicator data is transmitted to the big data processing platform that deployed on Tencent Cloud in real time through the 4G network.The collected monitoring data is analyzed,and the processing result is visualized by Baidu ECharts.The test results showed that the designed system could run well and will provide decision support for water resource protection.