Given a multi-turn conversational context and a raw user query,the goal of Conversational Query Reformulation(CQR)is to transform the query into a de-contextualized form that maximizes retrieval effectiveness for a do...Given a multi-turn conversational context and a raw user query,the goal of Conversational Query Reformulation(CQR)is to transform the query into a de-contextualized form that maximizes retrieval effectiveness for a downstream passage retriever.Conversational search seeks to retrieve relevant passages for the given questions in a conversational question answering system.Conversational Query Reformulation(CQR)improves conversational search by refining the original queries into de-contextualized forms to address issues such as omissions and coreferences.Previous CQR methods focus on imitating human-written queries,which may not always yield meaningful search results for the retriever.In this paper,we introduce GuideCQR,a framework that refines queries for CQR by leveraging key information from the initially retrieved documents.Specifically,GuideCQR extracts keywords and generates expected answers from the retrieved documents,then unifies them with the queries after filtering to add useful information that enhances the search process.Experimental results demonstrate that our proposed method achieves state-of-the-art performance across multiple datasets,outperforming previous CQR methods.Specifically,GuideCQR achieves MRR gains of 5.4%over LLM4CS on CAsT-19 and NDCG@3 gains of 29.2%on QReCC,and consistently improves retrieval across CAsT-19,CAsT-20,and QReCC benchmarks,demonstrating strong adaptability to various query types including human-rewritten queries.Additionally,we show that GuideCQR can get additional performance gains in conversational search using various types of queries,even for queries written by humans.展开更多
With the development of anti-virus technology,malicious documents have gradually become the main pathway of Advanced Persistent Threat(APT)attacks,therefore,the development of effective malicious document classifiers ...With the development of anti-virus technology,malicious documents have gradually become the main pathway of Advanced Persistent Threat(APT)attacks,therefore,the development of effective malicious document classifiers has become particularly urgent.Currently,detection methods based on document structure and behavioral features encounter challenges in feature engineering,these methods not only have limited accuracy,but also consume large resources,and usually can only detect documents in specific formats,which lacks versatility and adaptability.To address such problems,this paper proposes a novel malicious document detection method-visualizing documents as GGE images(Grayscale,Grayscale matrix,Entropy).The GGE method visualizes the original byte sequence of the malicious document as a grayscale image,the information entropy sequence of the document as an entropy image,and at the same time,the grayscale level co-occurrence matrix and the texture and spatial information stored in it are converted into grayscale matrix image,and fuses the three types of images to get the GGE color image.The Convolutional Block Attention Module-EfficientNet-B0(CBAM-EfficientNet-B0)model is then used for classification,combining transfer learning and applying the pre-trained model on the ImageNet dataset to the feature extraction process of GGE images.As shown in the experimental results,the GGE method has superior performance compared with other methods,which is suitable for detecting malicious documents in different formats,and achieves an accuracy of 99.44%and 97.39%on Portable Document Format(PDF)and office datasets,respectively,and consumes less time during the detection process,which can be effectively applied to the task of detecting malicious documents in real-time.展开更多
The rapid advancement of Large Language Models(LLMs)has enabled their application in diverse professional domains,including law.However,research on automatic judicial document generation remains limited,particularly f...The rapid advancement of Large Language Models(LLMs)has enabled their application in diverse professional domains,including law.However,research on automatic judicial document generation remains limited,particularly for taiwan region of China courts.This study proposes a keyword-guided training framework that enhances LLMs’ability to generate structured and semantically coherent judicial decisions in Chinese.The proposed method first employs LLMs to extract representative legal keywords from absolute court judgments.Then it integrates these keywords into Supervised Fine-Tuning(SFT)and Reinforcement Learning withHuman Feedback using Proximal Policy Optimization(RLHF-PPO).Experimental evaluations using models such as Chinese Alpaca 7B and TAIDE-LX-7B demonstrate that keyword-guided training significantly improves generation quality,achieving ROUGE-1,ROUGE-2,and ROUGE-L score gains of up to 17%,16%,and 20%,respectively.The results confirm that the proposed framework effectively aligns generated judgments with human-written legal logic and structural conventions.This research advances domainadaptive LLM fine-tuning strategies and establishes a technical foundation forAI-assisted judicial document generation in the taiwan region of China legal context.This research provides empirical evidence that domain-adaptive LLM fine-tuning strategies can significantly improve performance in complex,structured legal text generation.展开更多
In the international shipping industry, digital intelligence transformation has become essential, with both governments and enterprises actively working to integrate diverse datasets. The domain of maritime and shippi...In the international shipping industry, digital intelligence transformation has become essential, with both governments and enterprises actively working to integrate diverse datasets. The domain of maritime and shipping is characterized by a vast array of document types, filled with complex, large-scale, and often chaotic knowledge and relationships. Effectively managing these documents is crucial for developing a Large Language Model (LLM) in the maritime domain, enabling practitioners to access and leverage valuable information. A Knowledge Graph (KG) offers a state-of-the-art solution for enhancing knowledge retrieval, providing more accurate responses and enabling context-aware reasoning. This paper presents a framework for utilizing maritime and shipping documents to construct a knowledge graph using GraphRAG, a hybrid tool combining graph-based retrieval and generation capabilities. The extraction of entities and relationships from these documents and the KG construction process are detailed. Furthermore, the KG is integrated with an LLM to develop a Q&A system, demonstrating that the system significantly improves answer accuracy compared to traditional LLMs. Additionally, the KG construction process is up to 50% faster than conventional LLM-based approaches, underscoring the efficiency of our method. This study provides a promising approach to digital intelligence in shipping, advancing knowledge accessibility and decision-making.展开更多
Adult sex estimation is one of the first and most important steps in forensic examination.While dealing with disturbed burials,the most dimorphic anatomic areas of the skeleton(such as the coxae,the skull and the head...Adult sex estimation is one of the first and most important steps in forensic examination.While dealing with disturbed burials,the most dimorphic anatomic areas of the skeleton(such as the coxae,the skull and the head of femur and humerus)may be deteriorated or fragmented.In contrast,the minimum supero-inferior femoral neck diameter(SID)is generally much better preserved.The aim of the present research is to identify the discriminatory potential of SID for sex estimation and to test different formulae and mathematical procedures currently available in the forensic literature,on a sample of 295 contemporary individuals from the 21st Century Identified Skeletal Collection(University of Coimbra,Portugal),in order to identify its relevance for application in Portuguese forensic cases.Results showed that SID is a dimorphic variable,with high frequencies and probabilities of cases correctly estimated(0.82 and 0.83,respectively);statistically significant differences between females and males,and a high association between the metrics and sex,were identified.Posterior probabilities allow reliable estimations for all the measurements,excepting those between 31.0 and 31.5 mm,and the procedures that show the highest accuracy are those proposed by Seidemann et al.(1998),Curate et al.(2016),and Luna et al.(2021).Adult sex estimation from in a contemporary osteological sample from Buenos Aires,Argentina,with frequencies and probabilities between 0.82 and 0.83 for both sexes.The validation procedures implemented in this study highlight both the need to test quantitative models generated from diverse contemporary human populations,and the value of SID for obtaining reliable adult sex estimates,as they improve the quality of the biological profiles obtained in forensic contexts.展开更多
This critical review looks at the assessment of the application of artificial intelligence in handling legal documents with specific reference to medical negligence cases with a view of identifying its transformative ...This critical review looks at the assessment of the application of artificial intelligence in handling legal documents with specific reference to medical negligence cases with a view of identifying its transformative potentialities, issues and ethical concerns. The review consolidates findings that show the impact of AI in improving the efficiency, accuracy and justice delivery in the legal profession. The studies show increased efficiency in speed of document review and enhancement of the accuracy of the reviewed documents, with time efficiency estimates of 60% reduction of time. However, the review also outlines some of the problems that continue to characterize AI, such as data quality problems, biased algorithms and the problem of the opaque decision-making system. This paper assesses ethical issues related to patient autonomy, justice and non-malignant suffering, with particular focus on patient privacy and fair process, and on potential unfairness to patients. This paper’s review of AI innovations finds that regulations lag behind AI developments, leading to unsettled issues regarding legal responsibility for AI and user control over AI-generated results and findings in legal proceedings. Some of the future avenues that are presented in the study are the future of XAI for legal purposes, utilizing federated learning for resolving privacy issues, and the need to foster adaptive regulation. Finally, the review advocates for Legal Subject Matter Experts to collaborate with legal informatics experts, ethicists, and policy makers to develop the best solutions to implement AI in medical negligence claims. It reasons that there is great potential for AI to have a deep impact on the practice of law but when done, it must do so in a way that respects justice and on the Rights of Individuals.展开更多
In the process of building a new power system dominated by new energy sources,power storage is a key supporting technology that ensures the safe and stable operation of the power grid,enables the flexible regulation o...In the process of building a new power system dominated by new energy sources,power storage is a key supporting technology that ensures the safe and stable operation of the power grid,enables the flexible regulation of the system,and raises the level of new energy consumption.It is also key to achieving carbon peak and neutrality as well as energy transformation.展开更多
This video series is the first experimental psychology documentary made in China.It focuses on analyzing professional theories to raise people’s general understanding of basic psychology.By combining innovative audio...This video series is the first experimental psychology documentary made in China.It focuses on analyzing professional theories to raise people’s general understanding of basic psychology.By combining innovative audiovisual narrative with psychological experiments,it zooms in on real human nature through discussing social hotspots from the perspectives of social psychology,cognitive psychology,and personality psychology,in order to help people find answers for their current psychological difficulties.展开更多
What Are You Up To Today?Chief Director:Wu Zijuan Length:12 Episodes Producer:bilibili Broadcasting Platform:bilibili Produced by China’s YouTube-like video sharing platform bilibili,the film is a series of short doc...What Are You Up To Today?Chief Director:Wu Zijuan Length:12 Episodes Producer:bilibili Broadcasting Platform:bilibili Produced by China’s YouTube-like video sharing platform bilibili,the film is a series of short documentaries presenting people's daily life in different jobs.It follows 12 individuals in their respective jobs and trades that keep society functioning.By focusing on their daily lives,the documentary films capture the hustle and bustle of the days that make up a hopeful life.展开更多
In nursing practice,electronic nursing records(ENRs)are an important component of patient care documents,but they also significantly increase administrative burdens.With the development of artificial intelligence tech...In nursing practice,electronic nursing records(ENRs)are an important component of patient care documents,but they also significantly increase administrative burdens.With the development of artificial intelligence technology,it has become possible to use large text models to assist in generating nursing documents.This article explores the application of generative AI in nursing documentation.Research has shown that the application of generative AI in nursing documents demonstrates significant potential,but also faces challenges in terms of quality and implementation.In terms of efficiency,AI assisted document tools can significantly reduce the administrative burden on nurses by reallocating time to direct patient care.Studies have shown that they can reduce document time by 21-30%.However,there are variables in the quality of AI generated records,and the content is often described as'textbook style',lacking patient specific details and appropriate medical terminology.Successful implementation relies on a specialized framework that includes strong stakeholder engagement and adaptation to nursing specific workflows and regulatory standards.The conclusion points out that current AI systems are most suitable for assisting in drafting nursing documents,and clinical validation remains crucial for patient safety and document integrity.展开更多
This paper highlights the critical role of medical device design and development documents within the quality system,including their compliance with regulatory standards,their function as a traceable record,their supp...This paper highlights the critical role of medical device design and development documents within the quality system,including their compliance with regulatory standards,their function as a traceable record,their support for all stages,and their use in risk and change management.It also covers document template creation,review record association,information management,adverse event traceability,and the reconciliation of differences in international declarations.展开更多
Engineers often need to look for the right pieces of information by sifting through long engineering documents, It is a very tiring and time-consuming job. To address this issue, researchers are increasingly devoting ...Engineers often need to look for the right pieces of information by sifting through long engineering documents, It is a very tiring and time-consuming job. To address this issue, researchers are increasingly devoting their attention to new ways to help information users, including engineers, to access and retrieve document content. The research reported in this paper explores how to use the key technologies of document decomposition (study of document structure), document mark-up (with EXtensible Mark- up Language (XML), HyperText Mark-up Language (HTML), and Scalable Vector Graphics (SVG)), and a facetted classification mechanism. Document content extraction is implemented via computer programming (with Java). An Engineering Document Content Management System (EDCMS) developed in this research demonstrates that as information providers we can make document content in a more accessible manner for information users including engineers.The main features of the EDCMS system are: 1) EDCMS is a system that enables users, especially engineers, to access and retrieve information at content rather than document level. In other words, it provides the right pieces of information that answer specific questions so that engineers don't need to waste time sifting through the whole document to obtain the required piece of information. 2) Users can use the EDCMS via both the data and metadata of a document to access engineering document content. 3) Users can use the EDCMS to access and retrieve content objects, i.e. text, images and graphics (including engineering drawings) via multiple views and at different granularities based on decomposition schemes. Experiments with the EDCMS have been conducted on semi-structured documents, a textbook of CADCAM, and a set of project posters in the Engineering Design domain. Experimental results show that the system provides information users with a powerful solution to access document content.展开更多
Achieving a good recognition rate for degraded document images is difficult as degraded document images suffer from low contrast,bleedthrough,and nonuniform illumination effects.Unlike the existing baseline thresholdi...Achieving a good recognition rate for degraded document images is difficult as degraded document images suffer from low contrast,bleedthrough,and nonuniform illumination effects.Unlike the existing baseline thresholding techniques that use fixed thresholds and windows,the proposed method introduces a concept for obtaining dynamic windows according to the image content to achieve better binarization.To enhance a low-contrast image,we proposed a new mean histogram stretching method for suppressing noisy pixels in the background and,simultaneously,increasing pixel contrast at edges or near edges,which results in an enhanced image.For the enhanced image,we propose a new method for deriving adaptive local thresholds for dynamic windows.The dynamic window is derived by exploiting the advantage of Otsu thresholding.To assess the performance of the proposed method,we have used standard databases,namely,document image binarization contest(DIBCO),for experimentation.The comparative study on well-known existing methods indicates that the proposed method outperforms the existing methods in terms of quality and recognition rate.展开更多
We are delighted to have been invited to be guest editors for this special issue on forensic document examination,a forensic science discipline that has had a long history of use in both criminal and civil investigati...We are delighted to have been invited to be guest editors for this special issue on forensic document examination,a forensic science discipline that has had a long history of use in both criminal and civil investigations.The work of the forensic document examiner(FDE)can encompass handwriting(including signature)comparisons,and examinations and evaluations of physical components of questioned documents such as paper,ink,toner,and impressions(of writing and stamps)to answer questions of authenticity or source.展开更多
A document layout can be more informative than merely a document’s visual and structural appearance.Thus,document layout analysis(DLA)is considered a necessary prerequisite for advanced processing and detailed docume...A document layout can be more informative than merely a document’s visual and structural appearance.Thus,document layout analysis(DLA)is considered a necessary prerequisite for advanced processing and detailed document image analysis to be further used in several applications and different objectives.This research extends the traditional approaches of DLA and introduces the concept of semantic document layout analysis(SDLA)by proposing a novel framework for semantic layout analysis and characterization of handwritten manuscripts.The proposed SDLA approach enables the derivation of implicit information and semantic characteristics,which can be effectively utilized in dozens of practical applications for various purposes,in a way bridging the semantic gap and providingmore understandable high-level document image analysis and more invariant characterization via absolute and relative labeling.This approach is validated and evaluated on a large dataset ofArabic handwrittenmanuscripts comprising complex layouts.The experimental work shows promising results in terms of accurate and effective semantic characteristic-based clustering and retrieval of handwritten manuscripts.It also indicates the expected efficacy of using the capabilities of the proposed approach in automating and facilitating many functional,reallife tasks such as effort estimation and pricing of transcription or typing of such complex manuscripts.展开更多
Semantic segmentation is a crucial step for document understanding.In this paper,an NVIDIA Jetson Nano-based platform is applied for implementing semantic segmentation for teaching artificial intelligence concepts and...Semantic segmentation is a crucial step for document understanding.In this paper,an NVIDIA Jetson Nano-based platform is applied for implementing semantic segmentation for teaching artificial intelligence concepts and programming.To extract semantic structures from document images,we present an end-to-end dilated convolution network architecture.Dilated convolutions have well-known advantages for extracting multi-scale context information without losing spatial resolution.Our model utilizes dilated convolutions with residual network to represent the image features and predicting pixel labels.The convolution part works as feature extractor to obtain multidimensional and hierarchical image features.The consecutive deconvolution is used for producing full resolution segmentation prediction.The probability of each pixel decides its predefined semantic class label.To understand segmentation granularity,we compare performances at three different levels.From fine grained class to coarse class levels,the proposed dilated convolution network architecture is evaluated on three document datasets.The experimental results have shown that both semantic data distribution imbalance and network depth are import factors that influence the document’s semantic segmentation performances.The research is aimed at offering an education resource for teaching artificial intelligence concepts and techniques.展开更多
Background:Contrast enhancement plays an important role in the image processing field.Contrast correction has performed an adjustment on the darkness or brightness of the input image and increases the quality of the i...Background:Contrast enhancement plays an important role in the image processing field.Contrast correction has performed an adjustment on the darkness or brightness of the input image and increases the quality of the image.Objective:This paper proposed a novel method based on statistical data from the local mean and local standard deviation.Method:The proposed method modifies the mean and standard deviation of a neighbourhood at each pixel and divides it into three categories:background,foreground,and problematic(contrast&luminosity)region.Experimental results from both visual and objective aspects show that the proposed method can normalize the contrast variation problem effectively compared to Histogram Equalization(HE),Difference of Gaussian(DoG),and Butterworth Homomorphic Filtering(BHF).Seven(7)types of binarization methods were tested on the corrected image and produced a positive and impressive result.Result:Finally,a comparison in terms of Signal Noise Ratio(SNR),Misclassification Error(ME),F-measure,Peak Signal Noise Ratio(PSNR),Misclassification Penalty Metric(MPM),and Accuracy was calculated.Each binarization method shows an incremented result after applying it onto the corrected image compared to the original image.The SNR result of our proposed image is 9.350 higher than the three(3)other methods.The average increment after five(5)types of evaluation are:(Otsu=41.64%,Local Adaptive=7.05%,Niblack=30.28%,Bernsen=25%,Bradley=3.54%,Nick=1.59%,Gradient-Based=14.6%).Conclusion:The results presented in this paper effectively solve the contrast problem and finally produce better quality images.展开更多
Text extraction is an important initial step in digitizing the historical documents. In this paper, we present a text extraction method for historical Tibetan document images based on block projections. The task of te...Text extraction is an important initial step in digitizing the historical documents. In this paper, we present a text extraction method for historical Tibetan document images based on block projections. The task of text extraction is considered as text area detection and location problem. The images are divided equally into blocks and the blocks are filtered by the information of the categories of connected components and corner point density. By analyzing the filtered blocks' projections, the approximate text areas can be located, and the text regions are extracted. Experiments on the dataset of historical Tibetan documents demonstrate the effectiveness of the proposed method.展开更多
A rough set based corner classification neural network, the Rough-CC4, is presented to solve document classification problems such as document representation of different document sizes, document feature selection and...A rough set based corner classification neural network, the Rough-CC4, is presented to solve document classification problems such as document representation of different document sizes, document feature selection and document feature encoding. In the Rough-CC4, the documents are described by the equivalent classes of the approximate words. By this method, the dimensions representing the documents can be reduced, which can solve the precision problems caused by the different document sizes and also blur the differences caused by the approximate words. In the Rough-CC4, a binary encoding method is introduced, through which the importance of documents relative to each equivalent class is encoded. By this encoding method, the precision of the Rough-CC4 is improved greatly and the space complexity of the Rough-CC4 is reduced. The Rough-CC4 can be used in automatic classification of documents.展开更多
基金supported by the Institute of Information&Communications Technology Planning&Evaluation(IITP)grant funded by the Korea government(MSIT)[RS-2021-II211341,Artificial Intelligence Graduate School Program(Chung-Ang University)]Korea Institute for Advancement of Technology(KIAT)grant funded by the Korea Government(MOTIE)(RS-2025-25458133)supported by the Chung-Ang University Graduate Research Scholarship in 2025.
摘要Given a multi-turn conversational context and a raw user query,the goal of Conversational Query Reformulation(CQR)is to transform the query into a de-contextualized form that maximizes retrieval effectiveness for a downstream passage retriever.Conversational search seeks to retrieve relevant passages for the given questions in a conversational question answering system.Conversational Query Reformulation(CQR)improves conversational search by refining the original queries into de-contextualized forms to address issues such as omissions and coreferences.Previous CQR methods focus on imitating human-written queries,which may not always yield meaningful search results for the retriever.In this paper,we introduce GuideCQR,a framework that refines queries for CQR by leveraging key information from the initially retrieved documents.Specifically,GuideCQR extracts keywords and generates expected answers from the retrieved documents,then unifies them with the queries after filtering to add useful information that enhances the search process.Experimental results demonstrate that our proposed method achieves state-of-the-art performance across multiple datasets,outperforming previous CQR methods.Specifically,GuideCQR achieves MRR gains of 5.4%over LLM4CS on CAsT-19 and NDCG@3 gains of 29.2%on QReCC,and consistently improves retrieval across CAsT-19,CAsT-20,and QReCC benchmarks,demonstrating strong adaptability to various query types including human-rewritten queries.Additionally,we show that GuideCQR can get additional performance gains in conversational search using various types of queries,even for queries written by humans.
基金supported by the Natural Science Foundation of Henan Province(Grant No.242300420297)awarded to Yi Sun.
摘要With the development of anti-virus technology,malicious documents have gradually become the main pathway of Advanced Persistent Threat(APT)attacks,therefore,the development of effective malicious document classifiers has become particularly urgent.Currently,detection methods based on document structure and behavioral features encounter challenges in feature engineering,these methods not only have limited accuracy,but also consume large resources,and usually can only detect documents in specific formats,which lacks versatility and adaptability.To address such problems,this paper proposes a novel malicious document detection method-visualizing documents as GGE images(Grayscale,Grayscale matrix,Entropy).The GGE method visualizes the original byte sequence of the malicious document as a grayscale image,the information entropy sequence of the document as an entropy image,and at the same time,the grayscale level co-occurrence matrix and the texture and spatial information stored in it are converted into grayscale matrix image,and fuses the three types of images to get the GGE color image.The Convolutional Block Attention Module-EfficientNet-B0(CBAM-EfficientNet-B0)model is then used for classification,combining transfer learning and applying the pre-trained model on the ImageNet dataset to the feature extraction process of GGE images.As shown in the experimental results,the GGE method has superior performance compared with other methods,which is suitable for detecting malicious documents in different formats,and achieves an accuracy of 99.44%and 97.39%on Portable Document Format(PDF)and office datasets,respectively,and consumes less time during the detection process,which can be effectively applied to the task of detecting malicious documents in real-time.
摘要The rapid advancement of Large Language Models(LLMs)has enabled their application in diverse professional domains,including law.However,research on automatic judicial document generation remains limited,particularly for taiwan region of China courts.This study proposes a keyword-guided training framework that enhances LLMs’ability to generate structured and semantically coherent judicial decisions in Chinese.The proposed method first employs LLMs to extract representative legal keywords from absolute court judgments.Then it integrates these keywords into Supervised Fine-Tuning(SFT)and Reinforcement Learning withHuman Feedback using Proximal Policy Optimization(RLHF-PPO).Experimental evaluations using models such as Chinese Alpaca 7B and TAIDE-LX-7B demonstrate that keyword-guided training significantly improves generation quality,achieving ROUGE-1,ROUGE-2,and ROUGE-L score gains of up to 17%,16%,and 20%,respectively.The results confirm that the proposed framework effectively aligns generated judgments with human-written legal logic and structural conventions.This research advances domainadaptive LLM fine-tuning strategies and establishes a technical foundation forAI-assisted judicial document generation in the taiwan region of China legal context.This research provides empirical evidence that domain-adaptive LLM fine-tuning strategies can significantly improve performance in complex,structured legal text generation.
摘要In the international shipping industry, digital intelligence transformation has become essential, with both governments and enterprises actively working to integrate diverse datasets. The domain of maritime and shipping is characterized by a vast array of document types, filled with complex, large-scale, and often chaotic knowledge and relationships. Effectively managing these documents is crucial for developing a Large Language Model (LLM) in the maritime domain, enabling practitioners to access and leverage valuable information. A Knowledge Graph (KG) offers a state-of-the-art solution for enhancing knowledge retrieval, providing more accurate responses and enabling context-aware reasoning. This paper presents a framework for utilizing maritime and shipping documents to construct a knowledge graph using GraphRAG, a hybrid tool combining graph-based retrieval and generation capabilities. The extraction of entities and relationships from these documents and the KG construction process are detailed. Furthermore, the KG is integrated with an LLM to develop a Q&A system, demonstrating that the system significantly improves answer accuracy compared to traditional LLMs. Additionally, the KG construction process is up to 50% faster than conventional LLM-based approaches, underscoring the efficiency of our method. This study provides a promising approach to digital intelligence in shipping, advancing knowledge accessibility and decision-making.
摘要Adult sex estimation is one of the first and most important steps in forensic examination.While dealing with disturbed burials,the most dimorphic anatomic areas of the skeleton(such as the coxae,the skull and the head of femur and humerus)may be deteriorated or fragmented.In contrast,the minimum supero-inferior femoral neck diameter(SID)is generally much better preserved.The aim of the present research is to identify the discriminatory potential of SID for sex estimation and to test different formulae and mathematical procedures currently available in the forensic literature,on a sample of 295 contemporary individuals from the 21st Century Identified Skeletal Collection(University of Coimbra,Portugal),in order to identify its relevance for application in Portuguese forensic cases.Results showed that SID is a dimorphic variable,with high frequencies and probabilities of cases correctly estimated(0.82 and 0.83,respectively);statistically significant differences between females and males,and a high association between the metrics and sex,were identified.Posterior probabilities allow reliable estimations for all the measurements,excepting those between 31.0 and 31.5 mm,and the procedures that show the highest accuracy are those proposed by Seidemann et al.(1998),Curate et al.(2016),and Luna et al.(2021).Adult sex estimation from in a contemporary osteological sample from Buenos Aires,Argentina,with frequencies and probabilities between 0.82 and 0.83 for both sexes.The validation procedures implemented in this study highlight both the need to test quantitative models generated from diverse contemporary human populations,and the value of SID for obtaining reliable adult sex estimates,as they improve the quality of the biological profiles obtained in forensic contexts.
摘要This critical review looks at the assessment of the application of artificial intelligence in handling legal documents with specific reference to medical negligence cases with a view of identifying its transformative potentialities, issues and ethical concerns. The review consolidates findings that show the impact of AI in improving the efficiency, accuracy and justice delivery in the legal profession. The studies show increased efficiency in speed of document review and enhancement of the accuracy of the reviewed documents, with time efficiency estimates of 60% reduction of time. However, the review also outlines some of the problems that continue to characterize AI, such as data quality problems, biased algorithms and the problem of the opaque decision-making system. This paper assesses ethical issues related to patient autonomy, justice and non-malignant suffering, with particular focus on patient privacy and fair process, and on potential unfairness to patients. This paper’s review of AI innovations finds that regulations lag behind AI developments, leading to unsettled issues regarding legal responsibility for AI and user control over AI-generated results and findings in legal proceedings. Some of the future avenues that are presented in the study are the future of XAI for legal purposes, utilizing federated learning for resolving privacy issues, and the need to foster adaptive regulation. Finally, the review advocates for Legal Subject Matter Experts to collaborate with legal informatics experts, ethicists, and policy makers to develop the best solutions to implement AI in medical negligence claims. It reasons that there is great potential for AI to have a deep impact on the practice of law but when done, it must do so in a way that respects justice and on the Rights of Individuals.
摘要In the process of building a new power system dominated by new energy sources,power storage is a key supporting technology that ensures the safe and stable operation of the power grid,enables the flexible regulation of the system,and raises the level of new energy consumption.It is also key to achieving carbon peak and neutrality as well as energy transformation.
摘要This video series is the first experimental psychology documentary made in China.It focuses on analyzing professional theories to raise people’s general understanding of basic psychology.By combining innovative audiovisual narrative with psychological experiments,it zooms in on real human nature through discussing social hotspots from the perspectives of social psychology,cognitive psychology,and personality psychology,in order to help people find answers for their current psychological difficulties.
摘要What Are You Up To Today?Chief Director:Wu Zijuan Length:12 Episodes Producer:bilibili Broadcasting Platform:bilibili Produced by China’s YouTube-like video sharing platform bilibili,the film is a series of short documentaries presenting people's daily life in different jobs.It follows 12 individuals in their respective jobs and trades that keep society functioning.By focusing on their daily lives,the documentary films capture the hustle and bustle of the days that make up a hopeful life.
基金supported by the Tightly Integrated Health Consortium Research Project(Grant ynlglht202412)。
摘要In nursing practice,electronic nursing records(ENRs)are an important component of patient care documents,but they also significantly increase administrative burdens.With the development of artificial intelligence technology,it has become possible to use large text models to assist in generating nursing documents.This article explores the application of generative AI in nursing documentation.Research has shown that the application of generative AI in nursing documents demonstrates significant potential,but also faces challenges in terms of quality and implementation.In terms of efficiency,AI assisted document tools can significantly reduce the administrative burden on nurses by reallocating time to direct patient care.Studies have shown that they can reduce document time by 21-30%.However,there are variables in the quality of AI generated records,and the content is often described as'textbook style',lacking patient specific details and appropriate medical terminology.Successful implementation relies on a specialized framework that includes strong stakeholder engagement and adaptation to nursing specific workflows and regulatory standards.The conclusion points out that current AI systems are most suitable for assisting in drafting nursing documents,and clinical validation remains crucial for patient safety and document integrity.
摘要This paper highlights the critical role of medical device design and development documents within the quality system,including their compliance with regulatory standards,their function as a traceable record,their support for all stages,and their use in risk and change management.It also covers document template creation,review record association,information management,adverse event traceability,and the reconciliation of differences in international declarations.
基金This work was supported by the UK Engineering and Physical Sciences Research Council(EPSRC)(No.GR/R67507/01).
摘要Engineers often need to look for the right pieces of information by sifting through long engineering documents, It is a very tiring and time-consuming job. To address this issue, researchers are increasingly devoting their attention to new ways to help information users, including engineers, to access and retrieve document content. The research reported in this paper explores how to use the key technologies of document decomposition (study of document structure), document mark-up (with EXtensible Mark- up Language (XML), HyperText Mark-up Language (HTML), and Scalable Vector Graphics (SVG)), and a facetted classification mechanism. Document content extraction is implemented via computer programming (with Java). An Engineering Document Content Management System (EDCMS) developed in this research demonstrates that as information providers we can make document content in a more accessible manner for information users including engineers.The main features of the EDCMS system are: 1) EDCMS is a system that enables users, especially engineers, to access and retrieve information at content rather than document level. In other words, it provides the right pieces of information that answer specific questions so that engineers don't need to waste time sifting through the whole document to obtain the required piece of information. 2) Users can use the EDCMS via both the data and metadata of a document to access engineering document content. 3) Users can use the EDCMS to access and retrieve content objects, i.e. text, images and graphics (including engineering drawings) via multiple views and at different granularities based on decomposition schemes. Experiments with the EDCMS have been conducted on semi-structured documents, a textbook of CADCAM, and a set of project posters in the Engineering Design domain. Experimental results show that the system provides information users with a powerful solution to access document content.
基金funded by the Ministry of Higher Education,Malaysia for providing facilities and financial support under the Long Research Grant Scheme LRGS-1-2019-UKM-UKM-2-7.
摘要Achieving a good recognition rate for degraded document images is difficult as degraded document images suffer from low contrast,bleedthrough,and nonuniform illumination effects.Unlike the existing baseline thresholding techniques that use fixed thresholds and windows,the proposed method introduces a concept for obtaining dynamic windows according to the image content to achieve better binarization.To enhance a low-contrast image,we proposed a new mean histogram stretching method for suppressing noisy pixels in the background and,simultaneously,increasing pixel contrast at edges or near edges,which results in an enhanced image.For the enhanced image,we propose a new method for deriving adaptive local thresholds for dynamic windows.The dynamic window is derived by exploiting the advantage of Otsu thresholding.To assess the performance of the proposed method,we have used standard databases,namely,document image binarization contest(DIBCO),for experimentation.The comparative study on well-known existing methods indicates that the proposed method outperforms the existing methods in terms of quality and recognition rate.
摘要We are delighted to have been invited to be guest editors for this special issue on forensic document examination,a forensic science discipline that has had a long history of use in both criminal and civil investigations.The work of the forensic document examiner(FDE)can encompass handwriting(including signature)comparisons,and examinations and evaluations of physical components of questioned documents such as paper,ink,toner,and impressions(of writing and stamps)to answer questions of authenticity or source.
基金This research was supported and funded by KAU Scientific Endowment,King Abdulaziz University,Jeddah,Saudi Arabia.
摘要A document layout can be more informative than merely a document’s visual and structural appearance.Thus,document layout analysis(DLA)is considered a necessary prerequisite for advanced processing and detailed document image analysis to be further used in several applications and different objectives.This research extends the traditional approaches of DLA and introduces the concept of semantic document layout analysis(SDLA)by proposing a novel framework for semantic layout analysis and characterization of handwritten manuscripts.The proposed SDLA approach enables the derivation of implicit information and semantic characteristics,which can be effectively utilized in dozens of practical applications for various purposes,in a way bridging the semantic gap and providingmore understandable high-level document image analysis and more invariant characterization via absolute and relative labeling.This approach is validated and evaluated on a large dataset ofArabic handwrittenmanuscripts comprising complex layouts.The experimental work shows promising results in terms of accurate and effective semantic characteristic-based clustering and retrieval of handwritten manuscripts.It also indicates the expected efficacy of using the capabilities of the proposed approach in automating and facilitating many functional,reallife tasks such as effort estimation and pricing of transcription or typing of such complex manuscripts.
基金Project(61806107)supported by the National Natural Science Foundation of ChinaProject supported by the Shandong Key Laboratory of Wisdom Mine Information Technology,ChinaProject supported by the Opening Project of State Key Laboratory of Digital Publishing Technology,China。
摘要Semantic segmentation is a crucial step for document understanding.In this paper,an NVIDIA Jetson Nano-based platform is applied for implementing semantic segmentation for teaching artificial intelligence concepts and programming.To extract semantic structures from document images,we present an end-to-end dilated convolution network architecture.Dilated convolutions have well-known advantages for extracting multi-scale context information without losing spatial resolution.Our model utilizes dilated convolutions with residual network to represent the image features and predicting pixel labels.The convolution part works as feature extractor to obtain multidimensional and hierarchical image features.The consecutive deconvolution is used for producing full resolution segmentation prediction.The probability of each pixel decides its predefined semantic class label.To understand segmentation granularity,we compare performances at three different levels.From fine grained class to coarse class levels,the proposed dilated convolution network architecture is evaluated on three document datasets.The experimental results have shown that both semantic data distribution imbalance and network depth are import factors that influence the document’s semantic segmentation performances.The research is aimed at offering an education resource for teaching artificial intelligence concepts and techniques.
摘要Background:Contrast enhancement plays an important role in the image processing field.Contrast correction has performed an adjustment on the darkness or brightness of the input image and increases the quality of the image.Objective:This paper proposed a novel method based on statistical data from the local mean and local standard deviation.Method:The proposed method modifies the mean and standard deviation of a neighbourhood at each pixel and divides it into three categories:background,foreground,and problematic(contrast&luminosity)region.Experimental results from both visual and objective aspects show that the proposed method can normalize the contrast variation problem effectively compared to Histogram Equalization(HE),Difference of Gaussian(DoG),and Butterworth Homomorphic Filtering(BHF).Seven(7)types of binarization methods were tested on the corrected image and produced a positive and impressive result.Result:Finally,a comparison in terms of Signal Noise Ratio(SNR),Misclassification Error(ME),F-measure,Peak Signal Noise Ratio(PSNR),Misclassification Penalty Metric(MPM),and Accuracy was calculated.Each binarization method shows an incremented result after applying it onto the corrected image compared to the original image.The SNR result of our proposed image is 9.350 higher than the three(3)other methods.The average increment after five(5)types of evaluation are:(Otsu=41.64%,Local Adaptive=7.05%,Niblack=30.28%,Bernsen=25%,Bradley=3.54%,Nick=1.59%,Gradient-Based=14.6%).Conclusion:The results presented in this paper effectively solve the contrast problem and finally produce better quality images.
基金supported by the Innovation Platform Construction of Qinghai Province(No.2016-ZJ-Y04)the Basic Research Program of Qinghai Province(No.2016-ZJ-740)
摘要Text extraction is an important initial step in digitizing the historical documents. In this paper, we present a text extraction method for historical Tibetan document images based on block projections. The task of text extraction is considered as text area detection and location problem. The images are divided equally into blocks and the blocks are filtered by the information of the categories of connected components and corner point density. By analyzing the filtered blocks' projections, the approximate text areas can be located, and the text regions are extracted. Experiments on the dataset of historical Tibetan documents demonstrate the effectiveness of the proposed method.
基金The National Natural Science Foundation of China(No.60503020,60373066,60403016,60425206),the Natural Science Foundation of Jiangsu Higher Education Institutions ( No.04KJB520096),the Doctoral Foundation of Nanjing University of Posts and Telecommunication (No.0302).
摘要A rough set based corner classification neural network, the Rough-CC4, is presented to solve document classification problems such as document representation of different document sizes, document feature selection and document feature encoding. In the Rough-CC4, the documents are described by the equivalent classes of the approximate words. By this method, the dimensions representing the documents can be reduced, which can solve the precision problems caused by the different document sizes and also blur the differences caused by the approximate words. In the Rough-CC4, a binary encoding method is introduced, through which the importance of documents relative to each equivalent class is encoded. By this encoding method, the precision of the Rough-CC4 is improved greatly and the space complexity of the Rough-CC4 is reduced. The Rough-CC4 can be used in automatic classification of documents.