The influence of Tibetan characters on the visual recognition effects of Tibetan-Chinese bilingual guide signs based on drivers visual characteristics was studied.Four versions of Tibetan-Chinese bilingual guide signs...The influence of Tibetan characters on the visual recognition effects of Tibetan-Chinese bilingual guide signs based on drivers visual characteristics was studied.Four versions of Tibetan-Chinese bilingual guide signs with different heights and aspect ratios of Tibetan characters were designed,and corresponding road simulation models were established.10 Tibetan drivers and 10 Han drivers were selected to conduct driving simulation experiments using a driving simulator and eye tracker.The resultant data of the participant s pupil diameter and the visual recognition duration obtained from the eye tracker system were analyzed by analysis of variance.Combining results from the statistical analysis of driving simulator data and the questionnaire results on the visual recognition experience,it can be concluded that for Tibetan drivers,when the height of Tibetan characters was 2/3 of the height of Chinese characters,the visual recognition effect of the signs was better than that of 1/3 and 1/2 of the height of Chinese characters,indicating that increasing the height of Tibetan characters was conducive to improving the visual recognition effect of guide signs.The aspect ratio form of Tibetan had no significant effect on the level of difficulty encountered in drivers visual recognition,but it would affect the aesthetics of the bilingual guide signs.The recommended character height in Tibetan should be increased to improve the visual recognition process for Tibetan drivers.展开更多
AIM: To quantitatively evaluate the effect of a simulated smog environment on human visual function by psychophysical methods.METHODS: The smog environment was simulated in a 40×40×60 cm3 glass chamber fil...AIM: To quantitatively evaluate the effect of a simulated smog environment on human visual function by psychophysical methods.METHODS: The smog environment was simulated in a 40×40×60 cm3 glass chamber filled with a PM2.5 aerosol, and 14 subjects with normal visual function were examined by psychophysical methods with the foggy smog box placed in front of their eyes. The transmission of light through the smog box, an indication of the percentage concentration of smog, was determined with a luminance meter. Visual function under different smog concentrations was evaluated by the E-visual acuity, crowded E-visual acuity and contrast sensitivity.RESULTS: E-visual acuity, crowded E-visual acuity and contrast sensitivity were all impaired with a decrease in the transmission rate(TR) according to power functions, with invariable exponents of-1.41,-1.62 and-0.7, respectively, and R2 values of 0.99 for E and crowded E-visual acuity, 0.96 for contrast sensitivity. Crowded E-visual acuity decreased faster than E-visual acuity. There was a good correlation between the TR, extinction coefficient and visibility under heavy-smog conditions.CONCLUSION: Increases in smog concentration have a strong effect on visual function.展开更多
The methods of visual recognition,positioning and orienting with simple 3 D geometric workpieces are presented in this paper.The principle and operating process of multiple orientation run length coding based on gener...The methods of visual recognition,positioning and orienting with simple 3 D geometric workpieces are presented in this paper.The principle and operating process of multiple orientation run length coding based on general orientation run length coding and visual recognition method are described elaborately.The method of positioning and orientating based on the moment of inertia of the workpiece binary image is stated also.It has been applied in a research on flexible automatic coordinate measuring system formed by integrating computer aided design,computer vision and computer aided inspection planning,with a coordinate measuring machine.The results show that integrating computer vision with measurement system is a feasible and effective approach to improve their flexibility and automation.展开更多
A series of novel six-coordinated terpyridine zinc complexes,containing ammonium salts and thymine fragment at the two terminals,have been designed and synthesized,which can function as highly sensitive visualized sen...A series of novel six-coordinated terpyridine zinc complexes,containing ammonium salts and thymine fragment at the two terminals,have been designed and synthesized,which can function as highly sensitive visualized sensors for melamine detection via selective metallo-hydrogel formation.After fully characterization by various techniques,the complementary triple-hydrogen-bonding between the thymine fragment and melamine,as well as π-π stacking interactions may be responsible for the selective metallo-hydrogel formation.In light of the possible interference aroused by milk ingredients(proteins,peptides and amino acids) and legal/illegal additives(urine,sugars and vitamins),a series of control experiments are therefore involved.To our delight,this visual recognition is highly selective,no gelation was observed with the selected milk ingredients or additives.Remarkably,this new developed protocol enables convenient and highly selective visual recognition of melamine at a concentration as low as 10 ppm in raw milk without any tedious pretreatment.展开更多
Research on intelligent and robotic excavator has become a focus both at home and abroad, and this type of excavator becomes more and more important in application. In this paper, we developed a control system which c...Research on intelligent and robotic excavator has become a focus both at home and abroad, and this type of excavator becomes more and more important in application. In this paper, we developed a control system which can make the intelligent robotic excavator perform excavating operation autonomously. It can recognize the excavating targets by itself, program the operation automatically based on the original parameter, and finish all the tasks. Experimental results indicate the validity in real-time performance and precision of the control system. The intelligent robotic excavator can remarkably ease the labor intensity and enhance the working efficiency.展开更多
Visual speech recognition(VSR)aims to infer spoken content from visual observations of articulatory movements.Despite significant progress,it remains a challenging task in computer vision and speech processing.Its dif...Visual speech recognition(VSR)aims to infer spoken content from visual observations of articulatory movements.Despite significant progress,it remains a challenging task in computer vision and speech processing.Its difficulty arises from pronounced speaker-to-speaker variability,the presence of homophenes(phonemes that are visually indistinguishable),changes in illumination,and the intrinsically high-dimensional nature of spatiotemporal lip dynamics.In this work,we propose NestLipGNN,a graph-based framework that integrates Graph Neural Networks(GNNs)with a nested multi-granularity learning strategy for visual speech recognition.We construct dynamic lip graphs from facial landmarks to model both spatial relationships between lip regions and their temporal motion during speech articulation.The proposed nested learning architecture supports hierarchical feature extraction across several levels of linguistic abstraction,spanning phoneme-level articulatory units,viseme-level visual speech categories,and word-level semantic representations.We further introduce a Temporal Graph Attention mechanism(T-GAT)that adaptively reweights the importance of distinct lip regions over time.We also introduce a graph-based contrastive learning objective to improve the discrimination of visually similar speech patterns,directly confronting the challenge of homophene resolution.Experiments on the LRW,LRS2,LRS3,and GRID datasets show that NestLipGNN improves recognition accuracy compared with existing methods,obtaining 92.3%word-level accuracy on LRW and delivering a 2.1%absolute performance gain over prior methods.Comprehensive ablation analyses confirm the contribution of each architectural component.展开更多
How to create the scenery is the key issue in ancient towns.In this study,50 photos were collected and distributed through the Internet.First,456 online questionnaires with 25,080 data were got.Respondents'favorit...How to create the scenery is the key issue in ancient towns.In this study,50 photos were collected and distributed through the Internet.First,456 online questionnaires with 25,080 data were got.Respondents'favoritism was affected by gender,age,region,profession,and education.Second,SAM computer model was applied to image recognition of Wuzhen style photos,analyzing their visual elements.Third,SPSS software was used to analyze the correlation between subjective beauty degree score and objective landscape elements.Based on the coupled quantitative analysis of AI visual recognition and beauty degree score,it is found that the landscape elements that tourists cared most about are:water bodies,ancient buildings and boats.The proportions of the best landscape elements for the spatial sense of the ancient town are the sky ranged from 26.4%to 38.2%,water body ranged from 19.7%to 34.3%,and buildings ranged from 10.4%to 38.2%.This study reveals the pattern of different types of tourists'evaluation of the landscape to summarize the landscape construction strategy of ancient towns in Jiangnan accordingly.The results are not only beneflt to the cultural tourism of Wuzhen,but can also be applied to many ancient towns in Jiangnan.展开更多
Visual Place Recognition(VPR)technology aims to use visual information to judge the location of agents,which plays an irreplaceable role in tasks such as loop closure detection and relocation.It is well known that pre...Visual Place Recognition(VPR)technology aims to use visual information to judge the location of agents,which plays an irreplaceable role in tasks such as loop closure detection and relocation.It is well known that previous VPR algorithms emphasize the extraction and integration of general image features,while ignoring the mining of salient features that play a key role in the discrimination of VPR tasks.To this end,this paper proposes a Domain-invariant Information Extraction and Optimization Network(DIEONet)for VPR.The core of the algorithm is a newly designed Domain-invariant Information Mining Module(DIMM)and a Multi-sample Joint Triplet Loss(MJT Loss).Specifically,DIMM incorporates the interdependence between different spatial regions of the feature map in the cascaded convolutional unit group,which enhances the model’s attention to the domain-invariant static object class.MJT Loss introduces the“joint processing of multiple samples”mechanism into the original triplet loss,and adds a new distance constraint term for“positive and negative”samples,so that the model can avoid falling into local optimum during training.We demonstrate the effectiveness of our algorithm by conducting extensive experiments on several authoritative benchmarks.In particular,the proposed method achieves the best performance on the TokyoTM dataset with a Recall@1 metric of 92.89%.展开更多
In the Visual Place Recognition(VPR)task,existing research has leveraged large-scale pre-trained models to improve the performance of place recognition.However,when there are significant environmental differences betw...In the Visual Place Recognition(VPR)task,existing research has leveraged large-scale pre-trained models to improve the performance of place recognition.However,when there are significant environmental differences between query images and reference images,a large number of ineffective local features will interfere with the extraction of key landmark features,leading to the retrieval of visually similar but geographically different images.To address this perceptual aliasing problem caused by environmental condition changes,we propose a novel Visual Place Recognition method with Cross-Environment Robust Feature Enhancement(CerfeVPR).This method uses the GAN network to generate similar images of the original images under different environmental conditions,thereby enhancing the learning of robust features of the original images.This enables the global descriptor to effectively ignore appearance changes caused by environmental factors such as seasons and lighting,showing better place recognition accuracy than other methods.Meanwhile,we introduce a large kernel convolution adapter to fine tune the pre-trained model,obtaining a better image feature representation for subsequent robust feature learning.Then,we process the information of different local regions in the general features through a 3-layer pyramid scene parsing network and fuse it with a tag that retains global information to construct a multi-dimensional image feature representation.Based on this,we use the fused features of similar images to drive the robust feature learning of the original images and complete the feature matching between query images and retrieved images.Experiments on multiple commonly used datasets show that our method exhibits excellent performance.On average,CerfeVPR achieves the highest results,with all Recall@N values exceeding 90%.In particular,on the highly challenging Nordland dataset,the R@1 metric is improved by 4.6%,significantly outperforming other methods,which fully verifies the superiority of CerfeVPR in visual place recognition under complex environments.展开更多
Visual speech recognition is a central problem in computer vision,encompassing both lip reading(visual speech recognition)and sign language recognition.Although substantial progress has been achieved independently on ...Visual speech recognition is a central problem in computer vision,encompassing both lip reading(visual speech recognition)and sign language recognition.Although substantial progress has been achieved independently on each task,their complementary characteristics have rarely been explored jointly.In this work we propose UniModal-LSR(Unified Multimodal Lip and Sign Recognition),a novel deep learning framework that jointly addresses lip reading and sign language recognition within a single multimodal architecture.By exploiting shared properties of visual communication channels,namely temporal dynamics,spatial articulation structure,and contextual dependencies,the proposed model enables bidirectional transfer of knowledge between modalities.The framework incorporates a Hierarchical Temporal-Spatial Encoder that captures multi-scale temporal patterns through the combination of local convolutions and global self-attention.It also includes a Cross-Modal Attention Fusion module that performs dynamic,context-aware information exchange via bidirectional cross-attention and adaptive gating.Additionally,a Contrastive Semantic Alignment loss enforces semantic consistency across modality-specific representations.Overall,the architecture integrates three-dimensional convolutional neural networks for spatiotemporal feature extraction with graph neural networks for explicit hand-pose modeling.Extensive experiments on several public benchmarks show that UniModal-LSR improves performance compared with recent methods.The model attains a Word Error Rate(WER)of 33.2%on LRS2-BBC,representing a 12.4%relative gain.On PHOENIX-2014,it achieves 18.3%WER,a 13.7%relative gain.Moreover,the unified model reduces parameter count by 25.9%relative to two separate task-specific systems.These results indicate that unified multimodal modeling can improve visual speech recognition performance and may support future communication technologies.展开更多
Flatness pattern recognition is the key of the flatness control. The accuracy of the present flatness pattern recognition is limited and the shape defects cannot be reflected intuitively. In order to improve it, a nov...Flatness pattern recognition is the key of the flatness control. The accuracy of the present flatness pattern recognition is limited and the shape defects cannot be reflected intuitively. In order to improve it, a novel method via T-S cloud inference network optimized by genetic algorithm(GA) is proposed. T-S cloud inference network is constructed with T-S fuzzy neural network and the cloud model. So, the rapid of fuzzy logic and the uncertainty of cloud model for processing data are both taken into account. What's more, GA possesses good parallel design structure and global optimization characteristics. Compared with the simulation recognition results of traditional BP Algorithm, GA is more accurate and effective. Moreover, virtual reality technology is introduced into the field of shape control by Lab VIEW, MATLAB mixed programming. And virtual flatness pattern recognition interface is designed.Therefore, the data of engineering analysis and the actual model are combined with each other, and the shape defects could be seen more lively and intuitively.展开更多
Visual recognition is currently one of the most important and active research areas in computer vision,pattern recognition,and even the general field of artificial intelligence.It has great fundamental importance and ...Visual recognition is currently one of the most important and active research areas in computer vision,pattern recognition,and even the general field of artificial intelligence.It has great fundamental importance and strong industrial needs,particularly the modern deep neural networks(DNNs)and some brain-inspired methodologies,have largely boosted the recognition performance on many concrete tasks,with the help of large amounts of training data and new powerful computation resources.Although recognition accuracy is usually the first concern for new progresses,efficiency is actually rather important and sometimes critical for both academic research and industrial applications.Moreover,insightful views on the opportunities and challenges of efficiency are also highly required for the entire community.While general surveys on the efficiency issue have been done from various perspectives,as far as we are aware,scarcely any of them focused on visual recognition systematically,and thus it is unclear which progresses are applicable to it and what else should be concerned.In this survey,we present the review of recent advances with our suggestions on the new possible directions towards improving the efficiency of DNN-related and brain-inspired visual recognition approaches,including efficient network compression and dynamic brain-inspired networks.We investigate not only from the model but also from the data point of view(which is not the case in existing surveys)and focus on four typical data types(images,video,points,and events).This survey attempts to provide a systematic summary via a comprehensive survey that can serve as a valuable reference and inspire both researchers and practitioners working on visual recognition problems.展开更多
Lip-reading technologies are rapidly progressing following the breakthrough of deep learning.It plays a vital role in its many applications,such as:human-machine communication practices or security applications.In thi...Lip-reading technologies are rapidly progressing following the breakthrough of deep learning.It plays a vital role in its many applications,such as:human-machine communication practices or security applications.In this paper,we propose to develop an effective lip-reading recognition model for Arabic visual speech recognition by implementing deep learning algorithms.The Arabic visual datasets that have been collected contains 2400 records of Arabic digits and 960 records of Arabic phrases from 24 native speakers.The primary purpose is to provide a high-performance model in terms of enhancing the preprocessing phase.Firstly,we extract keyframes from our dataset.Secondly,we produce a Concatenated Frame Images(CFIs)that represent the utterance sequence in one single image.Finally,the VGG-19 is employed for visual features extraction in our proposed model.We have examined different keyframes:10,15,and 20 for comparing two types of approaches in the proposed model:(1)the VGG-19 base model and(2)VGG-19 base model with batch normalization.The results show that the second approach achieves greater accuracy:94%for digit recognition,97%for phrase recognition,and 93%for digits and phrases recognition in the test dataset.Therefore,our proposed model is superior to models based on CFIs input.展开更多
A two-stage algorithm based on deep learning for the detection and recognition of can bottom spray codes and numbers is proposed to address the problems of small character areas and fast production line speeds in can ...A two-stage algorithm based on deep learning for the detection and recognition of can bottom spray codes and numbers is proposed to address the problems of small character areas and fast production line speeds in can bottom spray code number recognition.In the coding number detection stage,Differentiable Binarization Network is used as the backbone network,combined with the Attention and Dilation Convolutions Path Aggregation Network feature fusion structure to enhance the model detection effect.In terms of text recognition,using the Scene Visual Text Recognition coding number recognition network for end-to-end training can alleviate the problem of coding recognition errors caused by image color distortion due to variations in lighting and background noise.In addition,model pruning and quantization are used to reduce the number ofmodel parameters to meet deployment requirements in resource-constrained environments.A comparative experiment was conducted using the dataset of tank bottom spray code numbers collected on-site,and a transfer experiment was conducted using the dataset of packaging box production date.The experimental results show that the algorithm proposed in this study can effectively locate the coding of cans at different positions on the roller conveyor,and can accurately identify the coding numbers at high production line speeds.The Hmean value of the coding number detection is 97.32%,and the accuracy of the coding number recognition is 98.21%.This verifies that the algorithm proposed in this paper has high accuracy in coding number detection and recognition.展开更多
Brood parasitic birds lay eggs in the nests of other birds,and the parasitized hosts can reduce the cost of raising unrelated offspring through the recognition of parasitic eggs.Hosts can adopt vision-based cognitive ...Brood parasitic birds lay eggs in the nests of other birds,and the parasitized hosts can reduce the cost of raising unrelated offspring through the recognition of parasitic eggs.Hosts can adopt vision-based cognitive mechanisms to recognize foreign eggs by comparing the colors of foreign and host eggs.However,there is currently no uniform conclusion as to whether this comparison involves the single or multiple threshold decision rules.In this study,we tested both hypotheses by adding model eggs of different colors to the nests of Barn Swallows(Hirundo rustica)of two geographical populations breeding in Hainan and Heilongjiang Provinces in China.Results showed that Barn Swallows rejected more white model eggs(moderate mimetic to their own eggs)and blue model eggs(highly non-mimetic eggs with shorter reflectance spectrum)than red model eggs(highly nonmimetic eggs with longer reflectance spectrum).There was no difference in the rejection rate of model eggs between the two populations of Barn Swallows,and clutch size was not a factor affecting egg recognition.Our results are consistent with the single rejection threshold model.This study provides strong experimental evidence that the color of model eggs can has an important effect on egg recognition in Barn Swallows,opening up new avenues to uncover the evolution of cuckoo egg mimicry and explore the cognitive mechanisms underlying the visual recognition of foreign eggs by hosts.展开更多
Visual Place Recognition(VPR) poses significant challenges due to the need for simultaneous comprehension of macro-level semantic layouts and micro-level discriminative details. Traditional singlescale feature represe...Visual Place Recognition(VPR) poses significant challenges due to the need for simultaneous comprehension of macro-level semantic layouts and micro-level discriminative details. Traditional singlescale feature representations struggle to meet these multi-granularity cognitive demands. To address this,we propose a novel multi-scale feature fusion strategy that effectively integrates high-level semantic context with spatially precise shallow-layer features, significantly enhancing recognition accuracy in structurally similar environments. Additionally, we overcome the computational inefficiency inherent in conventional Vision Transformers(Vi Ts) by introducing a specialized cross-attention mechanism augmented with memory expert modules. Inspired by human visual cognition, these modules selectively attend to key visual landmarks, progressively accumulating and transferring discriminative knowledge across tasks. This approach achieves superior recognition performance while substantially reducing computational complexity.展开更多
In this study, we propose an incremental learning approach based on a machine-machine interaction via relative attribute feedbacks that exploit comparative relationships among top level image categories. One machine a...In this study, we propose an incremental learning approach based on a machine-machine interaction via relative attribute feedbacks that exploit comparative relationships among top level image categories. One machine acts as 'Student (S)' with initially limited information and it endeavors to capture the task domain gradually by questioning its mentor on a pool of unlabeled data. The other machine is 'Teacher (T)' with the implicit knowledge for helping S on learning the class models. T initiates relative attributes as a communication channel by randomly sorting the classes on attribute space in an unsupervised manner. S starts modeling the categories in this intermediate level by using only a limited number of labeled data. Thereafter, it first selects an entropy-based sample from the pool of unlabeled data and triggers the conversation by propagating the selected image with its belief class in a query. Since T already knows the ground truth labels, it not only decides whether the belief is true or false, but it also provides an attribute-based feedback to S in each case without revealing the true label of the query sample if the belief is false. So the number of training data is increased virtually by dropping the falsely predicted sample back into the unlabeled pool. Next, S updates the attribute space which, in fact, has an impact on T's future responses, and then the category models are updated concurrently for the next run. We experience the weakly supervised algorithm on the real world datasets of faces and natural scenes in comparison with direct attribute prediction and semi-supervised learning approaches, and a noteworthy performance increase is achieved.展开更多
The optical character recognition for the right to left and cursive languages such as Arabic is challenging and received little attention from researchers in the past compared to the other Latin languages.Moreover,the...The optical character recognition for the right to left and cursive languages such as Arabic is challenging and received little attention from researchers in the past compared to the other Latin languages.Moreover,the absence of a standard publicly available dataset for several low-resource lan-guages,including the Pashto language remained a hurdle in the advancement of language processing.Realizing that,a clean dataset is the fundamental and core requirement of character recognition,this research begins with dataset generation and aims at a system capable of complete language understanding.Keeping in view the complete and full autonomous recognition of the cursive Pashto script.The first achievement of this research is a clean and standard dataset for the isolated characters of the Pashto script.In this paper,a database of isolated Pashto characters for forty four alphabets using various font styles has been introduced.In order to overcome the font style shortage,the graphical software Inkscape has been used to generate sufficient image data samples for each character.The dataset has been pre-processed and reduced in dimensions to 32×32 pixels,and further converted into the binary format with a black background and white text so that it resembles the Modified National Institute of Standards and Technology(MNIST)database.The benchmark database is publicly available for further research on the standard GitHub and Kaggle database servers both in pixel and Comma Separated Values(CSV)formats.展开更多
The continuing advances in deep learning have paved the way for several challenging ideas.One such idea is visual lip-reading,which has recently drawn many research interests.Lip-reading,often referred to as visual sp...The continuing advances in deep learning have paved the way for several challenging ideas.One such idea is visual lip-reading,which has recently drawn many research interests.Lip-reading,often referred to as visual speech recognition,is the ability to understand and predict spoken speech based solely on lip movements without using sounds.Due to the lack of research studies on visual speech recognition for the Arabic language in general,and its absence in the Quranic research,this research aims to fill this gap.This paper introduces a new publicly available Arabic lip-reading dataset containing 10490 videos captured from multiple viewpoints and comprising data samples at the letter level(i.e.,single letters(single alphabets)and Quranic disjoined letters)and in the word level based on the content and context of the book Al-Qaida Al-Noorania.This research uses visual speech recognition to recognize spoken Arabic letters(Arabic alphabets),Quranic disjoined letters,and Quranic words,mainly phonetic as they are recited in the Holy Quran according to Quranic study aid entitled Al-Qaida Al-Noorania.This study could further validate the correctness of pronunciation and,subsequently,assist people in correctly reciting Quran.Furthermore,a detailed description of the created dataset and its construction methodology is provided.This new dataset is used to train an effective pre-trained deep learning CNN model throughout transfer learning for lip-reading,achieving the accuracies of 83.3%,80.5%,and 77.5%on words,disjoined letters,and single letters,respectively,where an extended analysis of the results is provided.Finally,the experimental outcomes,different research aspects,and dataset collection consistency and challenges are discussed and concluded with several new promising trends for future work.展开更多
This paper presents a vision-based fingertip-writing character recognition system. The overall system is implemented through a CMOS image camera on a FPGA chip. A blue cover is mounted on the top of a finger to simpli...This paper presents a vision-based fingertip-writing character recognition system. The overall system is implemented through a CMOS image camera on a FPGA chip. A blue cover is mounted on the top of a finger to simplify fingertip detection and to enhance recognition accuracy. For each character stroke, 8 sample points (including start and end points) are recorded. 7 tangent angles between consecutive sampled points are also recorded as features. In addition, 3 features angles are extracted: angles of the triangle consisting of the start point, end point and average point of all (8 total) sampled points. According to these key feature angles, a simple template matching K-nearest-neighbor classifier is applied to distinguish each character stroke. Experimental result showed that the system can successfully recognize fingertip-writing character strokes of digits and small lower case letter alphabets with an accuracy of almost 100%. Overall, the proposed finger-tip-writing recognition system provides an easy-to-use and accurate visual character input method.展开更多
基金The National Natural Science Foundation of China(No.51768063,51868068)Shanxi Provincial Innovation Center Project for Digital Road Design Technology(No.202104010911019)。
摘要The influence of Tibetan characters on the visual recognition effects of Tibetan-Chinese bilingual guide signs based on drivers visual characteristics was studied.Four versions of Tibetan-Chinese bilingual guide signs with different heights and aspect ratios of Tibetan characters were designed,and corresponding road simulation models were established.10 Tibetan drivers and 10 Han drivers were selected to conduct driving simulation experiments using a driving simulator and eye tracker.The resultant data of the participant s pupil diameter and the visual recognition duration obtained from the eye tracker system were analyzed by analysis of variance.Combining results from the statistical analysis of driving simulator data and the questionnaire results on the visual recognition experience,it can be concluded that for Tibetan drivers,when the height of Tibetan characters was 2/3 of the height of Chinese characters,the visual recognition effect of the signs was better than that of 1/3 and 1/2 of the height of Chinese characters,indicating that increasing the height of Tibetan characters was conducive to improving the visual recognition effect of guide signs.The aspect ratio form of Tibetan had no significant effect on the level of difficulty encountered in drivers visual recognition,but it would affect the aesthetics of the bilingual guide signs.The recommended character height in Tibetan should be increased to improve the visual recognition process for Tibetan drivers.
基金Supported by National Nature Science Foundation of China (No. 81570880)
摘要AIM: To quantitatively evaluate the effect of a simulated smog environment on human visual function by psychophysical methods.METHODS: The smog environment was simulated in a 40×40×60 cm3 glass chamber filled with a PM2.5 aerosol, and 14 subjects with normal visual function were examined by psychophysical methods with the foggy smog box placed in front of their eyes. The transmission of light through the smog box, an indication of the percentage concentration of smog, was determined with a luminance meter. Visual function under different smog concentrations was evaluated by the E-visual acuity, crowded E-visual acuity and contrast sensitivity.RESULTS: E-visual acuity, crowded E-visual acuity and contrast sensitivity were all impaired with a decrease in the transmission rate(TR) according to power functions, with invariable exponents of-1.41,-1.62 and-0.7, respectively, and R2 values of 0.99 for E and crowded E-visual acuity, 0.96 for contrast sensitivity. Crowded E-visual acuity decreased faster than E-visual acuity. There was a good correlation between the TR, extinction coefficient and visibility under heavy-smog conditions.CONCLUSION: Increases in smog concentration have a strong effect on visual function.
摘要The methods of visual recognition,positioning and orienting with simple 3 D geometric workpieces are presented in this paper.The principle and operating process of multiple orientation run length coding based on general orientation run length coding and visual recognition method are described elaborately.The method of positioning and orientating based on the moment of inertia of the workpiece binary image is stated also.It has been applied in a research on flexible automatic coordinate measuring system formed by integrating computer aided design,computer vision and computer aided inspection planning,with a coordinate measuring machine.The results show that integrating computer vision with measurement system is a feasible and effective approach to improve their flexibility and automation.
基金Financial support from the State General Administration of the People’s Republic of China for Quality Supervision and Inspection and Quarantine (No.2016QK122)Shanghai Institute of Quality Inspection and Technical Research+1 种基金the National Natural Science Foundation of China (Nos.21572036 and 21861132002)the Department of Chemistry,Fudan University
摘要A series of novel six-coordinated terpyridine zinc complexes,containing ammonium salts and thymine fragment at the two terminals,have been designed and synthesized,which can function as highly sensitive visualized sensors for melamine detection via selective metallo-hydrogel formation.After fully characterization by various techniques,the complementary triple-hydrogen-bonding between the thymine fragment and melamine,as well as π-π stacking interactions may be responsible for the selective metallo-hydrogel formation.In light of the possible interference aroused by milk ingredients(proteins,peptides and amino acids) and legal/illegal additives(urine,sugars and vitamins),a series of control experiments are therefore involved.To our delight,this visual recognition is highly selective,no gelation was observed with the selected milk ingredients or additives.Remarkably,this new developed protocol enables convenient and highly selective visual recognition of melamine at a concentration as low as 10 ppm in raw milk without any tedious pretreatment.
摘要Research on intelligent and robotic excavator has become a focus both at home and abroad, and this type of excavator becomes more and more important in application. In this paper, we developed a control system which can make the intelligent robotic excavator perform excavating operation autonomously. It can recognize the excavating targets by itself, program the operation automatically based on the original parameter, and finish all the tasks. Experimental results indicate the validity in real-time performance and precision of the control system. The intelligent robotic excavator can remarkably ease the labor intensity and enhance the working efficiency.
基金funded by Ho Chi Minh City Open University(HCMCOU)the Ministry of Education and Training(Vietnam)under grant number B2025-MBS-01.
摘要Visual speech recognition(VSR)aims to infer spoken content from visual observations of articulatory movements.Despite significant progress,it remains a challenging task in computer vision and speech processing.Its difficulty arises from pronounced speaker-to-speaker variability,the presence of homophenes(phonemes that are visually indistinguishable),changes in illumination,and the intrinsically high-dimensional nature of spatiotemporal lip dynamics.In this work,we propose NestLipGNN,a graph-based framework that integrates Graph Neural Networks(GNNs)with a nested multi-granularity learning strategy for visual speech recognition.We construct dynamic lip graphs from facial landmarks to model both spatial relationships between lip regions and their temporal motion during speech articulation.The proposed nested learning architecture supports hierarchical feature extraction across several levels of linguistic abstraction,spanning phoneme-level articulatory units,viseme-level visual speech categories,and word-level semantic representations.We further introduce a Temporal Graph Attention mechanism(T-GAT)that adaptively reweights the importance of distinct lip regions over time.We also introduce a graph-based contrastive learning objective to improve the discrimination of visually similar speech patterns,directly confronting the challenge of homophene resolution.Experiments on the LRW,LRS2,LRS3,and GRID datasets show that NestLipGNN improves recognition accuracy compared with existing methods,obtaining 92.3%word-level accuracy on LRW and delivering a 2.1%absolute performance gain over prior methods.Comprehensive ablation analyses confirm the contribution of each architectural component.
基金fundings received by Chinese Ministry of Education Humanities and Social Sciences General Project(Grant No.24YJC760190)Zhejiang Province Philosophy and Social Sciences Planning Project(Grant No.25NDJC058YBM)Zhejiang Province Philosophy and Social Sciences Planning“Provincial and Municipal Cooperation”Project(Grant No.24SSHZ119YB)。
摘要How to create the scenery is the key issue in ancient towns.In this study,50 photos were collected and distributed through the Internet.First,456 online questionnaires with 25,080 data were got.Respondents'favoritism was affected by gender,age,region,profession,and education.Second,SAM computer model was applied to image recognition of Wuzhen style photos,analyzing their visual elements.Third,SPSS software was used to analyze the correlation between subjective beauty degree score and objective landscape elements.Based on the coupled quantitative analysis of AI visual recognition and beauty degree score,it is found that the landscape elements that tourists cared most about are:water bodies,ancient buildings and boats.The proportions of the best landscape elements for the spatial sense of the ancient town are the sky ranged from 26.4%to 38.2%,water body ranged from 19.7%to 34.3%,and buildings ranged from 10.4%to 38.2%.This study reveals the pattern of different types of tourists'evaluation of the landscape to summarize the landscape construction strategy of ancient towns in Jiangnan accordingly.The results are not only beneflt to the cultural tourism of Wuzhen,but can also be applied to many ancient towns in Jiangnan.
基金supported by the Natural Science Foundation of Xinjiang Uygur Autonomous Region under grant number 2022D01B186.
摘要Visual Place Recognition(VPR)technology aims to use visual information to judge the location of agents,which plays an irreplaceable role in tasks such as loop closure detection and relocation.It is well known that previous VPR algorithms emphasize the extraction and integration of general image features,while ignoring the mining of salient features that play a key role in the discrimination of VPR tasks.To this end,this paper proposes a Domain-invariant Information Extraction and Optimization Network(DIEONet)for VPR.The core of the algorithm is a newly designed Domain-invariant Information Mining Module(DIMM)and a Multi-sample Joint Triplet Loss(MJT Loss).Specifically,DIMM incorporates the interdependence between different spatial regions of the feature map in the cascaded convolutional unit group,which enhances the model’s attention to the domain-invariant static object class.MJT Loss introduces the“joint processing of multiple samples”mechanism into the original triplet loss,and adds a new distance constraint term for“positive and negative”samples,so that the model can avoid falling into local optimum during training.We demonstrate the effectiveness of our algorithm by conducting extensive experiments on several authoritative benchmarks.In particular,the proposed method achieves the best performance on the TokyoTM dataset with a Recall@1 metric of 92.89%.
基金supported by Postgraduate Scientific Research Innovation Project of Hunan Province CX20230915National Natural Science Foundation of China under Grant 62472440.
摘要In the Visual Place Recognition(VPR)task,existing research has leveraged large-scale pre-trained models to improve the performance of place recognition.However,when there are significant environmental differences between query images and reference images,a large number of ineffective local features will interfere with the extraction of key landmark features,leading to the retrieval of visually similar but geographically different images.To address this perceptual aliasing problem caused by environmental condition changes,we propose a novel Visual Place Recognition method with Cross-Environment Robust Feature Enhancement(CerfeVPR).This method uses the GAN network to generate similar images of the original images under different environmental conditions,thereby enhancing the learning of robust features of the original images.This enables the global descriptor to effectively ignore appearance changes caused by environmental factors such as seasons and lighting,showing better place recognition accuracy than other methods.Meanwhile,we introduce a large kernel convolution adapter to fine tune the pre-trained model,obtaining a better image feature representation for subsequent robust feature learning.Then,we process the information of different local regions in the general features through a 3-layer pyramid scene parsing network and fuse it with a tag that retains global information to construct a multi-dimensional image feature representation.Based on this,we use the fused features of similar images to drive the robust feature learning of the original images and complete the feature matching between query images and retrieved images.Experiments on multiple commonly used datasets show that our method exhibits excellent performance.On average,CerfeVPR achieves the highest results,with all Recall@N values exceeding 90%.In particular,on the highly challenging Nordland dataset,the R@1 metric is improved by 4.6%,significantly outperforming other methods,which fully verifies the superiority of CerfeVPR in visual place recognition under complex environments.
基金funded by Ho Chi Minh City Open University(HCMCOU)the Ministry of Education and Training(Vietnam)under grant number B2025-MBS-01.
摘要Visual speech recognition is a central problem in computer vision,encompassing both lip reading(visual speech recognition)and sign language recognition.Although substantial progress has been achieved independently on each task,their complementary characteristics have rarely been explored jointly.In this work we propose UniModal-LSR(Unified Multimodal Lip and Sign Recognition),a novel deep learning framework that jointly addresses lip reading and sign language recognition within a single multimodal architecture.By exploiting shared properties of visual communication channels,namely temporal dynamics,spatial articulation structure,and contextual dependencies,the proposed model enables bidirectional transfer of knowledge between modalities.The framework incorporates a Hierarchical Temporal-Spatial Encoder that captures multi-scale temporal patterns through the combination of local convolutions and global self-attention.It also includes a Cross-Modal Attention Fusion module that performs dynamic,context-aware information exchange via bidirectional cross-attention and adaptive gating.Additionally,a Contrastive Semantic Alignment loss enforces semantic consistency across modality-specific representations.Overall,the architecture integrates three-dimensional convolutional neural networks for spatiotemporal feature extraction with graph neural networks for explicit hand-pose modeling.Extensive experiments on several public benchmarks show that UniModal-LSR improves performance compared with recent methods.The model attains a Word Error Rate(WER)of 33.2%on LRS2-BBC,representing a 12.4%relative gain.On PHOENIX-2014,it achieves 18.3%WER,a 13.7%relative gain.Moreover,the unified model reduces parameter count by 25.9%relative to two separate task-specific systems.These results indicate that unified multimodal modeling can improve visual speech recognition performance and may support future communication technologies.
基金Project(LJRC013)supported by the University Innovation Team of Hebei Province Leading Talent Cultivation,China
摘要Flatness pattern recognition is the key of the flatness control. The accuracy of the present flatness pattern recognition is limited and the shape defects cannot be reflected intuitively. In order to improve it, a novel method via T-S cloud inference network optimized by genetic algorithm(GA) is proposed. T-S cloud inference network is constructed with T-S fuzzy neural network and the cloud model. So, the rapid of fuzzy logic and the uncertainty of cloud model for processing data are both taken into account. What's more, GA possesses good parallel design structure and global optimization characteristics. Compared with the simulation recognition results of traditional BP Algorithm, GA is more accurate and effective. Moreover, virtual reality technology is introduced into the field of shape control by Lab VIEW, MATLAB mixed programming. And virtual flatness pattern recognition interface is designed.Therefore, the data of engineering analysis and the actual model are combined with each other, and the shape defects could be seen more lively and intuitively.
基金supported by National Key R&D Program of China(No.2018AAA0102600)Beijing Natural Science Foundation,China(No.JQ21015)+1 种基金Beijing Academy of Artificial Intelligence(BAAI),ChinaPengcheng Laboratory,China。
摘要Visual recognition is currently one of the most important and active research areas in computer vision,pattern recognition,and even the general field of artificial intelligence.It has great fundamental importance and strong industrial needs,particularly the modern deep neural networks(DNNs)and some brain-inspired methodologies,have largely boosted the recognition performance on many concrete tasks,with the help of large amounts of training data and new powerful computation resources.Although recognition accuracy is usually the first concern for new progresses,efficiency is actually rather important and sometimes critical for both academic research and industrial applications.Moreover,insightful views on the opportunities and challenges of efficiency are also highly required for the entire community.While general surveys on the efficiency issue have been done from various perspectives,as far as we are aware,scarcely any of them focused on visual recognition systematically,and thus it is unclear which progresses are applicable to it and what else should be concerned.In this survey,we present the review of recent advances with our suggestions on the new possible directions towards improving the efficiency of DNN-related and brain-inspired visual recognition approaches,including efficient network compression and dynamic brain-inspired networks.We investigate not only from the model but also from the data point of view(which is not the case in existing surveys)and focus on four typical data types(images,video,points,and events).This survey attempts to provide a systematic summary via a comprehensive survey that can serve as a valuable reference and inspire both researchers and practitioners working on visual recognition problems.
摘要Lip-reading technologies are rapidly progressing following the breakthrough of deep learning.It plays a vital role in its many applications,such as:human-machine communication practices or security applications.In this paper,we propose to develop an effective lip-reading recognition model for Arabic visual speech recognition by implementing deep learning algorithms.The Arabic visual datasets that have been collected contains 2400 records of Arabic digits and 960 records of Arabic phrases from 24 native speakers.The primary purpose is to provide a high-performance model in terms of enhancing the preprocessing phase.Firstly,we extract keyframes from our dataset.Secondly,we produce a Concatenated Frame Images(CFIs)that represent the utterance sequence in one single image.Finally,the VGG-19 is employed for visual features extraction in our proposed model.We have examined different keyframes:10,15,and 20 for comparing two types of approaches in the proposed model:(1)the VGG-19 base model and(2)VGG-19 base model with batch normalization.The results show that the second approach achieves greater accuracy:94%for digit recognition,97%for phrase recognition,and 93%for digits and phrases recognition in the test dataset.Therefore,our proposed model is superior to models based on CFIs input.
摘要A two-stage algorithm based on deep learning for the detection and recognition of can bottom spray codes and numbers is proposed to address the problems of small character areas and fast production line speeds in can bottom spray code number recognition.In the coding number detection stage,Differentiable Binarization Network is used as the backbone network,combined with the Attention and Dilation Convolutions Path Aggregation Network feature fusion structure to enhance the model detection effect.In terms of text recognition,using the Scene Visual Text Recognition coding number recognition network for end-to-end training can alleviate the problem of coding recognition errors caused by image color distortion due to variations in lighting and background noise.In addition,model pruning and quantization are used to reduce the number ofmodel parameters to meet deployment requirements in resource-constrained environments.A comparative experiment was conducted using the dataset of tank bottom spray code numbers collected on-site,and a transfer experiment was conducted using the dataset of packaging box production date.The experimental results show that the algorithm proposed in this study can effectively locate the coding of cans at different positions on the roller conveyor,and can accurately identify the coding numbers at high production line speeds.The Hmean value of the coding number detection is 97.32%,and the accuracy of the coding number recognition is 98.21%.This verifies that the algorithm proposed in this paper has high accuracy in coding number detection and recognition.
基金supported by the National Natural Science Foundation of China(Nos.31970427 and 32270526 to W.L.)。
摘要Brood parasitic birds lay eggs in the nests of other birds,and the parasitized hosts can reduce the cost of raising unrelated offspring through the recognition of parasitic eggs.Hosts can adopt vision-based cognitive mechanisms to recognize foreign eggs by comparing the colors of foreign and host eggs.However,there is currently no uniform conclusion as to whether this comparison involves the single or multiple threshold decision rules.In this study,we tested both hypotheses by adding model eggs of different colors to the nests of Barn Swallows(Hirundo rustica)of two geographical populations breeding in Hainan and Heilongjiang Provinces in China.Results showed that Barn Swallows rejected more white model eggs(moderate mimetic to their own eggs)and blue model eggs(highly non-mimetic eggs with shorter reflectance spectrum)than red model eggs(highly nonmimetic eggs with longer reflectance spectrum).There was no difference in the rejection rate of model eggs between the two populations of Barn Swallows,and clutch size was not a factor affecting egg recognition.Our results are consistent with the single rejection threshold model.This study provides strong experimental evidence that the color of model eggs can has an important effect on egg recognition in Barn Swallows,opening up new avenues to uncover the evolution of cuckoo egg mimicry and explore the cognitive mechanisms underlying the visual recognition of foreign eggs by hosts.
基金supported by National Key Research and Development Program of China(No.2025YFE0199900)National Natural Science Foundation of China(Nos.62376261,U21A20487)+3 种基金Guangdong Basic and Applied Basic Research Foundation,China(No.2024A1515011754)Guangdong Technology Project(2023TX07Z126)Shenzhen Technology Project,China(JCYJ20220818101206014,KJZD20240903100000001)Yunnan Technology Project,China(No.202305AF150152)
摘要Visual Place Recognition(VPR) poses significant challenges due to the need for simultaneous comprehension of macro-level semantic layouts and micro-level discriminative details. Traditional singlescale feature representations struggle to meet these multi-granularity cognitive demands. To address this,we propose a novel multi-scale feature fusion strategy that effectively integrates high-level semantic context with spatially precise shallow-layer features, significantly enhancing recognition accuracy in structurally similar environments. Additionally, we overcome the computational inefficiency inherent in conventional Vision Transformers(Vi Ts) by introducing a specialized cross-attention mechanism augmented with memory expert modules. Inspired by human visual cognition, these modules selectively attend to key visual landmarks, progressively accumulating and transferring discriminative knowledge across tasks. This approach achieves superior recognition performance while substantially reducing computational complexity.
摘要In this study, we propose an incremental learning approach based on a machine-machine interaction via relative attribute feedbacks that exploit comparative relationships among top level image categories. One machine acts as 'Student (S)' with initially limited information and it endeavors to capture the task domain gradually by questioning its mentor on a pool of unlabeled data. The other machine is 'Teacher (T)' with the implicit knowledge for helping S on learning the class models. T initiates relative attributes as a communication channel by randomly sorting the classes on attribute space in an unsupervised manner. S starts modeling the categories in this intermediate level by using only a limited number of labeled data. Thereafter, it first selects an entropy-based sample from the pool of unlabeled data and triggers the conversation by propagating the selected image with its belief class in a query. Since T already knows the ground truth labels, it not only decides whether the belief is true or false, but it also provides an attribute-based feedback to S in each case without revealing the true label of the query sample if the belief is false. So the number of training data is increased virtually by dropping the falsely predicted sample back into the unlabeled pool. Next, S updates the attribute space which, in fact, has an impact on T's future responses, and then the category models are updated concurrently for the next run. We experience the weakly supervised algorithm on the real world datasets of faces and natural scenes in comparison with direct attribute prediction and semi-supervised learning approaches, and a noteworthy performance increase is achieved.
摘要The optical character recognition for the right to left and cursive languages such as Arabic is challenging and received little attention from researchers in the past compared to the other Latin languages.Moreover,the absence of a standard publicly available dataset for several low-resource lan-guages,including the Pashto language remained a hurdle in the advancement of language processing.Realizing that,a clean dataset is the fundamental and core requirement of character recognition,this research begins with dataset generation and aims at a system capable of complete language understanding.Keeping in view the complete and full autonomous recognition of the cursive Pashto script.The first achievement of this research is a clean and standard dataset for the isolated characters of the Pashto script.In this paper,a database of isolated Pashto characters for forty four alphabets using various font styles has been introduced.In order to overcome the font style shortage,the graphical software Inkscape has been used to generate sufficient image data samples for each character.The dataset has been pre-processed and reduced in dimensions to 32×32 pixels,and further converted into the binary format with a black background and white text so that it resembles the Modified National Institute of Standards and Technology(MNIST)database.The benchmark database is publicly available for further research on the standard GitHub and Kaggle database servers both in pixel and Comma Separated Values(CSV)formats.
基金This research was supported and funded by KAU Scientific Endowment,King Abdulaziz University,Jeddah,Saudi Arabia.
摘要The continuing advances in deep learning have paved the way for several challenging ideas.One such idea is visual lip-reading,which has recently drawn many research interests.Lip-reading,often referred to as visual speech recognition,is the ability to understand and predict spoken speech based solely on lip movements without using sounds.Due to the lack of research studies on visual speech recognition for the Arabic language in general,and its absence in the Quranic research,this research aims to fill this gap.This paper introduces a new publicly available Arabic lip-reading dataset containing 10490 videos captured from multiple viewpoints and comprising data samples at the letter level(i.e.,single letters(single alphabets)and Quranic disjoined letters)and in the word level based on the content and context of the book Al-Qaida Al-Noorania.This research uses visual speech recognition to recognize spoken Arabic letters(Arabic alphabets),Quranic disjoined letters,and Quranic words,mainly phonetic as they are recited in the Holy Quran according to Quranic study aid entitled Al-Qaida Al-Noorania.This study could further validate the correctness of pronunciation and,subsequently,assist people in correctly reciting Quran.Furthermore,a detailed description of the created dataset and its construction methodology is provided.This new dataset is used to train an effective pre-trained deep learning CNN model throughout transfer learning for lip-reading,achieving the accuracies of 83.3%,80.5%,and 77.5%on words,disjoined letters,and single letters,respectively,where an extended analysis of the results is provided.Finally,the experimental outcomes,different research aspects,and dataset collection consistency and challenges are discussed and concluded with several new promising trends for future work.
摘要This paper presents a vision-based fingertip-writing character recognition system. The overall system is implemented through a CMOS image camera on a FPGA chip. A blue cover is mounted on the top of a finger to simplify fingertip detection and to enhance recognition accuracy. For each character stroke, 8 sample points (including start and end points) are recorded. 7 tangent angles between consecutive sampled points are also recorded as features. In addition, 3 features angles are extracted: angles of the triangle consisting of the start point, end point and average point of all (8 total) sampled points. According to these key feature angles, a simple template matching K-nearest-neighbor classifier is applied to distinguish each character stroke. Experimental result showed that the system can successfully recognize fingertip-writing character strokes of digits and small lower case letter alphabets with an accuracy of almost 100%. Overall, the proposed finger-tip-writing recognition system provides an easy-to-use and accurate visual character input method.