Results 121 to 130 of about 34,020 (164)
Some of the next articles are maybe not open access.
Development of an Assamese OCR using Bangla OCR
Proceeding of the workshop on Document Analysis and Recognition, 2012This paper refers to the development of an OCR for the Assamese language by modifying an existing OCR for the Bangla language. This modification is feasible because the Assamese script is similar, except for a few characters, to the Bangla script. The OCR incorporates a two stage recognizer using SVM classifier with no post-processing.
Subhankar Ghosh +3 more
openaire +1 more source
Proceedings of the 29th ACM International Conference on Multimedia, 2021
Text-based visual question answering (TextVQA) requires analyzing both the visual contents and texts in an image to answer a question, which is more practical than general visual question answering (VQA). Existing efforts tend to regard optical character recognition (OCR) as a pre-processing and then combine it with a VQA framework.
Gangyan Zeng +3 more
openaire +1 more source
Text-based visual question answering (TextVQA) requires analyzing both the visual contents and texts in an image to answer a question, which is more practical than general visual question answering (VQA). Existing efforts tend to regard optical character recognition (OCR) as a pre-processing and then combine it with a VQA framework.
Gangyan Zeng +3 more
openaire +1 more source
Proceedings of the IEEE, 1992
It is argued that it is time for a major change of approach to optical character recognition (OCR) research. The traditional approach, focusing on the correct classification of isolated characters, has been exhausted. The demonstration of the superiority of a new classification method under operational conditions requires large experimental facilities ...
openaire +1 more source
It is argued that it is time for a major change of approach to optical character recognition (OCR) research. The traditional approach, focusing on the correct classification of isolated characters, has been exhausted. The demonstration of the superiority of a new classification method under operational conditions requires large experimental facilities ...
openaire +1 more source
OCR performance prediction using cross-OCR alignment
2015 13th International Conference on Document Analysis and Recognition (ICDAR), 2015Since 2006 the national library of France (BnF) has developed many mass digitization projects on its collections. The indexation of digital documents on Gallica (the digital library of the BnF) is done through their textual content obtained thanks to service providers that use Optical Character Recognition software (OCR). The modern technologies of OCR
Ben Salah, Ahmed +3 more
openaire +2 more sources
Journal of the American Society for Information Science, 1992
Optical Character Recognition (OCR) has become a highly demanded information transfer technology in recent years. A problem of current OCR technology is that texts produced by the state-of-the-art OCR software contain an unacceptable frequency of errors.
Wei Sun 0002 +3 more
openaire +1 more source
Optical Character Recognition (OCR) has become a highly demanded information transfer technology in recent years. A problem of current OCR technology is that texts produced by the state-of-the-art OCR software contain an unacceptable frequency of errors.
Wei Sun 0002 +3 more
openaire +1 more source
Proceedings of Sixth International Conference on Document Analysis and Recognition, 2002
Telugu is the language spoken by more than 100 million people of South India. Telugu has a complex orthography with a large number of distinct character shapes (estimated to be of the order of 10,000) composed of simple and compound characters formed from 16 vowels (called achchus) and 36 consonants (called hallus).
Atul Negi +2 more
openaire +1 more source
Telugu is the language spoken by more than 100 million people of South India. Telugu has a complex orthography with a large number of distinct character shapes (estimated to be of the order of 10,000) composed of simple and compound characters formed from 16 vowels (called achchus) and 36 consonants (called hallus).
Atul Negi +2 more
openaire +1 more source
2019 International Conference on Document Analysis and Recognition (ICDAR), 2019
Synthetic data generation for optical character recognition (OCR) promises unlimited training data at zero annotation cost. With enough fonts and seed text, we should be able to generate data to train a model that approaches or exceeds the performance with real annotated data. Unfortunately, this is not always the reality. Unconstrained image settings,
David Etter +3 more
openaire +1 more source
Synthetic data generation for optical character recognition (OCR) promises unlimited training data at zero annotation cost. With enough fonts and seed text, we should be able to generate data to train a model that approaches or exceeds the performance with real annotated data. Unfortunately, this is not always the reality. Unconstrained image settings,
David Etter +3 more
openaire +1 more source
2011 International Conference on Document Analysis and Recognition, 2011
Optical character recognition is carried out using techniques borrowed from statistical machine translation. In particular, the use of multiple simple feature functions in linear combination, along with minimum-error-rate training, integrated decoding, and $N$-gram language modeling is found to be remarkably effective, across several scripts and ...
Dmitriy Genzel +6 more
openaire +1 more source
Optical character recognition is carried out using techniques borrowed from statistical machine translation. In particular, the use of multiple simple feature functions in linear combination, along with minimum-error-rate training, integrated decoding, and $N$-gram language modeling is found to be remarkably effective, across several scripts and ...
Dmitriy Genzel +6 more
openaire +1 more source
Proceedings 15th International Conference on Pattern Recognition. ICPR-2000, 2002
We present a document-specific OCR system and apply it to a corpus of fixed business letters. Unsupervised classification of the segmented character bitmaps on each page, using a "clump" metric, typically yields several hundred clusters with highly skewed populations.
Tin Kam Ho, George Nagy
openaire +1 more source
We present a document-specific OCR system and apply it to a corpus of fixed business letters. Unsupervised classification of the segmented character bitmaps on each page, using a "clump" metric, typically yields several hundred clusters with highly skewed populations.
Tin Kam Ho, George Nagy
openaire +1 more source
Using OpenMP Directives to Accelerate OCR with Tesseract OCR
2021This paper is devoted the methods of speed-up optical character recognition which is used for transformation of the scanned image to the edited text format. The example of application of these methods are the systems of the automated search of fragment of text in the catalogues of electronic libraries, where as an entrance format both the entered
Barkovska, Olesia, Ryzhov, Ihor
openaire +1 more source

