Results 31 to 40 of about 34,020 (164)
OCR-IDL: OCR Annotations for Industry Document Library Dataset
Pretraining has proven successful in Document Intelligence tasks where deluge of documents are used to pretrain the models only later to be finetuned on downstream tasks. One of the problems of the pretraining approaches is the inconsistent usage of pretraining data with different OCR engines leading to incomparable results between models.
Ali Furkan Biten +4 more
openaire +2 more sources
The Page-Building a Pennsylvania German Thesaurus through the Correction of OCR Errors
The aim of our project is to build an online Thesaurus of Pennsylvania German. This North American minority language has been in contact with American English ever since its emergence around 1800 and is still lacking a received standard, producing thus ...
Camilla Balsamo, Barbara Hans-Bianchi
doaj +1 more source
Adapting the Tesseract open source OCR engine for multilingual OCR [PDF]
We describe efforts to adapt the Tesseract open source OCR engine for multiple scripts and languages. Effort has been concentrated on enabling generic multi-lingual operation such that negligible customization is required for a new language beyond providing a corpus of text.
Raymond Smith +2 more
openaire +1 more source
Recognition of Machine-Readable Zone in Identity Documents: A Review
Encoding personal identity document information into machine-readable formats is one of the most important approaches to automatic data processing. The Machine-Readable Zone (MRZ) of a document contains text data in the form of 2 or 3 long text lines of ...
Alexander V. Gayer +2 more
doaj +1 more source
Editors' preface to the monographic section of Iperstoria 13, titled "Digital Humanities: a cross-disciplinary approach to literature, language and education."
Sonia di Loreto, Annarita Taronna
doaj +1 more source
Aim: In oncology, thermal therapy is the application of external heat to fight cancer cells. The goal of whole-body thermal treatment (WBTT) is to raise the patient’s core temperature to 39–42 °C, and represents the only thermal treatment modality that ...
B. Wylleman +11 more
doaj +1 more source
Automatic Transcription of Organ Tablature Music Notation with Deep Neural Networks
Organ tablature music notation differs considerably in structure and form from the music notation used today. The manual transcription of organ tablature compositions to modern music notation is time-consuming and often prone to errors. In this paper, we
Daniel Schneider +4 more
doaj +1 more source
IIR-BinNet: An Ultra-Lightweight Neural Network for Document Image Binarization Using IIR Filters
Document image binarization task requires neural networks that can effectively analyze large image areas to separate text from complex backgrounds. The traditional approach of expanding the receptive field by increasing the depth of the network results ...
Daria Ershova +2 more
doaj +1 more source
Upcycle Your OCR: Reusing OCRs for Post-OCR Text Correction in Romanised Sanskrit
This paper has been accepted as a full paper in the SIGNLL Conference on Computational Natural Language Learning (CoNLL), 2018.
Amrith Krishna +3 more
openaire +2 more sources
OCR based slide retrieval [PDF]
This paper addresses the problem of acquiring, indexing and retrieving slides in the context of automatic oral presentation processing. Since the most suitable acquisition technique, in such a context, is the use of a framegrabber (a device capturing as images the slides displayed on a screen), the slides must be transcribed with an optical character ...
N. Daddaoua +2 more
openaire +2 more sources

