Results 21 to 30 of about 2,296,500 (293)
A survey of historical document image datasets [PDF]
This paper presents a systematic literature review of image datasets for document image analysis, focusing on historical documents, such as handwritten manuscripts and early prints.
Konstantina Nikolaidou +3 more
semanticscholar +1 more source
A Comprehensive Study of ImageNet Pre-Training for Historical Document Image Analysis [PDF]
Automatic analysis of scanned historical documents comprises a wide range of image analysis tasks, which are often challenging for machine learning due to a lack of human-annotated learning samples.
Linda Studer +7 more
semanticscholar +1 more source
Cross-Depicted Historical Motif Categorization and Retrieval with Deep Learning
In this paper, we tackle the problem of categorizing and identifying cross-depicted historical motifs using recent deep learning techniques, with aim of developing a content-based image retrieval system. As cross-depiction, we understand the problem that
Vinaychandran Pondenkandath +4 more
doaj +1 more source
DocSynth: A Layout Guided Approach for Controllable Document Image Synthesis [PDF]
Despite significant progress on current state-of-the-art image generation models, synthesis of document images containing multiple and complex object layouts is a challenging task.
Sanket Biswas +3 more
semanticscholar +1 more source
Fast Binarization of Unevenly Illuminated Document Images Based on Background Estimation for Optical Character Recognition Purposes [PDF]
One of the key operations during the image preprocessing step in Optical Character Recognition (OCR) algorithms is image binarization. Although for uniformly illuminated images, obtained typically by atbed scanners, the use of a single global threshold ...
Hubert Michalak, Krzysztof Okarma
doaj +3 more sources
Table detection is a preliminary step in extracting reliable information from tables in scanned document images. We present CasTabDetectoRS, a novel end-to-end trainable table detection framework that operates on Cascade Mask R-CNN, including Recursive ...
Khurram Azeem Hashmi +4 more
doaj +1 more source
Analysis of degraded ancient documents is challenging due to the severity and combination of degradation present in a single image. Ancient documents also suffer from additional noise during the digitalization process, particularly when digitalization is
Fitri Arnia +2 more
doaj +1 more source
This article presents a dataset of hyperspectral images of handwriting samples collected from 54 individuals. The purpose of the presented dataset is to further explore the use of hyperspectral imaging in document image analysis and to benchmark the ...
Ammad Ul Islam +4 more
doaj +1 more source
Automated analysis of images in documents for intelligent document search [PDF]
Authors use images to present a wide variety of important information in documents. For example, two-dimensional (2-D) plots display important data in scientific publications. Often, end-users seek to extract this data and convert it into a machine-processible form so that the data can be analyzed automatically or compared with other existing data ...
Xiaonan Lu +5 more
openaire +1 more source
PHTI: Pashto Handwritten Text Imagebase for Deep Learning Applications
Document Image Analysis (DIA) is one of the research areas of Artificial Intelligence (AI) that converts document images into machine-readable codes. In DIA systems, Optical Character Recognition (OCR) plays a key role in digitizing document images.
Ibrar Hussain +5 more
doaj +1 more source

