Handwriting-Based Text Line Segmentation from Malayalam Documents
Optical character recognition systems for Malayalam handwritten documents have become an open research area. A major hindrance in this research is the unavailability of a benchmark database. Therefore, a new database of 402 Malayalam handwritten document
Pearlsy P V, Deepa Sankar
doaj +1 more source
Text line segmentation of historical documents: a survey [PDF]
There is a huge amount of historical documents in libraries and in various National Archives that have not been exploited electronically. Although automatic reading of complete pages remains, in most cases, a long-term objective, tasks such as word spotting, text/image alignment, authentication and extraction of specific fields are in use today.
Laurence Likforman-Sulem +2 more
openaire +2 more sources
Combining Morphological and Histogram based Text Line Segmentation in the OCR Context [PDF]
Text line segmentation is one of the pre-stages of modern optical character recognition systems. The algorithmic approach proposed by this paper has been designed for this exact purpose.
Pit Schneider
doaj +1 more source
End-To-End Deep-Learning-Based Tamil Handwritten Document Recognition and Classification Model
Overview: Handwriting recognition (HR) involves converting handwritten text into machine-readable text. Tamil handwritten document recognition remains a challenging process in various text real-world applications owing to the differences in the sizes ...
C. Vinotheni, S. Lakshmana Pandian
doaj +1 more source
TranSentCut - transformer based Thai sentence segmentation [PDF]
We propose TranSentCut, a sentence segmentation model for Thai based on the transformer architecture. Sentence segmentation for Thai is a problem because there is no end of sentence marker like in other languages. Existing methods make use of POS tags,
Sumeth Yuenyong +1 more
doaj +1 more source
Segmentation and Recognition for Historical Tibetan Document Images
As a shining pearl in traditional Tibetan culture, historical Tibetan documents have received extensive attention from historians, linguists and Buddhist scholars.
Longlong Ma +5 more
doaj +1 more source
You Actually Look Twice At it (YALTAi): using an object detection approach instead of region segmentation within the Kraken engine [PDF]
Layout Analysis (the identification of zones and their classification) is the first step along line segmentation in Optical Character Recognition and similar tasks.
Thibault Clérice
doaj +1 more source
End-to-End Transcript Alignment of 17th Century Manuscripts: The Case of Moccia Code
The growth of digital libraries has yielded a large number of handwritten historical documents in the form of images, often accompanied by a digital transcription of the content. The ability to track the position of the words of the digital transcription
Giuseppe De Gregorio +2 more
doaj +1 more source
AT-Text: Assembling Text Components for Efficient Dense Scene Text Detection
Text detection is a prerequisite for text recognition in scene images. Previous segmentation-based methods for detecting scene text have already achieved a promising performance.
Haiyan Li, Hongtao Lu
doaj +1 more source
Text Line Extraction in Historical Documents Using Mask R-CNN
Text line extraction is an essential preprocessing step in many handwritten document image analysis tasks. It includes detecting text lines in a document image and segmenting the regions of each detected line.
Ahmad Droby +5 more
doaj +1 more source

