Results 11 to 20 of about 2,520 (188)

The Lemmatization of Copulatives in Northern Sotho *

open access: yesLexikos, 2011
<p>Abstract: For learners of Northern Sotho as a second or even foreign language, the copulative system is probably the most complicated grammatical system to master. The encoding needs of such learners, i.e.
D.J. Prinsloo
doaj   +4 more sources

Universal Lemmatizer: A sequence-to-sequence model for lemmatizing Universal Dependencies treebanks [PDF]

open access: yesNatural Language Engineering, 2020
AbstractIn this paper, we present a novel lemmatization method based on a sequence-to-sequence neural network architecture and morphosyntactic context representation. In the proposed method, our context-sensitive lemmatizer generates the lemma one character at a time based on the surface form characters and its morphosyntactic features obtained from a ...
Jenna Kanerva   +2 more
openaire   +2 more sources

Hybrid lemmatization in HuSpaCy

open access: yesCoRR, 2023
Lemmatization is still not a trivial task for morphologically rich languages. Previous studies showed that hybrid architectures usually work better for these languages and can yield great results. This paper presents a hybrid lemmatizer utilizing both a neural model, dictionaries and hand-crafted rules.
Péter Berkecz   +4 more
openaire   +2 more sources

Automatic Lemmatizer Construction with Focus on OOV Words Lemmatization [PDF]

open access: yes, 2005
This paper deals with the automatic construction of a lemmatizer from a Full Form – Lemma (FFL) training dictionary and with lemmatization of new, in the FFL dictionary unseen, i.e. out-of-vocabulary (OOV) words. Three methods of lemmatization of three kinds of OOV words (missing full forms, unknown words, and compound words) are introduced.
Jakub Kanis, Ludek Müller
openaire   +2 more sources

Neural Lemmatization of Multiword Expressions [PDF]

open access: yesProceedings of the Joint Workshop on Multiword Expressions and WordNet (MWE-WN 2019), 2019
This article focuses on the lemmatization of multiword expressions (MWEs). We propose a deep encoder-decoder architecture generating for every MWE word its corresponding part in the lemma, based on the internal context of the MWE. The encoder relies on recurrent networks based on (1) the character sequence of the individual words to capture their ...
Marine Schmitt, Mathieu Constant
openaire   +2 more sources

Translating Speech to Indian Sign Language Using Natural Language Processing

open access: yesFuture Internet, 2022
Language plays a vital role in the communication of ideas, thoughts, and information to others. Hearing-impaired people also understand our thoughts using a language known as sign language.
Purushottam Sharma   +4 more
doaj   +1 more source

A novel Arabic lemmatization algorithm [PDF]

open access: yesProceedings of the second workshop on Analytics for noisy unstructured text data, 2008
Tokenization is a fundamental step in processing textual data preceding the tasks of information retrieval, text mining, and natural language processing. Tokenization is a language-dependent approach, including normalization, stop words removal, lemmatization and stemming.Both stemming and lemmatization share a common goal of reducing a word to its ...
Eiman Tamah Al-Shammari   +1 more
openaire   +1 more source

Lemmatization of Polish person names [PDF]

open access: yesProceedings of the Workshop on Balto-Slavonic Natural Language Processing Information Extraction and Enabling Technologies - ACL '07, 2007
The paper presents two techniques for lemmatization of Polish person names. First, we apply a rule-based approach which relies on linguistic information and heuristics. Then, we investigate an alternative knowledge-poor method which employs string distance measures. We provide an evaluation of the adopted techniques using a set of newspaper texts.
Piskorski, Jakub   +2 more
openaire   +2 more sources

Morphological Tagging and Lemmatization in the Albanian Language

open access: yesSEEU Review, 2021
An important element of Natural Language Processing is parts of speech tagging. With fine-grained word-class annotations, the word forms in a text can be enhanced and can also be used in downstream processes, such as dependency parsing.
Mati Diellza Nagavci   +2 more
doaj   +1 more source

Developing Core Technologies for Resource-Scarce Nguni Languages

open access: yesInformation, 2021
The creation of linguistic resources is crucial to the continued growth of research and development efforts in the field of natural language processing, especially for resource-scarce languages.
Jakobus S. du Toit, Martin J. Puttkammer
doaj   +1 more source

Home - About - Disclaimer - Privacy