Results 151 to 160 of about 2,520 (188)
Towards an Optimal Solution to Lemmatization in Arabic
Abstract Lemmatization—computing the canonical forms of words in running text—is an important component in any NLP system and a key preprocessing step for most applications that rely on natural language understanding. In the case of Arabic, lemmatization is a complex task because of the rich morphology, agglutinative aspects, and lexical ambiguity ...
Gabor Bella +2 more
exaly +4 more sources
Some of the next articles are maybe not open access.
Related searches:
Related searches:
Lemmatization of Inflected Nouns
2021In this chapter, we describe a process of lemmatization of inflected nouns in Bengali as a part of lexical processing. Inflected nouns are used at a very high frequency in Bengali texts. We first collect a large number of inflected nouns from a Bengali corpus and compile a noun database.
Niladri Sekhar Dash, Dash Niladri Sekhar
exaly +2 more sources
A hybrid approach for Arabic lemmatization
International Journal of Speech Technology, 2018We present in this article an Arabic lemmatizer that assigns to each word of an Arabic sentence, a single lemma taking into account the word context. The proposed system comprises two modules. The first one consists in an analysis out of context, based on the morphosyntactic analyser Alkhalil Morpho Sys 2.
Mohamed Boudchiche, Azzeddine Mazroui
openaire +1 more source
Hybrid Lemmatizer for Estonian
2014In this paper, we present a lemmatizer for the Estonian language, which employs a hybrid approach to handle both in- and out-of-vocabulary words. Our method uses only publicly available data and does not require any external tools such as a POS tagger. In the process of experimentation, we achieved the accuracy of 91%.
Alexander Tkachenko +2 more
openaire +1 more source
Automatic lemmatization of Persian words*
Journal of Quantitative Linguistics, 2006Abstract This study presents a rather novel method for suffix and prefix stripping of Persian words. The method presented is a language independent one and mostly relies on a specially arranged corpus composed of a list of roots, word-forms, prefixes, and suffixes which has been manually compiled.
openaire +1 more source
A Lemmatizer Tool for Assamese Language
2019Word Sense Disambiguation (WSD) requires sense tagged corpora. Words in a corpus appear in inflected or morphed forms. Sense tagging can only be done with words in their root or lemmatized forms. Similarly for Part of Speech Tagging (POS), the words in a corpus are required to be available in their root forms.
Arindam Roy +2 more
openaire +1 more source
A Neural Lemmatizer for Bengali
Proceedings of the Language Resources and Evaluation Conference, 2016Abhisek Chakrabarty +2 more
openaire +2 more sources
Lemmatization and Headword Structure
1997Abstract We now come to the actual structure and presentation of Palsgrave’s word list, and, as we can see from the following examples, his lexicographical method of presentation differs greatly from modem practice in bilingual English-French dictionaries.
openaire +1 more source

