Results 21 to 30 of about 2,520 (188)
Joint Lemmatization and Morphological Tagging with Lemming [PDF]
We present LEMMING, a modular log-linear model that jointly models lemmatization and tagging and supports the integration of arbitrary global features. It is trainable on corpora annotated with gold standard tags and lemmata and does not rely on morphological dictionaries or analyzers.
Thomas Müller 0009 +3 more
openaire +2 more sources
Legal Terms in General Dictionaries of English: The Civil Procedure Mystery
Many general language dictionaries contain specialized terms, including legal terms relating to civil lawsuits. The existing literature provides general discussions of scientific and technical terms in ordinary dictionaries but does not specifically ...
Sandro Nielsen
doaj +1 more source
Die Lemmatisierung von Goethes "Wahlverwandtschaften"
In German studies computer-oriented methods are taking on an increasing significance and consequently should be considered in detail together with all the problems connected with them. This article deals with the automatic lemmatization of Goethe's novel
Martina Schwanke
doaj +1 more source
A Memory-Based Lemmatizer for Ancient Greek [PDF]
In this paper we present the lemmatizer that we developed for Ancient Greek: GLEM. As far as we know, GLEM is the first publicly available lemmatizer for Ancient Greek that uses POS information to disambiguate and that also assigns output to unseen words, words that are not yet in the lexicon.As the basis for the lemmatizer we used an existing memory ...
Corien Bary, Peter Berck, Iris Hendrickx
openaire +2 more sources
Detection of Sensitive Data to Counter Global Terrorism
Global terrorism has created challenges to the criminal justice system due to its abnormal activities, which lead to financial loss, cyberwar, and cyber-crime.
Binod Kumar Adhikari +4 more
doaj +1 more source
The LiLa Lemma Bank: A Knowledge Base of Latin Canonical Forms
The dataset contains a list of 215,102 Latin dictionary forms (known as canonical forms or lemmas). The dataset is a set of 1,699,687 Resource Description Framework (RDF) triples that describe, using a series of Web Ontology Language (OWL) ontologies for
Francesco Mambrini +1 more
doaj +1 more source
Annif Analyzer Shootout: Comparing text lemmatization methods for automated subject indexing
Automated text classification is an important function for many AI systems relevant to libraries, including automated subject indexing and classification.
Osma Suominen, Ilkka Koskenniemi
doaj
The regularization of Old English weak verbs
This article deals with the regularization of non-standard spellings of the verbal forms extracted from a corpus. It addresses the question of what the limits of regularization are when lemmatizing Old English weak verbs.
Marta Tío Sáenz
doaj +1 more source
Lemmatization and parsing with TACT preprocessing programs
None
R G Siemens
doaj +1 more source
BanglaLem: A Transformer-based Bangla Lemmatizer with an Enhanced Dataset
Lemmatization plays a crucial role in various natural language processing (NLP) tasks, such as information retrieval, sentiment analysis, text summarization, and text classification.
Md Fuadul Islam +4 more
doaj +1 more source

