Results 41 to 50 of about 645,699 (249)
Developing a POS Tagged Corpus of Urdu Tweets
Processing of social media text like tweets is challenging for traditional Natural Language Processing (NLP) tools developed for well-edited text due to the noisy nature of such text.
Amber Baig +3 more
doaj +1 more source
Old Catalan Morphosyntax: Developing an Annotated Corpus
This paper presents a full procedure for the development of a Part-of-Speech (POS) tagged corpus of Old Catalan. As an extremely low-resource language with rich inflection and frequent homographs, Old Catalan poses non-trivial problems in the development
Marieke Meelen, Afra Pujol i Campeny
doaj +1 more source
Improving Data Driven Part-of-Speech Tagging by Morphologic Knowledge Induction [PDF]
We present a Markov part-of-speech tagger for which the P (w|t) emission probabilities of word w given tag t are replaced by a linear interpolation of tag emission probabilities given a list of representations of w.
Reichel, Uwe D., Uwe D. Reichel
core +1 more source
Improving part-of-speech tagging in Amharic language using deep neural network
To date, several POS taggers have been introduced to facilitate the success of semantic analysis for different languages. However, the task of POS tagging becomes a bit intricate in morphologically complex languages, like Amharic.
Sintayehu Hirpassa, G.S. Lehal
doaj +1 more source
Automatic natural language processing of large texts often presents recurring challenges in multiple languages: even for most advanced tasks, the texts are first processed by basic processing steps – from tokenization to parsing.
Milan Straka, Jan Hajic, Jana Straková
semanticscholar +1 more source
POS Tagging and Its Applications for Mathematics [PDF]
Content analysis of scientific publications is a nontrivial task, but a useful and important one for scientific information services. In the Gutenberg era it was a domain of human experts; in the digital age many machine-based methods, e.g., graph analysis tools and machine-learning techniques, have been developed for it.
Ulf Schöneberg, Wolfram Sperber
openaire +2 more sources
Chakma Language POS Tagging Dataset [PDF]
The Chakma Language POS Tagging Dataset is a valuable linguistic resource designed for the analysis and understanding of the Chakma language. Chakma is a member of the Indo-Aryan language family and is primarily spoken by the Chakma people in the ...
Rahman, M (via Mendeley Data)
core +1 more source
POS Tagging Bahasa Madura dengan Menggunakan Algoritma Brill Tagger
Bahasa Madura adalah bahasa daerah yang selain digunakan di Pulau Madura juga digunakan di daerah lainnya seperti di kota Jember, Pasuruan, dan Probolinggo.
Nindian Puspa Dewi, Ubaidi Ubaidi
doaj +1 more source
Simple Semi-Supervised POS Tagging [PDF]
We tackle the question: how much supervision is needed to achieve state-of-the-art performance in part-of-speech (POS) tagging, if we leverage lexical representations given by the model of Brown et al. (1992)? It has become a standard practice to use automatically induced “Brown clusters” in place of POS tags.
Karl Stratos, Michael Collins 0001
openaire +1 more source
The challenge of POS tagging and lemmatization in morphologically rich languages is examined by comparing German and Latin. We start by defining an NLP evaluation roadmap to model the combination of tools and resources guiding our experiments.
Rüdiger Gleim +8 more
doaj +1 more source

