Results 1 to 10 of about 160 (101)

MulTed: a multilingual aligned and tagged parallel corpus [PDF]

open access: yesApplied Computing and Informatics, 2022
Recently, more data-driven approaches are demanding multilingual parallel resources primarily in the cross-language studies. To meet these demands, building multilingual parallel corpora are becoming the focus of many Natural Language Processing (NLP ...
Abdelhak Lakhouaja, Imad Zeroual
exaly   +2 more sources

COMPILACIÓN Y ANÁLISIS DE UN CORPUS PARALELO PARA LA INVESTIGACIÓN EN TRADUCCIÓN: PROYECTO CON DÉJÀ VU, TREETAGGER E IMS OPEN CORPUS WORKBENCH [PDF]

open access: yesRLA Revista De Lingüística Teórica Y Aplicada, 2016
[ES] Aunque en los últimos años la lingüística de corpus ha experimentado una gran evolución y en la actualidad cuenta con una creciente presencia en proyectos de investigación en tor- no a estudios de Lingüística y Traducción (por ejemplo: Kübler y Foucou, 2003; Laroche y Langlais, 2010), los procedimientos técnicos más avanzados enfocados a la ...
TERESA Molés-Cases
exaly   +4 more sources

Stylistic Authorship Comparison and Attribution of Spanish News Forum Messages Based on the TreeTagger POS Tagger

open access: yesProcedia, Social and Behavioral Sciences, 2015
AbstractElectronic texts from emails, social networks or mobile phones are currently of interest in Forensic Linguistics. Many of these texts analyzed are well under 200 words long. This work aims at identifying text authorship by using part-of-speech tags over short texts. Our corpus consists of 28 texts taken from forum messages.
Crespo, Mario, Frías, Antonio
exaly   +2 more sources

Towards a standard Part of Speech tagset for the Arabic language

open access: yesJournal of King Saud University - Computer and Information Sciences, 2017
Part of Speech (PoS) tagging is still not very well investigated with respect to the Arabic language. Determining the PoS tags of a word in a particular context is difficult, primarily because there is no use of diacritics in most of contemporary texts ...
Abdelhak Lakhouaja, Imad Zeroual
exaly   +3 more sources

Evaluating a pivot-based approach for bilingual lexicon extraction. [PDF]

open access: yesComput Intell Neurosci, 2015
A pivot‐based approach for bilingual lexicon extraction is based on the similarity of context vectors represented by words in a pivot language like English. In this paper, in order to show validity and usability of the pivot‐based approach, we evaluate the approach in company with two different methods for estimating context vectors: one estimates them
Kim JH, Kwon HS, Seo HW.
europepmc   +2 more sources

Issues in training the TreeTagger for Georgian

open access: yesCorpora
The paper describes the process of retraining the TreeTagger program ( Schmid, 1994 ) for the Georgian language. This includes considering some general procedures such as designing a training corpus, creating a tagging lexicon, and training the TreeTagger on Georgian texts.
exaly   +3 more sources

An auxiliary Part‐of‐Speech tagger for blog and microblog cyber‐slang

open access: yesStatistical Analysis and Data Mining: The ASA Data Science Journal, Volume 16, Issue 1, Page 65-79, February 2023., 2023
Abstract The increasing impact of Web 2.0 involves a growing usage of slang, abbreviations, and emphasized words, which limit the performance of traditional natural language processing models. The state‐of‐the‐art Part‐of‐Speech (POS) taggers are often unable to assign a meaningful POS tag to all the words in a Web 2.0 text.
Silvia Golia, Paola Zola
wiley   +1 more source

DeepEva: A deep neural network architecture for assessing sentence complexity in Italian and English languages

open access: yesArray, 2021
Automatic Text Complexity Evaluation (ATE) is a research field that aims at creating new methodologies to make autonomous the process of the text complexity evaluation, that is the study of the text-linguistic features (e.g., lexical, syntactical ...
Giosué Lo Bosco   +2 more
doaj   +1 more source

Effects of Availability, Contingency, and Formulaicity on the Accuracy of English Grammatical Morphemes in Second Language Writing

open access: yesLanguage Learning, Volume 72, Issue 4, Page 899-940, December 2022., 2022
Abstract We investigated whether the accuracy of grammatical morphemes in second language (L2) learners’ writing is associated with usage‐based distributional factors. Specifically, we examined whether the accuracy of L2 English inflectional morphemes is associated with the availability (i.e., token frequency) and contingency (i.e., token frequency ...
Akira Murakami, Nick C. Ellis
wiley   +1 more source

MT Evaluation in the Context of Language Complexity

open access: yesComplexity, Volume 2021, Issue 1, 2021., 2021
The paper focuses on investigating the impact of artificial agent (machine translator) on human agent (posteditor) using a proposed methodology, which is based on language complexity measures, POS tags, frequent tagsets, association rules, and their summarization. We examine this impact from the point of view of language complexity in terms of word and
Dasa Munkova   +4 more
wiley   +1 more source

Home - About - Disclaimer - Privacy