Results 1 to 10 of about 160 (101)
MulTed: a multilingual aligned and tagged parallel corpus [PDF]
Recently, more data-driven approaches are demanding multilingual parallel resources primarily in the cross-language studies. To meet these demands, building multilingual parallel corpora are becoming the focus of many Natural Language Processing (NLP ...
Abdelhak Lakhouaja, Imad Zeroual
exaly +2 more sources
COMPILACIÓN Y ANÁLISIS DE UN CORPUS PARALELO PARA LA INVESTIGACIÓN EN TRADUCCIÓN: PROYECTO CON DÉJÀ VU, TREETAGGER E IMS OPEN CORPUS WORKBENCH [PDF]
[ES] Aunque en los últimos años la lingüística de corpus ha experimentado una gran evolución y en la actualidad cuenta con una creciente presencia en proyectos de investigación en tor- no a estudios de Lingüística y Traducción (por ejemplo: Kübler y Foucou, 2003; Laroche y Langlais, 2010), los procedimientos técnicos más avanzados enfocados a la ...
TERESA Molés-Cases
exaly +4 more sources
AbstractElectronic texts from emails, social networks or mobile phones are currently of interest in Forensic Linguistics. Many of these texts analyzed are well under 200 words long. This work aims at identifying text authorship by using part-of-speech tags over short texts. Our corpus consists of 28 texts taken from forum messages.
Crespo, Mario, Frías, Antonio
exaly +2 more sources
Towards a standard Part of Speech tagset for the Arabic language
Part of Speech (PoS) tagging is still not very well investigated with respect to the Arabic language. Determining the PoS tags of a word in a particular context is difficult, primarily because there is no use of diacritics in most of contemporary texts ...
Abdelhak Lakhouaja, Imad Zeroual
exaly +3 more sources
Evaluating a pivot-based approach for bilingual lexicon extraction. [PDF]
A pivot‐based approach for bilingual lexicon extraction is based on the similarity of context vectors represented by words in a pivot language like English. In this paper, in order to show validity and usability of the pivot‐based approach, we evaluate the approach in company with two different methods for estimating context vectors: one estimates them
Kim JH, Kwon HS, Seo HW.
europepmc +2 more sources
Issues in training the TreeTagger for Georgian
The paper describes the process of retraining the TreeTagger program ( Schmid, 1994 ) for the Georgian language. This includes considering some general procedures such as designing a training corpus, creating a tagging lexicon, and training the TreeTagger on Georgian texts.
exaly +3 more sources
An auxiliary Part‐of‐Speech tagger for blog and microblog cyber‐slang
Abstract The increasing impact of Web 2.0 involves a growing usage of slang, abbreviations, and emphasized words, which limit the performance of traditional natural language processing models. The state‐of‐the‐art Part‐of‐Speech (POS) taggers are often unable to assign a meaningful POS tag to all the words in a Web 2.0 text.
Silvia Golia, Paola Zola
wiley +1 more source
Automatic Text Complexity Evaluation (ATE) is a research field that aims at creating new methodologies to make autonomous the process of the text complexity evaluation, that is the study of the text-linguistic features (e.g., lexical, syntactical ...
Giosué Lo Bosco +2 more
doaj +1 more source
Abstract We investigated whether the accuracy of grammatical morphemes in second language (L2) learners’ writing is associated with usage‐based distributional factors. Specifically, we examined whether the accuracy of L2 English inflectional morphemes is associated with the availability (i.e., token frequency) and contingency (i.e., token frequency ...
Akira Murakami, Nick C. Ellis
wiley +1 more source
MT Evaluation in the Context of Language Complexity
The paper focuses on investigating the impact of artificial agent (machine translator) on human agent (posteditor) using a proposed methodology, which is based on language complexity measures, POS tags, frequent tagsets, association rules, and their summarization. We examine this impact from the point of view of language complexity in terms of word and
Dasa Munkova +4 more
wiley +1 more source

