Results 31 to 40 of about 2,065,709 (290)

Sense discrimination with parallel corpora [PDF]

open access: yesProceedings of the ACL-02 workshop on Word sense disambiguation recent successes and future directions -, 2002
This paper describes an experiment that uses translation equivalents derived from parallel corpora to determine sense distinctions that can be used for automatic sense-tagging and other disambiguation tasks. Our results show that sense distinctions derived from cross-lingual information are at least as reliable as those made by human annotators ...
Nancy Ide, Tomaz Erjavec, Dan Tufis
openaire   +1 more source

Aligning sentences in parallel corpora [PDF]

open access: yesProceedings of the 29th annual meeting on Association for Computational Linguistics -, 1991
In this paper we describe a statistical technique for aligning sentences with their translations in two parallel corpora. In addition to certain anchor points that are available in our data, the only information about the sentences that we use for calculating alignments is the number of tokens that they contain.
Peter F. Brown   +2 more
openaire   +1 more source

Paraphrasing with bilingual parallel corpora [PDF]

open access: yesProceedings of the 43rd Annual Meeting on Association for Computational Linguistics - ACL '05, 2005
Previous work has used monolingual parallel corpora to extract and generate paraphrases. We show that this task can be done using bilingual parallel corpora, a much more commonly available resource. Using alignment techniques from phrasebased statistical machine translation, we show how paraphrases in one language can be identified using a phrase in ...
Bannard, Colin, Callison-Burch, Chris
openaire   +1 more source

Dual Conditional Cross-Entropy Filtering of Noisy Parallel Corpora [PDF]

open access: yesConference on Machine Translation, 2018
In this work we introduce dual conditional cross-entropy filtering for noisy parallel data. For each sentence pair of the noisy parallel corpus we compute cross-entropy scores according to two inverse translation models trained on clean data. We penalize
Marcin Junczys-Dowmunt
semanticscholar   +1 more source

Feasibility of using corpora as a tool in translation practice

open access: yesAustralian Journal of Applied Linguistics, 2023
Professional translators commonly employ various tools to streamline and ensure the accuracy and consistency of their work. One such tool is corpora, which becomes particularly crucial when dealing with authentic texts like those from the United Nations
Noureldin Abdelaal
doaj   +1 more source

Pièges méthodologiques des corpus parallèles et comment les éviter

open access: yesCorela, 2017
In this article, we present methodological problems pertaining to the exploitation of parallel corpora, i.e. corpora composed of translations and their respective originals, and we try to propose the principles and rules helping to avoid said problems ...
Olga Nádvorníková
doaj   +1 more source

Voice conversion from non-parallel corpora using variational auto-encoder [PDF]

open access: yesAsia-Pacific Signal and Information Processing Association Annual Summit and Conference, 2016
We propose a flexible framework for spectral conversion (SC) that facilitates training with unaligned corpora. Many SC frameworks require parallel corpora, phonetic alignments, or explicit frame-wise correspondence for learning conversion functions or ...
Chin-Cheng Hsu   +4 more
semanticscholar   +1 more source

The use of English, Czech and French punctuation marks in reference, parallel and comparable web corpora: a question of methodology [PDF]

open access: yesLinguistica Pragensia, 2020
This paper analyses the frequency of six punctuation marks (the comma, period, colon, semicolon, question mark and exclamation mark) in three languages (English, French and Czech) in three different types of corpora — comparable web corpora, large ...
Olga Nádvorníková
doaj  

UPC: An Open Word-Sense Annotated Parallel Corpora for Machine Translation Study

open access: yesApplied Sciences, 2020
Machine translation (MT) has recently attracted much research on various advanced techniques (i.e., statistical-based and deep learning-based) and achieved great results for popular languages.
Van-Hai Vu   +3 more
doaj   +1 more source

OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles

open access: yesInternational Conference on Language Resources and Evaluation, 2016
We present a new major release of the OpenSubtitles collection of parallel corpora. The release is compiled from a large database of movie and TV subtitles and includes a total of 1689 bitexts spanning 2.6 billion sentences across 60 languages.
Pierre Lison, J. Tiedemann
semanticscholar   +1 more source

Home - About - Disclaimer - Privacy