Results 31 to 40 of about 2,065,709 (290)
Sense discrimination with parallel corpora [PDF]
This paper describes an experiment that uses translation equivalents derived from parallel corpora to determine sense distinctions that can be used for automatic sense-tagging and other disambiguation tasks. Our results show that sense distinctions derived from cross-lingual information are at least as reliable as those made by human annotators ...
Nancy Ide, Tomaz Erjavec, Dan Tufis
openaire +1 more source
Aligning sentences in parallel corpora [PDF]
In this paper we describe a statistical technique for aligning sentences with their translations in two parallel corpora. In addition to certain anchor points that are available in our data, the only information about the sentences that we use for calculating alignments is the number of tokens that they contain.
Peter F. Brown +2 more
openaire +1 more source
Paraphrasing with bilingual parallel corpora [PDF]
Previous work has used monolingual parallel corpora to extract and generate paraphrases. We show that this task can be done using bilingual parallel corpora, a much more commonly available resource. Using alignment techniques from phrasebased statistical machine translation, we show how paraphrases in one language can be identified using a phrase in ...
Bannard, Colin, Callison-Burch, Chris
openaire +1 more source
Dual Conditional Cross-Entropy Filtering of Noisy Parallel Corpora [PDF]
In this work we introduce dual conditional cross-entropy filtering for noisy parallel data. For each sentence pair of the noisy parallel corpus we compute cross-entropy scores according to two inverse translation models trained on clean data. We penalize
Marcin Junczys-Dowmunt
semanticscholar +1 more source
Feasibility of using corpora as a tool in translation practice
Professional translators commonly employ various tools to streamline and ensure the accuracy and consistency of their work. One such tool is corpora, which becomes particularly crucial when dealing with authentic texts like those from the United Nations
Noureldin Abdelaal
doaj +1 more source
Pièges méthodologiques des corpus parallèles et comment les éviter
In this article, we present methodological problems pertaining to the exploitation of parallel corpora, i.e. corpora composed of translations and their respective originals, and we try to propose the principles and rules helping to avoid said problems ...
Olga Nádvorníková
doaj +1 more source
Voice conversion from non-parallel corpora using variational auto-encoder [PDF]
We propose a flexible framework for spectral conversion (SC) that facilitates training with unaligned corpora. Many SC frameworks require parallel corpora, phonetic alignments, or explicit frame-wise correspondence for learning conversion functions or ...
Chin-Cheng Hsu +4 more
semanticscholar +1 more source
The use of English, Czech and French punctuation marks in reference, parallel and comparable web corpora: a question of methodology [PDF]
This paper analyses the frequency of six punctuation marks (the comma, period, colon, semicolon, question mark and exclamation mark) in three languages (English, French and Czech) in three different types of corpora — comparable web corpora, large ...
Olga Nádvorníková
doaj
UPC: An Open Word-Sense Annotated Parallel Corpora for Machine Translation Study
Machine translation (MT) has recently attracted much research on various advanced techniques (i.e., statistical-based and deep learning-based) and achieved great results for popular languages.
Van-Hai Vu +3 more
doaj +1 more source
OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles
We present a new major release of the OpenSubtitles collection of parallel corpora. The release is compiled from a large database of movie and TV subtitles and includes a total of 1689 bitexts spanning 2.6 billion sentences across 60 languages.
Pierre Lison, J. Tiedemann
semanticscholar +1 more source

