Results 11 to 20 of about 43,026 (265)
The EuroPat Corpus: A Parallel Corpus of European Patent Data
We present the EuroPat corpus of patent-specific parallel data for 6 official European languages paired with English: German, Spanish, French, Croatian, Norwegian, and Polish. The filtered parallel corpora range in size from 51 million sentences (Spanish-English) to 154k sentences (Croatian-English), with the unfiltered (raw) corpora being up to 2 ...
Heafield, Kenneth +4 more
openaire +3 more sources
A Data Augmentation Method for English-Vietnamese Neural Machine Translation
The translation quality of machine translation systems depends on the parallel corpus used for training, particularly on the quantity and quality of the corpus.
Nghia Luan Pham +2 more
doaj +1 more source
Dutch Parallel Corpus: A Balanced Copyright-Cleared Parallel Corpus [PDF]
This paper presents the Dutch Parallel Corpus, a high-quality parallel corpus for Dutch, French and English consisting of more than ten million words. The corpus contains five different text types and is balanced with respect to text type and translation direction. All texts included in the corpus have been cleared from copyright.
Macken, Lieve +2 more
openaire +1 more source
Consumer Eroski parallel corpus
This paper introduces the Consumer Eroski Parallel Corpus, a collection of articles originally written in Spanish and later translated to three languages also spoken in Spain: Basque, Catalan and Galician. The articles have been correlated in the four
Asier Alcázar
doaj +1 more source
The Bulgarian-Polish-Russian parallel corpus
The Bulgarian-Polish-Russian parallel corpus The Semantics Laboratory Team of Institute of Slavic Studies of Polish Academy of Sciences is planning to begin work on the creation of a Bulgarian-Polish-Russian parallel corpus.
Maksim Duškin +1 more
doaj +1 more source
Introduction – Languages in contrast 20 years on
This special issue of the Nordic Journal of English Studies comprises papers from the symposium Languages in Contrast held in Lund 5 December 2014 in celebration of the 20th anniversary of the Nordic Parallel Corpus Project which led to the English ...
Lene Nordrum +2 more
doaj +1 more source
TweetMT: A Parallel Microblog Corpus
We introduce TweetMT, a parallel corpus of tweets in four language pairs that combine five languages (Spanish from/to Basque, Catalan, Galician and Portuguese), all of which have an official status in the Iberian Peninsula. The corpus has been created by combining automatic collection and crowdsourcing approaches, and it is publicly available.
Iñaki San Vicente +8 more
openaire +2 more sources
PARALLEL CORPUS OF TEXTS: THEORETICAL, METHODOLOGICAL AND LEXICOGRAPHICAL ANALYSIS OF PRINCIPLES
The article deals with theoretical and methodological aspects of processing two or more texts from the corpus translation, indicates a positive lexicographical level of parallel corpus translation (text enrichment stable phrases, idioms, specification ...
Ю. І. Дем’янчук
doaj +1 more source
Korpus paralel memiliki peran yang sangat penting dalam mesin penerjemah statistik (MPS). Korpus paralel yang diperoleh berbagai sumber biasanya memiliki kualitas yang kurang baik, sedangkan kuantitas korpus paralel merupakan tuntutan utama bagi hasil ...
Herry Sujaini
doaj +1 more source
Sign languages are used by the deaf and mute community of the world. These are gesture based languages where the subjects use hands and facial expressions to perform different gestures.
Uzma Farooq +4 more
doaj +1 more source

