Results 21 to 30 of about 2,065,709 (290)
Bootstrapping parallel corpora [PDF]
We present two methods for the automatic creation of parallel corpora. Whereas previous work into the automatic construction of parallel corpora has focused on harvesting them from the web, we examine the use of existing parallel corpora to bootstrap data for new language pairs.
Chris Callison-Burch, Miles Osborne
openaire +2 more sources
Building English – Punjabi Aligned Parallel Corpora of Nouns from Comparable Corpora
Comparable corpora are the right resources for extracting parallel data due to their abundant availability. It is of great importance where parallel data are scarce.
Kaur Dilshad, Singh Satwinder
doaj +1 more source
The neural machine translation models for the low-resource Kazakh–English language pair [PDF]
The development of the machine translation field was driven by people’s need to communicate with each other globally by automatically translating words, sentences, and texts from one language into another.
Vladislav Karyukin +4 more
doaj +2 more sources
Automatic alignment in parallel corpora [PDF]
This paper addresses the alignment issue in the framework of exploitation of large bimultilingual corpora for translation purposes. A generic alignment scheme is proposed that can meet varying requirements of different applications. Depending on the level at which alignment is sought, appropriate surface linguistic information is invoked coupled with ...
Harris Papageorgiou +2 more
openaire +1 more source
Automatic Generation of Exercises for Second Language Learning from Parallel Corpus Data [PDF]
Creating language learning exercises is a time-consuming task and made-up sample sentences frequently lack authenticity. Authentic samples can be obtained from corpora, but it is necessary to identify material that is suitable for language learners ...
Arianna Zanetti +2 more
doaj +1 more source
WebCrawl African : A Multilingual Parallel Corpora for African Languages
WebCrawl African is a mixed domain multilingual parallel corpora for a pool of African languages compiled by ANVITA machine translation team of Centre for Artificial Intelligence and Robotics Lab, primarily for accelerating research on low-resource and ...
Pavanpankaj Vegi +6 more
semanticscholar +1 more source
SciPar: A Collection of Parallel Corpora from Scientific Abstracts
This paper presents SciPar, a new collection of parallel corpora created from openly available metadata of bachelor theses, master theses and doctoral dissertations hosted in institutional repositories, digital libraries of universities and national ...
Dimitrios Roussis +4 more
semanticscholar +1 more source
MulTed: a multilingual aligned and tagged parallel corpus [PDF]
Recently, more data-driven approaches are demanding multilingual parallel resources primarily in the cross-language studies. To meet these demands, building multilingual parallel corpora are becoming the focus of many Natural Language Processing (NLP ...
Imad Zeroual, Abdelhak Lakhouaja
doaj +1 more source
Research has shown that parallel corpora have potential benefits for translator training and education. Most of the current available Arabic corpora, modern standard or dialectical, are monolingual in nature and there is an apparent lack in the Arabic ...
Awad Alhassan, Y. M. Sabtan, L. I. Omar
semanticscholar +1 more source
This paper surveys the strategies that the Contrastive, Typological, and Translation Mining parallel corpus traditions rely on to deal with the issue of target language representativeness of translations.
Bert Le Bruyn +6 more
doaj +1 more source

