Results 221 to 230 of about 10,051 (260)
Some of the next articles are maybe not open access.
International Journal of Corpus Linguistics, 2001
Corpus linguistics lacks strategies for describing and comparing corpora. Currently most descriptions of corpora are textual, and questions such as ‘what sort of a corpus is this?’, or ‘how does this corpus compare to that?’ can only be answered impressionistically.
openaire +1 more source
Corpus linguistics lacks strategies for describing and comparing corpora. Currently most descriptions of corpora are textual, and questions such as ‘what sort of a corpus is this?’, or ‘how does this corpus compare to that?’ can only be answered impressionistically.
openaire +1 more source
Comparing web-crawled and traditional corpora
Language Resources and Evaluation, 2020Using a multi-dimensional (MD) analysis of register variability, the study compares two corpora of Czech: Koditex, a “traditional” corpus carefully designed using various sources with rich metadata, and Araneum Bohemicum Maximum, a web-crawled corpus with an opportunistic composition representative of the “searchable” web.
Václav Cvrcek +6 more
openaire +1 more source
2019
The availability of parallel corpora is limited, especially for under-resourced languages and narrow domains. On the other hand, the number of comparable documents in these areas that are freely available on the Web is continuously increasing. Algorithmic approaches to identify these documents from the Web are needed for the purpose of automatically ...
Monica Lestari Paramita +11 more
openaire +1 more source
The availability of parallel corpora is limited, especially for under-resourced languages and narrow domains. On the other hand, the number of comparable documents in these areas that are freely available on the Web is continuously increasing. Algorithmic approaches to identify these documents from the Web are needed for the purpose of automatically ...
Monica Lestari Paramita +11 more
openaire +1 more source
Twitter As a Multilingual Source of Comparable Corpora
Proceedings of the 12th International Conference on Advances in Mobile Computing and Multimedia, 2014This article describes a new method to build comparable corpora from Twitter. Our strategy relies on the fact that Twitter is one of the most popular online social microblog allowing large audiences to express their thoughts and reactions about specific events or breaking news in various languages. Given two languages and a particular topic, We propose
Malek Hajjem +2 more
openaire +1 more source
Comparative study on corpora for speech translation
IEEE Transactions on Audio, Speech and Language Processing, 2006This paper investigates issues in preparing corpora for developing speech-to-speech translation (S2ST). It is impractical to create a broad-coverage parallel corpus only from dialog speech. An alternative approach is to have bilingual experts write conversational-style texts in the target domain, with translations.
Gen-ichiro Kikui +3 more
openaire +1 more source
2013
One of the major bottlenecks in the development of Statistical Machine Translation systems for most language pairs is the lack of bilingual parallel training data. Currently available parallel corpora span relatively few language pairs and very few domains; building new ones of sufficiently large size and high quality is time-consuming and expensive.
Dragos Stefan Munteanu, Daniel Marcu
openaire +1 more source
One of the major bottlenecks in the development of Statistical Machine Translation systems for most language pairs is the lack of bilingual parallel training data. Currently available parallel corpora span relatively few language pairs and very few domains; building new ones of sufficiently large size and high quality is time-consuming and expensive.
Dragos Stefan Munteanu, Daniel Marcu
openaire +1 more source
Semi-Automatic Parallel Corpora Extraction from Comparable News Corpora
Polibits, 2010The parallel corpus is a necessary resource in many multi/cross lingual natural language processing applications that include Machine Translation and Cross Lingual Information Retreival. Preparation of large scale parallel corpus takes time and also demands the linguistics skill.
Thoudam Doren Singh +1 more
openaire +1 more source
How Comparable Can 'Comparable Corpora' Be?
Target. International Journal of Translation Studies, 1997AbstractThe development of a coherent methodology for corpus-based work in translation studies is essential for the evolution of this newfield of research into a fully-fledged paradigm within the discipline. The design of a monolingual, multi-source-language comparable corpus of English as a resource for the systematic study of the nature of translated
openaire +1 more source
Building and Using Comparable Corpora
2013The 1990s saw a paradigm change in the use of corpus-driven methods in NLP. In the field of multilingual NLP (such as machine translation and terminology mining) this implied the use of parallel corpora. However, parallel resources are relatively scarce: many more texts are produced daily by native speakers of any given language than translated.
openaire +1 more source
Repetition and language models and comparable corpora
Proceedings of the 2nd Workshop on Building and Using Comparable Corpora from Parallel to Non-parallel Corpora - BUCC '09, 2009I will discuss a couple of non-standard features that I believe could be useful for working with comparable corpora. Dotplots have been used in biology to find interesting DNA sequences. Biology is interested in ordered matches, which show up as (possibly broken) diagonals in dot-plots. Information Retrieval is more interested in unordered matches (e.g.
openaire +2 more sources

