Results 21 to 30 of about 25,957 (254)

A learner corpus is born this way: From raw data to processed dataset

open access: yesData in Brief, 2022
This data article presents the development of a learner corpus (i.e. a systematic computerized web-based repository of written texts produced by language learners) from the initial phase of the development where written assignments were collected from ...
Chung Hong Danny Leung   +2 more
doaj   +1 more source

A Web Corpus and Word Sketches for Japanese

open access: yesJournal of Natural Language Processing, 2008
Of all the major world languages, Japanese is lagging behind in terms of publicly accessible and searchable corpora. In this paper we describe the development of JpWaC (Japanese Web as Corpus), a large corpus of 400 million words of Japanese web text, and its encoding for the Sketch Engine.
Erjavec, Irena Srdanovic   +2 more
openaire   +3 more sources

The slWaC Corpus of the Slovene Web

open access: yesInformatica (Slovenia), 2015
The availability of large collections of text (language corpora) is crucial for empirically supported linguis- tic investigations of various languages ; however, such corpora are complicated and expensive to collect. In recent years corpora made from texts on the World Wide Web have become an attractive alternative to traditional corpora, as they can ...
Erjavec, Tomaž   +2 more
openaire   +5 more sources

Future (and not-so-future) trends in the teaching of translation technology

open access: yesRevista Tradumàtica, 2013
This paper proposes an approach to teaching translation technology that focus less on exposing students to ever more types of CAT tools than on two sets of meta-competences—revising skills and documentary research skills—and on the technologies that ...
Frank Austermuehl
doaj   +1 more source

Annotated Lexicon for Sentiment Analysis in the Bosnian Language

open access: yesSlovenščina 2.0: Empirične, aplikativne in interdisciplinarne raziskave, 2023
The paper presents the first sentiment-annotated lexicon of the Bosnian language. The annotation process and methodology are presented along with a usability study, which concentrates on language coverage. The composition of the starting base was done by
Sead Jahić, Jernej Vičič
doaj   +1 more source

Providing Web Archive News Articles as Corpus Data

open access: yesJournal of Open Humanities Data
While the huge data repositories of web archives carry big potential for knowledge production in academia, researchers have described significant challenges when trying to access and make use of web archives in research.
Jon Carlstedt Tønnessen   +1 more
doaj   +1 more source

Les humanités numériques pour repenser les catégories d’analyse

open access: yesRevue Française des Sciences de l’Information et de la Communication, 2016
Researchers analysing web corpus in digital humanities may have trouble categorising entities belonging to studied concepts. By combining two research fields – a territorial web and reputation of an online service – we will try to demonstrate that user ...
Mariannig Le Béchec, Camille Alloing
doaj   +1 more source

A Prototype Theory-Based Study of Crimes in the English and Arabic Societies Using Web-as-Corpus [PDF]

open access: yesThe Egyptian Journal of Language Engineering, 2017
Drawing on the classified prototypes of semantic categories, this paper uses the principles of the prototype theoryto cross-culturally explore the hierarchical prototypes of crimes in web-booted Arabic and English corpora.
Fadia Badawi   +2 more
doaj   +1 more source

Korpus XIX w. Uniwersytetu Warszawskiego i IJP PAN

open access: yesLingVaria, 2023
CORPUS OF THE 19TH CENTURY OF THE WARSAW UNIVERSITY AND IJP PAN The article describes a historical corpus which documents the 19th and early 20th century.
Marek Łaziński   +2 more
doaj   +1 more source

Concepções de Corpus de Análise na Pesquisa em Educação em Ciências Naturais: Uma Investigação em Dissertações e Teses de um Programa de Pós-Graduação

open access: yesRevista Brasileira de Pesquisa em Educação em Ciências, 2020
Investigaram-se neste trabalho compreensões e concepções sobre corpus de análise em dissertações e teses de um programa de pós-graduação em Educação em Ciências e Matemática.
Julio Murilo Trevas dos Santos   +1 more
doaj   +1 more source

Home - About - Disclaimer - Privacy