Results 31 to 40 of about 25,957 (254)
Current Challenges in Web Corpus Building
In this paper we discuss some of the current challenges in web corpus building that we faced in the recent years when expanding the corpora in Sketch Engine. The purpose of the paper is to provide an overview and raise discussion on possible solutions, rather than bringing ready solutions to the readers.
Milos Jakubícek +3 more
openaire +3 more sources
EuroGOV: Engineering a Multilingual Web Corpus [PDF]
EuroGOV is a multilingual web corpus that was created to serve as the document collection for WebCLEF, the CLEF 2005 web retrieval task. EuroGOV is a collection of web pages crawled from the European Union portal, European Union member state governmental web sites, and Russian government web sites.
Börkur Sigurbjörnsson +2 more
openaire +4 more sources
Fostering Innovation: Streamlining Magnetocaloric Materials Research by Digitalization
Magnetocaloric cooling (MCE) is an environmentally friendly refrigeration method with great potential. Optimizing MCE materials involves the preparation and screening of large quantities of samples, which in turn generates a large amount of data. A digitalization approach is presented that uses ontologies, knowledge graphs, and digital workflows to ...
Simon Bekemeier +17 more
wiley +1 more source
The paper describes three software packages – the main components of a software system for processing and web-presentation of Bulgarian language resources – parallel corpora and bilingual dictionaries. The author briefly prese nts current versions of the
Ralitsa Dutsova
doaj +1 more source
Building machine‐readable vocabularies for materials science is slow, expert‐driven work. This study benchmarks 13 large language models on two of its first steps: finding candidate terms in engineering articles and deciding where they belong in a class hierarchy.
Thomas Bjarsch +3 more
wiley +1 more source
The aim of this study is to analyze drug mentions in web forums to evaluate the utility of this data source for drug post-marketing studies. We automatically annotated over 60 million posts extracted from 21 French web forums.
Bissan Audeh +6 more
doaj +1 more source
Reproduction of stacking fault energy calculations from literature with a semi‐automated large language model‐assisted extraction procedure: extraction of simulation protocol, atomistic structures, computational parameters, and reported results, ontology alignment, knowledge graph construction and, finally, recomputation forvalidation.
Sepideh Baghaee Ravari +5 more
wiley +1 more source
Digitalizing electroplating requires both domain knowledge and interoperability. This work introduces PlatOn, a domain ontology for trivalent chromium plating and coating characterization, and a hybrid pipeline that aligns it to a mid‐level reference ontology by combining eight similarity metrics with language model reasoning. Expert‐validated mappings
Janik Harter +10 more
wiley +1 more source
A Comprehensive Corpus-based Study of the Use of Evaluative Adjectives in Promotional Hotel Websites
: The purpose of this paper is to analyse and, if necessary, qualify the high incidence of evaluative adjectives presumably to be expected in promotional hotel websites.
Nuria Edo Marzá
doaj +1 more source
The SPECIES and ORGANISMS Resources for Fast and Accurate Identification of Taxonomic Names in Text. [PDF]
The exponential growth of the biomedical literature is making the need for efficient, accurate text-mining tools increasingly clear. The identification of named biological entities in text is a central and difficult task.
Evangelos Pafilis +7 more
doaj +1 more source

