Results 21 to 30 of about 9,795,875 (280)

The Cardamom workbench for historical and under-resourced languages

open access: yes, 2023
This paper describes the creation of a workbench tool designed to make technologies developed throughout the lifespan of the Cardamom project easily accessible to researchers who could most benefit from them, but who may not have the technical expertise to apply bleeding edge technologies to their own datasets.
Doyle, Adrian   +5 more
openaire   +4 more sources

Author identification for Under-Resourced language (KadazanDusun)

open access: yesIndonesian Journal of Electrical Engineering and Computer Science, 2020
<span>This paper presents the task of Author Identification for KadazanDusun language by using tweets as the source of data to perform Author Identification task of short text on KadazanDusun, which is considered as one the under-resourced language in Malaysia.
Nursyahirah Tarmizi   +2 more
openaire   +2 more sources

Creating language resources for under-resourced languages: methodologies, and experiments with Arabic [PDF]

open access: yesLanguage Resources and Evaluation, 2014
Language resources are important for those working on computational methods to analyse and study languages. These resources are needed to help advancing the research in fields such as natural language processing, machine learning, information retrieval and text analysis in general.
Mahmoud El-Haj   +2 more
openaire   +4 more sources

Building Speech Recognition Systems for Language Documentation: The CoEDL Endangered Language Pipeline and Inference System (ELPIS) [PDF]

open access: yes, 2018
Machine learning has revolutionised speech technologies for major world languages, but these technologies have generally not been available for the roughly 4,000 languages with populations of fewer than 10,000 speakers.
František Kratochvíl   +48 more
core   +1 more source

A Python package for text processing for Serbian: nlpheart [PDF]

open access: yesScientific Technical Review, 2020
Within the past two decades, text processing became an important part of most state-of-the-art advanced automation systems. However, for many under-resourced languages it is still challenging to perform textual data preparation, due to the lack of ...
Ostrogonac Stevan   +2 more
doaj   +1 more source

Automatic Speech Recognition Using Limited Vocabulary: A Survey

open access: yesApplied Artificial Intelligence, 2022
Automatic Speech Recognition (ASR) is an active field of research due to its large number of applications and the proliferation of interfaces or computing devices that can support speech processing.
Jean Louis K. E Fendji   +3 more
doaj   +1 more source

Offensive Language Detection in Under-Resourced Algerian Dialectal Arabic Language

open access: yes, 2023
This paper addresses the problem of detecting the offensive and abusive content in Facebook comments, where we focus on the Algerian dialectal Arabic which is one of under-resourced languages. The latter has a variety of dialects mixed with different languages (i.e. Berber, French and English). In addition, we deal with texts written in both Arabic and
Oussama Boucherit, Kheireddine Abainia
openaire   +3 more sources

Code-Switching in Automatic Speech Recognition: The Issues and Future Directions

open access: yesApplied Sciences, 2022
Code-switching (CS) in spoken language is where the speech has two or more languages within an utterance. It is an unsolved issue in automatic speech recognition (ASR) research as ASR needs to recognise speech in bilingual and multilingual settings ...
Mumtaz Begum Mustafa   +6 more
doaj   +1 more source

MaCoCu:Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages [PDF]

open access: yes, 2022
We introduce the project MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages, funded by the Connecting Europe Facility, which is aimed at building monolingual and parallel corpora for under ...
Ramírez-Sánchez, Gema   +13 more
core   +10 more sources

Extractive summarization of Malayalam documents using latent Dirichlet allocation: An experience

open access: yesJournal of Intelligent Systems, 2022
Automatic text summarization (ATS) extracts information from a source text and presents it to the user in a condensed form while preserving its primary content.
Kondath Manju   +2 more
doaj   +1 more source

Home - About - Disclaimer - Privacy