Results 51 to 60 of about 1,363 (163)

AraReasoner: Evaluating Reasoning-Based LLMs for Arabic NLP

open access: yesFindings of the Association for Computational Linguistics: EMNLP 2025
Large language models (LLMs) have shown remarkable progress in reasoning abilities and general natural language processing (NLP) tasks, yet their performance on Arabic data, characterized by rich morphology, diverse dialects, and complex script, remains underexplored.
Ahmed Hasanaath   +5 more
openaire   +3 more sources

The role of first and second language reading, first language low‐level skills, and working memory in second language writing

open access: yesThe Modern Language Journal, Volume 110, Issue 1, Page 159-187, Spring 2026.
Abstract Our research investigated how L2 and L1 reading, L1 low‐level skills and working memory are related to ratings and the linguistic characteristics (productivity, cohesion, lexical sophistication and diversity, syntactic complexity, and accuracy) of argumentative and narrative texts. The research was conducted in Hungary with 95 secondary school
Judit Kormos, Csilla Bartha
wiley   +1 more source

PROCESSING TOOLS FOR CORPUS LINGUISTICS: A CASE STUDY ON ARABIC HISTORICAL CORPUS

open access: yesJordanian Journal of Computers and Information Technology
This paper explores the development, design, and reconstruction of a Historical Arabic Corpus (HAC), which covers more than 1600 years of uninterrupted language use.
Bassam Hasan Hammo, Sane Yagi
doaj   +1 more source

Mitigating the Negative Transfer in Multi‐Task Learning for Harmful Language Detection in Spanish and Arabic

open access: yesExpert Systems, Volume 43, Issue 2, February 2026.
ABSTRACT Negative transfer continues to limit the benefits of multi‐task learning (MTL) in harmful language detection, where related tasks must share representations without diluting task‐specific nuances. We introduce task awareness (TA), a methodological framework that explicitly conditions MTL models on the task they must solve.
Angel Felipe Magnossão de Paula   +3 more
wiley   +1 more source

Dataset of Arabic spam and ham tweets

open access: yesData in Brief
This data article provides a dataset of 132421 posts and their corresponding information collected from Twitter social media. The data has two classes, ham or spam, where ham indicates non-spam clean tweets.
Sanaa Kaddoura, Safaa Henno
doaj   +1 more source

Recent Advancements in Topic Modeling Techniques for Healthcare, Bioinformatics, and Other Potential Applications

open access: yesAdvanced Intelligent Systems, Volume 8, Issue 1, January 2026.
This article offers a comprehensive review of topic modeling techniques, tracing their evolution from inception to recent developments. It explores methods such as latent Dirichlet allocation, latent semantic analysis, non‐negative matrix factorization, probabilistic latent semantic analysis, Top2Vec, and BERTopic, highlighting their strengths ...
Pratima Kumari   +6 more
wiley   +1 more source

PAAD: POLITICAL ARABIC ARTICLES DATASET FOR AUTOMATIC TEXT CATEGORIZATION

open access: yesIraqi Journal for Computers and Informatics, 2020
Now day’s text Classification and Sentiment analysis is considered as one of the popular Natural Language Processing (NLP) tasks. This kind of technique plays significant role in human activities and has impact on the daily behaviours.
Dhafar Hamed Abd   +2 more
doaj   +1 more source

AI‐Driven SMS Intelligence Framework for Detecting Fraud, Disinformation, and Threat Patterns in Pakistan’s Communication Networks

open access: yesApplied Computational Intelligence and Soft Computing, Volume 2026, Issue 1, 2026.
Fraudulent SMS messages are a significant threat in Pakistan, impacting financial security, identity theft, and misinformation and enabling exploitation by adversaries. This study proposes a localized AI‐based framework for SMS fraud detection, utilizing machine learning (logistic regression [LR], random forest [RF], support vector machine [SVM ...
Maryam Zaman   +8 more
wiley   +1 more source

Morphologically-analyzed and syntactically-annotated Quran datasetMendeley Data

open access: yesData in Brief
This paper introduces the Morphologically-Analyzed and Syntactically-Annotated Quran (MASAQ) dataset, a comprehensive resource designed to address the scarcity of annotated Quranic Arabic corpora and facilitate the development of advanced Natural ...
Majdi Sawalha   +6 more
doaj   +1 more source

Relative Chaoticity of Natural Languages

open access: yesComplexity, Volume 2026, Issue 1, 2026.
This paper presents a novel approach to analyzing and grouping natural languages based on the degree of their chaoticity. It clusters 52 languages from 18 language families, according to the value of the entropy–complexity pair, to reveal the chaotic properties of semantic trajectories.
Assel S. Yerbolova   +6 more
wiley   +1 more source

Home - About - Disclaimer - Privacy