Results 51 to 60 of about 1,363 (163)
AraReasoner: Evaluating Reasoning-Based LLMs for Arabic NLP
Large language models (LLMs) have shown remarkable progress in reasoning abilities and general natural language processing (NLP) tasks, yet their performance on Arabic data, characterized by rich morphology, diverse dialects, and complex script, remains underexplored.
Ahmed Hasanaath +5 more
openaire +3 more sources
Abstract Our research investigated how L2 and L1 reading, L1 low‐level skills and working memory are related to ratings and the linguistic characteristics (productivity, cohesion, lexical sophistication and diversity, syntactic complexity, and accuracy) of argumentative and narrative texts. The research was conducted in Hungary with 95 secondary school
Judit Kormos, Csilla Bartha
wiley +1 more source
PROCESSING TOOLS FOR CORPUS LINGUISTICS: A CASE STUDY ON ARABIC HISTORICAL CORPUS
This paper explores the development, design, and reconstruction of a Historical Arabic Corpus (HAC), which covers more than 1600 years of uninterrupted language use.
Bassam Hasan Hammo, Sane Yagi
doaj +1 more source
ABSTRACT Negative transfer continues to limit the benefits of multi‐task learning (MTL) in harmful language detection, where related tasks must share representations without diluting task‐specific nuances. We introduce task awareness (TA), a methodological framework that explicitly conditions MTL models on the task they must solve.
Angel Felipe Magnossão de Paula +3 more
wiley +1 more source
Dataset of Arabic spam and ham tweets
This data article provides a dataset of 132421 posts and their corresponding information collected from Twitter social media. The data has two classes, ham or spam, where ham indicates non-spam clean tweets.
Sanaa Kaddoura, Safaa Henno
doaj +1 more source
This article offers a comprehensive review of topic modeling techniques, tracing their evolution from inception to recent developments. It explores methods such as latent Dirichlet allocation, latent semantic analysis, non‐negative matrix factorization, probabilistic latent semantic analysis, Top2Vec, and BERTopic, highlighting their strengths ...
Pratima Kumari +6 more
wiley +1 more source
PAAD: POLITICAL ARABIC ARTICLES DATASET FOR AUTOMATIC TEXT CATEGORIZATION
Now day’s text Classification and Sentiment analysis is considered as one of the popular Natural Language Processing (NLP) tasks. This kind of technique plays significant role in human activities and has impact on the daily behaviours.
Dhafar Hamed Abd +2 more
doaj +1 more source
Fraudulent SMS messages are a significant threat in Pakistan, impacting financial security, identity theft, and misinformation and enabling exploitation by adversaries. This study proposes a localized AI‐based framework for SMS fraud detection, utilizing machine learning (logistic regression [LR], random forest [RF], support vector machine [SVM ...
Maryam Zaman +8 more
wiley +1 more source
Morphologically-analyzed and syntactically-annotated Quran datasetMendeley Data
This paper introduces the Morphologically-Analyzed and Syntactically-Annotated Quran (MASAQ) dataset, a comprehensive resource designed to address the scarcity of annotated Quranic Arabic corpora and facilitate the development of advanced Natural ...
Majdi Sawalha +6 more
doaj +1 more source
Relative Chaoticity of Natural Languages
This paper presents a novel approach to analyzing and grouping natural languages based on the degree of their chaoticity. It clusters 52 languages from 18 language families, according to the value of the entropy–complexity pair, to reveal the chaotic properties of semantic trajectories.
Assel S. Yerbolova +6 more
wiley +1 more source

