Results 11 to 20 of about 187,791 (308)

An Adaptive LDA Optimal Topic Number Selection Method in News Topic Identification [PDF]

open access: yesIEEE Access, 2023
Nowadays, news text information is exploding, and people need more and more heterogeneous news content. Therefore, news text topic identification is needed to help viewers quickly and accurately screen and filter news related to their interests to save ...
Mingming Zheng   +3 more
doaj   +2 more sources

Diversity Over Size: On the Effect of Sample and Topic Sizes for Topic-Dependent Argument Mining Datasets [PDF]

open access: yesProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
The task of Argument Mining, that is extracting and classifying argument components for a specific topic from large document sources, is an inherently difficult task for machine learning models and humans alike, as large Argument Mining datasets are rare and recognition of argument components requires expert knowledge.
Benjamin Schiller   +3 more
core   +5 more sources

A Multi-Faceted Approach to Trending Topic Attack Detection Using Semantic Similarity and Large-Scale Datasets [PDF]

open access: yesIEEE Access
Twitter’s widespread popularity has made it a prime target for malicious actors exploiting trending hashtags to disseminate harmful content. This study marks the first systematic exploration of semantic consistency in tweets to detect trending ...
Insaf Kraidia   +2 more
doaj   +2 more sources

MasakhaNEWS: News Topic Classification for African languages [PDF]

open access: yes, 2023
African languages are severely under-represented in NLP research due to lack of datasets covering several NLP tasks. While there are individual language specific datasets that are being expanded to different tasks, only a handful of NLP tasks (e.g. named
Adeeko, Adetola   +64 more
core   +1 more source

Selection of the Optimal Number of Topics for LDA Topic Model—Taking Patent Policy Analysis as an Example

open access: yesEntropy, 2021
This study constructs a comprehensive index to effectively judge the optimal number of topics in the LDA topic model. Based on the requirements for selecting the number of topics, a comprehensive judgment index of perplexity, isolation, stability, and ...
Jingxian Gan, Yong Qi
doaj   +1 more source

Unsupervised Text Topic-Related Gene Extraction for Large Unbalanced Datasets [PDF]

open access: yes, 2020
There is a common notion that traditional unsupervised feature extraction algorithms follow the assumption that the distribution of the different clusters in a dataset is balanced.
Jing-Tao, Sun   +11 more
core   +2 more sources

Investigating topic bias in emotion classification [PDF]

open access: yes, 2023
In emotion classification, texts are assigned a conceptual emotion representation such as discrete labels or dimensions of cognitive appraisal. Emotion classifiers are typically not universally applicable, but base their classification decisions on ...
Wegge, Maximilian
core   +1 more source

Retrieval Topic Recurrent Memory Network for Remote Sensing Image Captioning

open access: yesIEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2020
Remote sensing image (RSI) captioning aims to generate sentences to describe the content of RSIs. Generally, five sentences are used to describe the RSI in caption datasets.
Binqiang Wang   +3 more
doaj   +1 more source

An Improved BiLSTM Approach for User Stance Detection Based on External Commonsense Knowledge and Environment Information

open access: yesApplied Sciences, 2022
In the age of social networks, the number of tweets sent by users has led to a sharp rise in public opinion. Public opinions are closely related to user stances. User stance detection has become an important task in the field of public opinion.
Peng Jia   +5 more
doaj   +1 more source

INTERACTIVE TOOL FOR VISUALIZATION OF TOPIC MODELS [PDF]

open access: yesActa Electrotechnica et Informatica, 2019
Digital data are all around us and occurs in various forms as videos, pictures or texts. Digital documents represent the vast majority of such data. It can be e-news, social media contributions and so on.
Miroslav SMATANA   +3 more
doaj   +1 more source

Home - About - Disclaimer - Privacy