Results 1 to 10 of about 187,791 (308)
A social and news media benchmark dataset for topic modeling [PDF]
Topic modeling is an active research area with several unanswered questions. The focus of recent research in this area is on the use of a vector embedding representation of the input text with both generative and evolutionary topic modeling techniques ...
Samuel Miles +4 more
doaj +4 more sources
Topic-driven Clustering for Document Datasets [PDF]
In this paper, we define the problem of topic-driven clustering, which organizes a document collection according to a given set of topics (either from domain experts, or as a requirement satisfying users' needs). We propose three topic-driven schemes that consider the similarity between the document to its topic and the relationship among the documents
Ying Zhao 0008, George Karypis
openaire +3 more sources
Comparison of Artificial Intelligence Tools With Human Coding for Sentiment, Topic, and Thematic Analysis Tasks of Public Health Datasets During the COVID-19 Pandemic in Australia: Case Study [PDF]
BackgroundPublic opinion, which may be influenced by personal experiences, news, and social media, can impact compliance with public health measures (PHMs) during health emergencies.
Danielle Hutchinson +5 more
doaj +2 more sources
Incorporating topical stance into signed bipartite networks for user retweet prediction. [PDF]
Social networks accelerate information dissemination, and retweet behavior is an important way of user interaction. User retweet prediction analyzes user characteristics to predict retweet behavior and emotional polarity, which can help platforms ...
Lixia Li +3 more
doaj +2 more sources
Analysis and tuning of hierarchical topic models based on Renyi entropy approach [PDF]
Hierarchical topic modeling is a potentially powerful instrument for determining topical structures of text collections that additionally allows constructing a hierarchy representing the levels of topic abstractness.
Sergei Koltcov +3 more
doaj +2 more sources
Operational Challenges in the Use of Structured Secondary Data for Health Research
Background: In Brazil, secondary data for epidemiology are largely available. However, they are insufficiently prepared for use in research, even when it comes to structured data since they were often designed for other purposes.
Kelsy N. Areco +15 more
doaj +1 more source
Purpose To construct a standard dataset of contrast-enhanced CT images of liver tumors to test the performance and safety of artificial intelligence (AI)-based algorithms for clinical decision support systems (CDSSs). Materials and Methods A consensus
Seung-seob Kim +6 more
doaj +1 more source
Topic2features: a novel framework to classify noisy and sparse textual data using LDA topic distributions [PDF]
In supervised machine learning, specifically in classification tasks, selecting and analyzing the feature vector to achieve better results is one of the most important tasks.
Junaid Abdul Wahid +6 more
doaj +2 more sources
TACO: Topics in Algorithmic COde generation dataset
We introduce TACO, an open-source, large-scale code generation dataset, with a focus on the optics of algorithms, designed to provide a more challenging training dataset and evaluation benchmark in the field of code generation models. TACO includes competition-level programming questions that are more challenging, to enhance or evaluate problem ...
Rongao Li +8 more
openaire +2 more sources
Hot topic trends have become increasingly important in the era of social media, as these trends can spread rapidly through online platforms and significantly impact public discourse and behavior.
Zohaib Ahmad Khan +6 more
doaj +1 more source

