Results 11 to 20 of about 136,207 (260)
Multilingual Search with Subword TF-IDF
Multilingual search can be achieved with subword tokenization. The accuracy of traditional TF-IDF approaches depend on manually curated tokenization, stop words and stemming rules, whereas subword TF-IDF (STF-IDF) can offer higher accuracy without such heuristics.
Wangperawong, Artit
openaire +3 more sources
Deriving TF-IDF as a Fisher Kernel [PDF]
The Dirichlet compound multinomial (DCM) distribution has recently been shown to be a good model for documents because it captures the phenomenon of word burstiness, unlike standard models such as the multinomial distribution. This paper investigates the DCM Fisher kernel, a function for comparing documents derived from the DCM.
Charles Elkan
openaire +2 more sources
Probabilistic retrieval models - relationships, context-specific application, selection and implementation [PDF]
PhDRetrieval models are the core components of information retrieval systems, which guide the document and query representations, as well as the document ranking schemes.
Wang, Jun
core +4 more sources
Analysis of network public opinion can help to effectively predict the public emotion and the multi-level government behaviors. Due to the massive and multidimensional characteristics of network public opinion data, the in-depth value mining of public ...
Hangfeng Lin, Naiqing Bu
doaj +1 more source
The problem of classification of scientific articles is considered.
N. O. Ukhanov +3 more
doaj +1 more source
Hadis adalah sumber rujukan agama Islam kedua setelah Al-Qur’an. Teks Hadis saat ini diteliti dalam bidang teknologi untuk dapat ditangkap nilai-nilai yang terkandung di dalamnya secara pegetahuan teknologi. Dengan adanya penelitian terhadap Kitab Hadis,
Ana Tsalitsatun Ni'mah +1 more
doaj +1 more source
Sentiment analysis is the extraction and categorization of sentiments that have been expressed in text data using text analysis techniques. Manifested by earlier studies, sentiment analysis of drug reviews has a large potential for providing valuable ...
Eysha Saad +6 more
doaj +1 more source
USTW Vs. STW: A Comparative Analysis for Exam Question Classification based on Bloom’s Taxonomy
Bloom’s Taxonomy (BT) is widely used in educational institutions to produce high-quality exam papers to evaluate students’ knowledge at different cognitive levels.
Mohammed Osman Gani +3 more
doaj +1 more source
Uzbek text summarization based on TF-IDF
The volume of information is increasing at an incredible rate with the rapid development of the Internet and electronic information services. Due to time constraints, we don't have the opportunity to read all this information. Even the task of analyzing textual data related to one field requires a lot of work. The text summarization task helps to solve
Khabibulla Madatov +2 more
openaire +3 more sources
A comparative study of keyword extraction algorithms for English texts
This study mainly analyzed the keyword extraction of English text. First, two commonly used algorithms, the term frequency–inverse document frequency (TF–IDF) algorithm and the keyphrase extraction algorithm (KEA), were introduced.
Li Jinye
doaj +1 more source

