Results 281 to 290 of about 28,490,308 (324)
Some of the next articles are maybe not open access.
Interspeech 2010, 2010
This paper describes an extension of the n-gram language model: the similar n-gram language model. The estimation of the probability P(s) of a string s by the classical model of order n is computed using statistics of occurrences of the last n words of the string in the corpus, whereas the proposed model further uses all the strings s' for which the ...
Gillot, Christian +3 more
openaire +2 more sources
This paper describes an extension of the n-gram language model: the similar n-gram language model. The estimation of the probability P(s) of a string s by the classical model of order n is computed using statistics of occurrences of the last n words of the string in the corpus, whereas the proposed model further uses all the strings s' for which the ...
Gillot, Christian +3 more
openaire +2 more sources
KiloGrams: Very Large N-Grams for Malware Classification
arXiv.org, 2019N-grams have been a common tool for information retrieval and machine learning applications for decades. In nearly all previous works, only a few values of $n$ are tested, with $n > 6$ being exceedingly rare.
Edward Raff +6 more
semanticscholar +1 more source
Polish N-Grams and Their Correction Process
2010 4th International Conference on Multimedia and Ubiquitous Engineering, 2010Word n-gram statistics collected from over 1 300 000 000 words are presented. Eventhough they were collected from various good sources, they contain several types of errors. The paper focuses on the process of partly supervised correction of the n- grams. Types of errors are described as well as our software allowing efficient and fast corrections.
Bartosz Ziólko +2 more
openaire +2 more sources
Growing an n-gram language model
Interspeech 2005, 2005Traditionally, when building an n-gram model, we decide the span of the model history, collect the relevant statistics and estimate the model. The model can be pruned down to a smaller size by manipulating the statistics or the estimated model. This paper shows how an n-gram model can be built by adding suitable sets of n-grams to a unigram model until
Vesa Siivola, Bryan L. Pellom
openaire +2 more sources
Discriminative n-gram language modeling
Computer Speech & Language, 2007This paper describes discriminative language modeling for a large vocabulary speech recognition task. We contrast two parameter estimation methods: the perceptron algorithm, and a method based on maximizing the regularized conditional log-likelihood. The models are encoded as deterministic weighted finite state automata, and are applied by intersecting
Brian Roark +2 more
openaire +2 more sources
ACM SIGIR Forum, 2011
The Web N-gram Workshop was held on July 23, 2010 in Geneva, Switzerland, in conjunction with the 33rd Annual ACM SIGIR Conference. The workshop brought together leaders in information retrieval and language modeling to discuss the challenges in information retrieval and how language modeling approaches may help address some of these challenges, with a
Chengxiang Zhai +4 more
openaire +1 more source
The Web N-gram Workshop was held on July 23, 2010 in Geneva, Switzerland, in conjunction with the 33rd Annual ACM SIGIR Conference. The workshop brought together leaders in information retrieval and language modeling to discuss the challenges in information retrieval and how language modeling approaches may help address some of these challenges, with a
Chengxiang Zhai +4 more
openaire +1 more source
A succinct N-gram language model
Proceedings of the ACL-IJCNLP 2009 Conference Short Papers on - ACL-IJCNLP '09, 2009Efficient processing of tera-scale text data is an important research topic. This paper proposes lossless compression of N-gram language models based on LOUDS, a succinct data structure. LOUDS succinctly represents a trie with M nodes as a 2M + 1 bit string. We compress it further for the N-gram language model structure.
Taro Watanabe +2 more
openaire +2 more sources
Algorithmically generated malicious domain names detection based on n-grams features
Expert systems with applications, 2021A. Cucchiarelli +3 more
semanticscholar +1 more source
Distributing N-Gram Graphs for Classification
2017N-gram models have been an established choice for language modeling in machine translation, summarization and other tasks. Recently n-gram graphs managed to capture significant language characteristics that go beyond mere vocabulary and grammar, for tasks such as text classification.
Ioannis Kontopoulos +2 more
openaire +1 more source
An Indic Language N-gram Viewer
Proceedings of the 8th Annual Meeting of the Forum for Information Retrieval Evaluation, 2016In this paper, we introduce culturomic studies on diachronic corpora for three languages -- Hindi, Bengali, and English, in the same spirit as that of the Google Culturomics Project [13]. Based on Hindi, Bengali, and English text from news blogs and newspapers, we are able to extract trajectories of salient words that are of importance in contemporary ...
Shanta Phani +3 more
openaire +1 more source

