Results 281 to 290 of about 28,490,308 (324)
Some of the next articles are maybe not open access.

Similar n-gram language model

Interspeech 2010, 2010
This paper describes an extension of the n-gram language model: the similar n-gram language model. The estimation of the probability P(s) of a string s by the classical model of order n is computed using statistics of occurrences of the last n words of the string in the corpus, whereas the proposed model further uses all the strings s' for which the ...
Gillot, Christian   +3 more
openaire   +2 more sources

KiloGrams: Very Large N-Grams for Malware Classification

arXiv.org, 2019
N-grams have been a common tool for information retrieval and machine learning applications for decades. In nearly all previous works, only a few values of $n$ are tested, with $n > 6$ being exceedingly rare.
Edward Raff   +6 more
semanticscholar   +1 more source

Polish N-Grams and Their Correction Process

2010 4th International Conference on Multimedia and Ubiquitous Engineering, 2010
Word n-gram statistics collected from over 1 300 000 000 words are presented. Eventhough they were collected from various good sources, they contain several types of errors. The paper focuses on the process of partly supervised correction of the n- grams. Types of errors are described as well as our software allowing efficient and fast corrections.
Bartosz Ziólko   +2 more
openaire   +2 more sources

Growing an n-gram language model

Interspeech 2005, 2005
Traditionally, when building an n-gram model, we decide the span of the model history, collect the relevant statistics and estimate the model. The model can be pruned down to a smaller size by manipulating the statistics or the estimated model. This paper shows how an n-gram model can be built by adding suitable sets of n-grams to a unigram model until
Vesa Siivola, Bryan L. Pellom
openaire   +2 more sources

Discriminative n-gram language modeling

Computer Speech & Language, 2007
This paper describes discriminative language modeling for a large vocabulary speech recognition task. We contrast two parameter estimation methods: the perceptron algorithm, and a method based on maximizing the regularized conditional log-likelihood. The models are encoded as deterministic weighted finite state automata, and are applied by intersecting
Brian Roark   +2 more
openaire   +2 more sources

Web N-gram workshop 2010

ACM SIGIR Forum, 2011
The Web N-gram Workshop was held on July 23, 2010 in Geneva, Switzerland, in conjunction with the 33rd Annual ACM SIGIR Conference. The workshop brought together leaders in information retrieval and language modeling to discuss the challenges in information retrieval and how language modeling approaches may help address some of these challenges, with a
Chengxiang Zhai   +4 more
openaire   +1 more source

A succinct N-gram language model

Proceedings of the ACL-IJCNLP 2009 Conference Short Papers on - ACL-IJCNLP '09, 2009
Efficient processing of tera-scale text data is an important research topic. This paper proposes lossless compression of N-gram language models based on LOUDS, a succinct data structure. LOUDS succinctly represents a trie with M nodes as a 2M + 1 bit string. We compress it further for the N-gram language model structure.
Taro Watanabe   +2 more
openaire   +2 more sources

Algorithmically generated malicious domain names detection based on n-grams features

Expert systems with applications, 2021
A. Cucchiarelli   +3 more
semanticscholar   +1 more source

Distributing N-Gram Graphs for Classification

2017
N-gram models have been an established choice for language modeling in machine translation, summarization and other tasks. Recently n-gram graphs managed to capture significant language characteristics that go beyond mere vocabulary and grammar, for tasks such as text classification.
Ioannis Kontopoulos   +2 more
openaire   +1 more source

An Indic Language N-gram Viewer

Proceedings of the 8th Annual Meeting of the Forum for Information Retrieval Evaluation, 2016
In this paper, we introduce culturomic studies on diachronic corpora for three languages -- Hindi, Bengali, and English, in the same spirit as that of the Google Culturomics Project [13]. Based on Hindi, Bengali, and English text from news blogs and newspapers, we are able to extract trajectories of salient words that are of importance in contemporary ...
Shanta Phani   +3 more
openaire   +1 more source

Home - About - Disclaimer - Privacy