gsufsort: constructing suffix arrays, LCP arrays and BWTs for string collections
Background The construction of a suffix array for a collection of strings is a fundamental task in Bioinformatics and in many other applications that process strings.
Felipe A. Louza +4 more
doaj +1 more source
Compression and information entropy of binary strings from the collision history of three hard balls
We investigate how to measure and define the entropy of a simple chaotic system, three hard spheres on a ring. A novel approach is presented, which does not assume the ergodic hypothesis.
M Vedak, G J Ackland
doaj +1 more source
Calculating Kolmogorov complexity from the output frequency distributions of small Turing machines. [PDF]
Drawing on various notions from theoretical computer science, we present a novel numerical approach, motivated by the notion of algorithmic probability, to the problem of approximating the Kolmogorov-Chaitin complexity of short strings.
Fernando Soler-Toscano +3 more
doaj +1 more source
Linear space string correction algorithm using the Damerau-Levenshtein distance
Background The Damerau-Levenshtein (DL) distance metric has been widely used in the biological science. It tries to identify the similar region of DNA,RNA and protein sequences by transforming one sequence to the another using the substitution, insertion,
Chunchun Zhao, Sartaj Sahni
doaj +1 more source
Efficient and effectiveness retrieval of information using some of the approximate string matching algorithms [PDF]
The research aims at buliding integral computer database for sales ,by using six algorithms of approximate string matching with practicable example;soundex,information metaphone,longest common subsequence ,dice cofficient,levenshtein distance and fuzzy ...
Anhar Mohammed, Suhiar Essa
doaj +1 more source
Permuted Pattern Matching Algorithms on Multi-Track Strings
A multi-track string is a tuple of strings of the same length. Given the pattern and text of two multi-track strings, the permuted pattern matching problem is to find the occurrence positions of all permutations of the pattern in the text. In this paper,
Diptarama Hendrian +4 more
doaj +1 more source
String correction using the Damerau-Levenshtein distance
Background In the string correction problem, we are to transform one string into another using a set of prescribed edit operations. In string correction using the Damerau-Levenshtein (DL) distance, the permissible edit operations are: substitution ...
Chunchun Zhao, Sartaj Sahni
doaj +1 more source
Toward Efficient Similarity Search under Edit Distance on Hybrid Architectures
Edit distance is the most widely used method to quantify similarity between two strings. We investigate the problem of similarity search under edit distance.
Madiha Khalid +2 more
doaj +1 more source
Kendall tau sequence distance: Extending Kendall tau from ranks to sequences [PDF]
An edit distance is a measure of the minimum cost sequence of edit operations to transform one structureinto another. Edit distance can be used as a measure of similarity as part of a pattern recognition system, withlower values of edit distance implying
Vincent Cicirello
doaj +1 more source
An efficient rank based approach for closest string and closest substring. [PDF]
This paper aims to present a new genetic approach that uses rank distance for solving two known NP-hard problems, and to compare rank distance with other distance measures for strings.
Liviu P Dinu, Radu Ionescu
doaj +1 more source

