Results 21 to 30 of about 1,556 (153)

Designing small universal k-mer hitting sets for improved analysis of high-throughput sequencing. [PDF]

open access: yesPLoS Computational Biology, 2017
With the rapidly increasing volume of deep sequencing data, more efficient algorithms and data structures are needed. Minimizers are a central recent paradigm that has improved various sequence analysis tasks, including hashing for faster read overlap ...
Yaron Orenstein   +4 more
doaj   +1 more source

Fast mapping of short sequences with mismatches, insertions and deletions using index structures. [PDF]

open access: yesPLoS Computational Biology, 2009
With few exceptions, current methods for short read mapping make use of simple seed heuristics to speed up the search. Most of the underlying matching models neglect the necessity to allow not only mismatches, but also insertions and deletions.
Steve Hoffmann   +7 more
doaj   +1 more source

Accelerated preprocessing in task of searching substrings in a string

open access: yesAdvanced Engineering Research, 2019
Introduction. A rapid development of the systems such as Yandex, Google, etc., has predetermined the relevance of the task of searching substrings in a string, and approaches to its solution are actively investigated. This task is used to create database
A. V. Mazurenko, N. V. Boldyrikhin
doaj   +1 more source

These are not the k-mers you are looking for: efficient online k-mer counting using a probabilistic data structure. [PDF]

open access: yesPLoS ONE, 2014
K-mer abundance analysis is widely used for many purposes in nucleotide sequence analysis, including data preprocessing for de novo assembly, repeat detection, and sequencing coverage estimation.
Qingpeng Zhang   +4 more
doaj   +1 more source

Sampling the Suffix Array with Minimizers [PDF]

open access: yes, 2015
Sampling (evenly) the suffixes from the suffix array is an old idea trading the pattern search time for reduced index space. A few years ago Claude et al. showed an alphabet sampling scheme allowing for more efficient pattern searches compared to the sparse suffix array, for long enough patterns.
Szymon Grabowski, Marcin Raniszewski
openaire   +2 more sources

Gclust: A Parallel Clustering Tool for Microbial Genomic Data

open access: yesGenomics, Proteomics & Bioinformatics, 2019
The accelerating growth of the public microbial genomic data imposes substantial burden on the research community that uses such resources. Building databases for non-redundant reference sequences from massive microbial genomic data based on clustering ...
Ruilin Li   +17 more
doaj   +1 more source

Computing Maximal Lyndon Substrings of a String

open access: yesAlgorithms, 2020
There are two reasons to have an efficient algorithm for identifying all right-maximal Lyndon substrings of a string: firstly, Bannai et al. introduced in 2015 a linear algorithm to compute all runs of a string that relies on knowing all right-maximal ...
Frantisek Franek, Michael Liut
doaj   +1 more source

Approximate String Matching with Compressed Indexes

open access: yesAlgorithms, 2009
A compressed full-text self-index for a text T is a data structure requiring reduced space and able to search for patterns P in T. It can also reproduce any substring of T, thus actually replacing T. Despite the recent explosion of interest on compressed
Pedro Morales   +3 more
doaj   +1 more source

Dynamic extended suffix arrays

open access: yesJournal of Discrete Algorithms, 2010
zbMATH Open Web Interface contents unavailable due to conflicting licenses.
Salson, Mikael   +3 more
openaire   +5 more sources

Counting Suffix Arrays and Strings

open access: yesTheoretical Computer Science, 2005
zbMATH Open Web Interface contents unavailable due to conflicting licenses.
Schürmann, Klaus-Bernd, Stoye, Jens
openaire   +3 more sources

Home - About - Disclaimer - Privacy