Abstract
Spaced seeds are important tools for similarity search in bioinformatics, and using several seeds together often significantly improves their performance. With existing approaches, however, for each seed we keep a separate linear-size data structure, either a hash table or a spaced suffix array (SSA). In this paper we show how to compress SSAs relative to normal suffix arrays (SAs) and still support fast random access to them. We first prove a theoretical upper bound on the space needed to store an SSA when we already have the SA. We then present experiments indicating that our approach works even better in practice.
| Lingua originale | Inglese |
|---|---|
| pagine (da-a) | 37-45 |
| Numero di pagine | 9 |
| Rivista | CEUR Workshop Proceedings |
| Volume | 1146 |
| Stato di pubblicazione | Pubblicato - 2014 |
| Pubblicato esternamente | Sì |
| Evento | 2nd International Conference on Algorithms for Big Data, ICABD 2014 - Palermo, Italy Durata: 7 apr 2014 → 9 apr 2014 |
Fingerprint
Entra nei temi di ricerca di 'Compressed spaced suffix arrays'. Insieme formano una fingerprint unica.Cita questo
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver