Researchers Reduced Protein Database Index Memory
A new lossless compression algorithm lowers memory requirements by over 33 percent for large-scale protein sequence classification.
Updated on Sept. 24, 2026 in Biotech

Live Poll
Do you believe new data compression technologies will lead to significant improvements in public health research?
Researchers have developed a lossless compression algorithm that significantly shrinks the memory footprint required to index vast protein databases. This research-stage development enables the classification of sequences against a database of 250 billion amino acid characters using only 182 GB of memory.
Why it matters
By reducing the memory burden of indexing, this method allows for faster and more accessible taxonomic classification of biological sequences. It provides a more efficient alternative to existing tools for analyzing complex datasets, such as viral transcriptome profiles.
The new index occupies 182 GB of memory while covering the nr database of 250 billion amino acid characters, representing a reduction of over 33 percent compared to the Kaiju method. The algorithm also outperforms the Kraken2 classification method in accuracy.
The details
The method employs a new scheme for the run-block compression algorithm to reduce the physical size of the FM-index, a data structure that allows for fast full-text searching. This approach scales with alphabet size, maintaining efficiency as protein datasets expand. The Centrifuger method uses this index to perform taxonomic classification and identify viral transcriptome profiles in human cells.
Timeline
September 2026: The paper describing the algorithm was published.
The Tech Race
The study advances sequence classification performance relative to the nr protein sequence database, which serves as a global benchmark for protein search efficiency. This development updates the performance standards for analyzing large biological datasets by lowering the hardware requirements compared to existing tools like Kaiju and Kraken2.
Researchers working with large-scale genomic or proteomic datasets can expect to see reduced hardware requirements when performing sequence classification. The tool allows for the analysis of complex viral profiles on more modest computing clusters than previously required.
The takeaway
This algorithm demonstrates that sophisticated compression schemes can drastically improve the accessibility of large biological databases. Future researchers should monitor subsequent performance benchmarks for this method against even larger metagenomic datasets.
Further reading
For more on the latest computational tools in biology, explore our Biotech section.
More information
Read the complete research paper and algorithm details.
Live Poll
Do you believe new data compression technologies will lead to significant improvements in public health research?






