S+ Trees Claim 40x Faster Search Than Binary Search
20 Jul 2026
A cache-aware layout aims to speed up sorted-data search
A newly published implementation of a static search tree (S+ tree) claims to dramatically outperform binary search on sorted data — by as much as 40x, according to the source blog post and accompanying code. The project also reports that an intermediate technique, the Eytzinger layout, runs roughly 4x faster than binary search on large arrays, positioning the S+ tree as a further leap beyond that approach.
The core idea is to design a search-tree layout that is explicitly aware of CPU cache hierarchies, rather than relying on the generic comparison-based binary search most systems use today.
The technical details behind the claim
The reported numbers are tied to specific hardware cache parameters used in testing:
- L3 cache size: 12MB
- L2 cache size: 256kB
- Hugepage size: 2MB
- Standard page size: 4kB
- Hugepage allocation threshold: 32MB (allocations below this size do not automatically become hugepages)
The approach leans on hugepages and cache-size-specific tuning to achieve its performance gains — a design choice that ties the results closely to the tested system's memory architecture.
Why genomics is the motivating use case
The project is explicitly motivated by the challenge of indexing DNA sequences — including full human genomes, which contain roughly 3 billion basepairs. The author frames the S+ tree as a first step toward speeding up suffix array searches, a data structure commonly used in genome indexing and sequence alignment. However, the suffix array search application itself is described as future work, meaning the genomics use case is not yet fully realized or benchmarked.
Source code for the implementation is publicly available at github:RagnarGrootKoerkamp/static-search-tree.
What's still unclear
Several open questions remain about the claims:
- The exact hardware and benchmark setup used to produce the 40x figure isn't specified in detail.
- The array or dataset sizes tested to obtain both the 40x and 4x results are not disclosed.
- There's no comparison to other modern search structures beyond binary search and the Eytzinger layout.
- No benchmarks specific to actual DNA or genome indexing workloads have been published yet — the genomics angle remains motivational context rather than demonstrated performance.
Risks to keep in mind
- Performance gains may be specific to the tested hardware and cache configuration, and may not generalize to other systems.
- Reliance on hugepages and cache-size-specific tuning could make the technique harder to port or maintain across different environments.
- The genome-indexing use case — the stated motivation for the project — is not yet fully realized, since suffix array search is still described as a future step.
Why founders should care
For founders building search-heavy or data-indexing infrastructure, this development could plausibly signal a path toward meaningfully faster throughput on large sorted datasets — particularly relevant for teams working with billions of records, such as in bioinformatics or large-scale search backends. The fact that the source code is openly available likely lowers the barrier for teams to prototype or evaluate the technique in-house without waiting for a packaged product.
That said, because the benchmarks are tied to a specific cache and hugepage configuration, founders should be cautious about assuming these speedups will transfer directly to their own production hardware. It would be prudent to validate performance on representative workloads before making architectural decisions based on these figures. Teams in genomics or sequence-search tooling may want to watch for follow-up work on the suffix array application, since that piece — the stated end goal of the project — has not yet been demonstrated.