BORSUK · Docs

Research & Analysis

How the internals behave when you push on the knobs.

Canonical standard-dataset evidence, every leaf method, configuration ablations, CPU/RAM/disk timelines, scale studies, external comparisons, and publication positioning live here. The default docs contain only production architecture and guidance.

Read this first

Evidence classes stay separate

Current qualification is in progress. The exact-vector storage boundary changed to standard typed Arrow IPC record batches on 23 July 2026. Every earlier descriptor-v8 latency, CPU, RAM, disk, request, and cost row is therefore historical and cannot support the current defaults. Fresh empty-prefix runs will be promoted only after they satisfy the expanded DBpedia/Cohere/LAION/MS MARCO/filter/multitenant/hybrid/late-interaction matrix.

Full-corpus AWS results, controlled local method experiments, synthetic scale studies, direct external measurements, and vendor-reported context are never merged into one ranking. Every chart is backed by a checked-in CSV, and missing method/dataset cells remain explicitly unmeasured. The canonical reviewable index is docs/research/.

Canonical corpus

What has been evaluated

Historical corpus: Fashion-MNIST, GloVe, SIFT, NYTimes, GIST, and Deep-Image retain descriptor-v8 source recreations and archived curves for regression research. Current matrix: DBpedia OpenAI 1M, Cohere 1M/10M, LAION 100M, MS MARCO V2 138M, filters, multitenancy, BEIR hybrid, ColBERT late interaction, and mixed cache coverage are now required before publication.

Parquet versus Vortex

Corrected real-segment replay keeps Parquet as the default

The first Vortex latency comparison was invalid because it stopped before materializing compressed Vortex arrays to the Arrow values consumed downstream. The corrected AWS replay covers all 17 normal-segment objects from a real Fashion-MNIST-784 index, with 3 warmups and 30 measured repetitions for every object, format, and workload. All 7,650 timed rows use the same materialized_arrow boundary.

Vortex coalesced selective native-S3 scans and won projection, point, range, and filter workloads in that historical reader. Parquet won full materialization and was much smaller: 96.84 MiB versus 253.78 MiB for Vortex default and 319.56 MiB for compact. That selective advantage was not the production query path, and the later frozen object-role qualifications rejected promotion. The unreleased backend and its dependency graph have since been removed; this section remains revision-bound research evidence.

Corrected real BORSUK segment storage footprint and materialized Arrow latency distributions for Parquet and Vortex
17 real segments · mean ± standard deviation · p50/p95/p99 · no compressed-result timing
Combined corrected format replay CPU RAM disk and network resource timeline
Combined campaign resources · 1.287 GiB peak RSS · not attributed to one query format

The aggregate CSVs and full methodology live in the table-format A/B report. Vortex compact is not a default candidate on this schema: it is 3.30× the Parquet footprint and its pooled full-scan p95 is 637.8 ms versus 129.8 ms for Parquet.

Public ANN corpora

Archived pre-Arrow-sidecar results

Every table row in this section predates the standard typed Arrow IPC sidecar and is historical. It remains visible for auditability, but none is a current production/default claim. The replacement runs report startup separately from uncached, zero-backing-GET disk-cached, and 0/25/50/75/100% mixed-coverage p50/p95/p99/mean/stddev.

Guarantee boundary: production pq-scan reports empirical recall on each complete query set. Exact/guaranteed_recall is the formal 1.0 path and remains RAM-bounded, but it may read and score the full eligible corpus when safe lower bounds cannot prune. The NYTimes control measured 405.6 MB/query and 5.19 s disk-cached p95 for exact search; an empirical 1.000 ANN result is never relabelled as guaranteed.

Datasetrecall@10uncached p95disk-cached p95peak RSS
Fashion-MNIST · descriptor-v6 first pass0.988178.2 ms1.40 ms276 MiB build+sweep
Fashion-MNIST · descriptor-v8 packed default · 3-run median0.989197.9 ms1.667 ms75.0 MiB worst serving
GloVe · descriptor-v7 adaptive rebuild0.985495.1 ms13.3 ms350 MiB build+sweep
NYTimes · descriptor-v7 adaptive rebuild0.958701.8 ms9.3 ms262 MiB build+sweep
GloVe · descriptor-v8 packed rebuild1.000*448.5 ms25.0 ms360 MiB build+sweep
GloVe · descriptor-v8 packed default · 3-run median1.000*467.3 ms25.0 ms148.8 MiB worst serving
NYTimes · descriptor-v8 packed rebuild0.958312.5 ms12.3 ms280 MiB build+sweep
NYTimes · descriptor-v8 code128 default · 3-run median0.993368.9 ms19.7 ms118.9 MiB worst serving
SIFT · descriptor-v8 packed default · 3-run median0.997213.8 ms4.58 ms104.9 MiB worst serving
GIST · descriptor-v8 code256 default · 3-run median0.995427.3 ms29.2 ms304.3 MiB worst serving
Deep-Image · descriptor-v8 default · 3-run median0.990508.2 ms26.0 ms271.9 MiB worst serving
Deep-Image · descriptor-v8 hierarchy balanced0.987403.2 ms21.1 ms461 MiB build+sweep
Deep-Image · descriptor-v8 hierarchy high recall0.998889.0 ms57.0 ms461 MiB build+sweep
Fashion-MNIST · historical pre-fixed-page v80.989191.1 ms2.45 ms229 MiB median
Fashion-MNIST · historical v7 global PQ0.97264.9 ms2.49 ms99.2 MiB
GloVe · historical v7 global PQ0.95693.5 ms7.08 ms262 MiB
SIFT · historical v7 global PQ0.97666.1 ms5.41 ms240 MiB
NYTimes · historical v60.959477.3 ms78.3 ms759 MiB
GIST · historical v60.967543.7 ms75.2 ms724 MiB
Deep-Image · historical v60.956368.0 ms45.7 ms712 MiB

These rows use the intended cache definitions but the superseded physical format. The old Deep-Image 41.9/44.6/46.4 ms row was disk-cached p50/p95/p99, not an S3-uncached result.

Quality versus time

Recall/latency curves for every production corpus

Each point changes nprobe while keeping the dataset and engine fixed. Higher recall usually requires more cells, bytes, GETs, and exact rerank work; the selected production point is the lowest tested configuration at or above recall 0.95, not the fastest point at arbitrary quality.

Fresh descriptor-v7 adaptive GloVe recall versus uncached and disk-cached p95
GloVe · descriptor-v7 adaptive flat-256 · fresh rebuild diagnostic
Fresh descriptor-v8 packed GloVe recall versus uncached and disk-cached p95
GloVe · descriptor-v8 packed flat-256 · fewer S3 GETs
Fresh descriptor-v6 Fashion-MNIST hierarchical recall versus latency curve
Fashion-MNIST · descriptor-v6 hierarchy · first-pass diagnostic
Fresh descriptor-v8 packed Fashion-MNIST recall versus uncached and disk-cached p95
Fashion-MNIST · descriptor-v8 packed adaptive hierarchy · 0.989 default recall
Rejected 4096-leaf hierarchical GloVe recall versus latency curve
GloVe · rejected 4,096-leaf hierarchy · fast pages, insufficient recall
Diagnostic fixed-page Deep-Image recall versus uncached and disk-cached p95
Deep-Image · fixed pages/code64 · diagnostic
Fresh descriptor-v8 hierarchical Deep-Image recall versus uncached and disk-cached p95
Deep-Image · descriptor-v8 full-dimensional hierarchy · 9.99M rows
Fresh descriptor-v8 GIST code256 recall versus uncached and disk-cached p95
GIST · descriptor-v8 code256 routing curve · up to 0.999 empirical recall
GIST code256 24-probe candidate recall and latency curve
GIST · candidate curve · selected 96-row exact shortlist
Deep-Image 512-probe candidate-width recall versus latency control
Deep-Image · 512 probes · candidate widening rejected at a 0.999 plateau
Fresh descriptor-v7 adaptive NYTimes recall versus uncached and disk-cached p95
NYTimes · descriptor-v7 adaptive flat-256 · fresh rebuild diagnostic
Fresh descriptor-v8 packed NYTimes recall versus uncached and disk-cached p95
NYTimes · descriptor-v8 packed flat-256 · fewer S3 GETs
Rejected NYTimes code256 candidate recall versus latency curve
NYTimes · rejected code256 · 0.993 plateau through 1,024 candidates
Fresh descriptor-v8 packed SIFT recall versus uncached and disk-cached p95
SIFT · descriptor-v8 packed adaptive hierarchy · 0.997 default recall
Fashion-MNIST recall versus disk-cached latency curve
Fashion-MNIST · 784D
GloVe recall versus disk-cached latency curve
GloVe · 100D
SIFT recall versus disk-cached latency curve
SIFT · 128D
NYTimes recall versus disk-cached latency curve
NYTimes · 256D
GIST recall versus disk-cached latency curve
GIST · 960D
Deep-Image recall versus disk-cached latency curve
Deep-Image · 96D
NYTimes v7 global PQ candidate recall versus uncached latency
NYTimes · v7 global PQ candidate sweep
GIST v7 64-subspace global PQ candidate recall curve with latency contaminated by concurrent build
GIST · v7 64-subspace recall sweep; latency not promotable because a Deep-Image build overlapped

The packed descriptor-v8 Deep-Image recreation closes the large-angular routing question. The full-dimensional hierarchy reaches 0.987 recall at 96 probes with 403/21 ms uncached/disk-cached p95, 0.994 at 192 with 568/34 ms, and 0.998 at 384 with 889/57 ms. The 512 point reaches 0.999 but crosses 1.07 s and 126 MB/query. Holding 512 probes and widening exact candidates from 256 to 512 does not change recall; it raises uncached p95 from 1.10 to 1.24 s and GETs from 344 to 393. Build plus sweep peaks at about 461 MiB RSS and 4.20 GiB scratch, which is deleted after publication. The older product router needed 983 ms to reach only 0.993, so product routing remains a rejected ablation and large normalized corpora now use the bounded full-dimensional hierarchy. Exact mode remains the only formal 1.0 guarantee.

Experimental graph path

From 1.95 seconds to 28 milliseconds at unchanged recall

On the same graph-enabled Fashion-MNIST index and all 100 queries, a score-once best-first frontier reduced graph p95 from 1,951.5 to 208.8 ms. Sharing the immutable decoded/validated graph in the byte-accounted cell cache reduced it to 28.5 ms. Moving graph preparation into warm() removed the first-query metadata tail: 28.0 ms p95, 28.5 ms p99, zero measured-query GETs, and the same 0.970 recall at width 512.

The recall-matched point is wider: nprobe=32, 2,560 candidates/cell reaches 0.986 recall. Across three fresh repetitions its p95 is 56.4–57.7 ms (mean 56.9), versus pq-scan at the same 0.986 recall and 89.2 ms, and Vamana-PQ at 0.988 recall and 96.2–96.4 ms. These are memory-preloaded experimental graph results on one corpus—not a replacement for the six-corpus graph-free production default.

MethodCells / candidatesRecall@10p95Worst RSSQuery GETs
graph32 / 2,5600.98656.4–57.7 ms378.5 MiB0
pq-scan22 / 120.98689.2 ms379.0 MiB0
Vamana-PQ22 / 320.98896.2–96.4 ms379.0 MiB0
Optimized Fashion-MNIST graph recall versus p95 latency curve
Graph recall/latency sweep · all 100 queries
Optimized graph CPU RAM disk and cache timeline
Recall-matched graph · repetition 2
PQ scan CPU RAM disk and cache timeline on the graph-enabled comparison index
Recall-matched pq-scan · repetition 2
Vamana PQ CPU RAM disk and cache timeline
Recall-matched Vamana-PQ · repetition 2

The full-coverage 4,096/8,192 candidate rows intentionally short-circuit graph traversal and exact-scan the routed cells, so their non-monotonic latency is labelled as a fallback. A prior 0.990/0.995 claim at widths 512/1,024 came from 20 profiling queries; the publication curve above uses all 100. Method, profile, and raw-artifact details →

Scale baseline · AWS eu-central-1

100M vectors stay memory-bounded; S3 latency is the next target

The first complete 100M×96D angular recreation reached 0.992 recall at 64 probes with 910.5 ms uncached and 44.4 ms disk-cached p95. Doubling probes to 128 and 256 did not change recall; it raised uncached p95 to 1.24/1.70 s and disk-cached p95 to 79/146 ms, so both rows are rejected.

The counters isolate the failure: 76.19 MB/query is the logical code span, while only 0.48 MB/query is additional uncached exact-rerank and final-record traffic. The v8 packed reader fetched from the first to the last selected slice in an object, including unrelated cells between them. A balanced 64-probe payload is about 26.56 MB before headers, so the measured span was 2.87× that lower bound. The v9 reader uses bounded, gap-aware ranges and never folds more than 64 KiB of unselected data into one code read; promotion waits for its fresh lower-probe AWS curve.

This historical result is not current evidence. The present pq-scan, srht-pq-scan, structured fast-turboquant-scan, and dense TurboQuant reference controls are rebuilt and labeled separately.

100M vector SRHT-PQ recall versus uncached and disk-cached p95
100M×96D · recall plateaus at 64 probes
100M vector build and query CPU RAM disk scratch and cache timeline
4.16-hour whole-run resource envelope

Cache-tier research · local synthetic control

Requested hot queries are not assumed to be cache hits

The cache benchmark records each query's requested class (hot or outside_hot_set) separately from its observed decoded-RAM, local-disk, and backing-storage access fractions. Shared cells and LRU eviction mean those are different facts: an outside query may reuse a hot cell, while a seeded query may miss after eviction.

On a fixed 100k-vector, 96D, 25-cell synthetic graph index, 128 MiB retained only 85% of accesses for the 100%-requested-hot mix and measured 117.3 ms steady p95. The smallest tested complete cap was 256 MiB: 100% decoded access, 0.285 ms steady p95, 1.831 ms p95 with 16 callers, and 227.3 MiB peak RSS. A 512 MiB cap did not improve that resident set. Queries outside the seeded set still paid a 112–129 ms p95 miss, so publication charts show all-query, hot-query, and outside-hot-set tails separately.

The matching auto control retained only 12/26 segment graphs at 64 MiB, reported coverage_complete=false, and served through srht-pq-scan without failing. At 256 MiB it retained 26/26, reported complete coverage, and selected the graph. These are local controls; the six-corpus AWS matrix must reproduce the envelope before any cache size becomes a default.

Whole-process telemetry

CPU, RAM, process disk, and cache growth

These production traces include the embedded engine, language/runtime overhead, query buffers, and cache—not just an artifact-size estimate. Uncapped multi-user traces remain a separate throughput-ceiling experiment. The v7 GloVe trace uses four admitted searches and the shared exact-rerank read gate.

Fresh descriptor-v7 GloVe build and recall sweep CPU RAM disk and cache timeline
GloVe · descriptor-v7 adaptive rebuild · 350 MiB build+sweep peak RSS
Fresh descriptor-v7 NYTimes build and recall sweep CPU RAM disk and cache timeline
NYTimes · descriptor-v7 adaptive rebuild · 262 MiB build+sweep peak RSS
Fresh descriptor-v8 packed GloVe CPU RAM scratch disk and cache timeline
GloVe · descriptor-v8 packed rebuild · 360 MiB peak RSS
Fresh descriptor-v8 packed NYTimes CPU RAM scratch disk and cache timeline
NYTimes · descriptor-v8 packed rebuild · 280 MiB peak RSS
Rejected NYTimes code256 build and sweep CPU RAM scratch disk and cache timeline
NYTimes · rejected code256 · 284 MiB peak RSS, doubled scan bytes
Rejected GloVe product-2x64 build and sweep CPU RAM scratch disk and cache timeline
GloVe · rejected product-2x64 · 335 MiB peak RSS, insufficient recall
Rejected NYTimes product-2x64 build and sweep CPU RAM scratch disk and cache timeline
NYTimes · rejected product-2x64 · 257 MiB peak RSS, insufficient recall
Diagnostic fixed-page GloVe width-16 CPU RAM process disk and cache timeline
GloVe · fixed-page width-16 query-only diagnostic · 139 MiB peak RSS
Fresh descriptor-v6 Fashion-MNIST CPU RAM scratch disk and cache timeline
Fashion-MNIST · descriptor-v6 first pass · 276 MiB build+sweep peak RSS
Fresh descriptor-v8 Fashion-MNIST CPU RAM scratch disk and cache timeline
Fashion-MNIST · descriptor-v8 packed rebuild · 258 MiB peak RSS · 163 MiB scratch
Fashion-MNIST descriptor-v8 selected production repetition CPU RAM process disk and cache timeline
Fashion-MNIST · selected production repetition · 73 MiB peak RSS
Rejected 4096-leaf GloVe hierarchy CPU RAM scratch disk and cache timeline
GloVe · rejected 4,096-leaf hierarchy · 324 MiB build-plus-query peak RSS
Deep-Image fixed-page build CPU RAM scratch disk and cache timeline
Deep-Image · fresh 9.99M build · 349 MiB peak RSS · 4.21 GiB sampled scratch
Deep-Image descriptor-v8 hierarchical build and recall sweep CPU RAM scratch disk and cache timeline
Deep-Image · descriptor-v8 hierarchy · 457 MiB peak RSS · 4.20 GiB scratch
Deep-Image candidate-width control CPU RAM process disk and cache timeline
Deep-Image · candidate-width control · 256–384 candidates
Deep-Image isolated 512-candidate control CPU RAM process disk and cache timeline
Deep-Image · candidate-width control · isolated 512 candidates
SIFT descriptor-v8 build and recall sweep CPU RAM scratch disk and cache timeline
SIFT · descriptor-v8 packed rebuild · 306 MiB peak RSS · 538 MiB scratch
SIFT descriptor-v8 selected production repetition CPU RAM process disk and cache timeline
SIFT · selected production repetition · 105 MiB peak RSS
Deep-Image descriptor-v8 selected production repetition CPU RAM process disk and cache timeline
Deep-Image · selected production repetition · 263 MiB peak RSS
Historical pre-fixed-page vector-level IVF Fashion-MNIST CPU RAM scratch disk and cache timeline
Fashion-MNIST · historical pre-fixed-page v8 · fresh-prefix repetition r6
Superseded vector-level IVF Fashion-MNIST CPU RAM scratch disk and cache timeline
Fashion-MNIST · vector-level IVF r1 · recall-qualified, resource profile superseded
Rejected 4096-cell vector-level IVF GloVe CPU RAM scratch disk and cache timeline
GloVe · rejected 4,096-cell scale topology · 0.709 recall at 1.03 s p95
Rejected checkpoint-centroid v8 Fashion-MNIST CPU RAM disk and cache timeline
Fashion-MNIST · rejected checkpoint-centroid router · diagnostic only
Rejected checkpoint-centroid v8 GloVe CPU RAM disk and cache timeline
GloVe · rejected router · 244 MiB peak but 0.623 default recall
Rejected checkpoint-centroid v8 Deep-Image CPU RAM disk and cache timeline
Deep-Image · rejected router · 368 MiB peak but 0.576 default recall
Fashion-MNIST v7 global PQ CPU RAM disk and cache timeline
Fashion-MNIST · v7 global PQ · repetition 2
GloVe v7 global PQ CPU RAM disk and cache timeline
GloVe · v7 global PQ · repetition 2
SIFT v7 global PQ CPU RAM disk and cache timeline
SIFT · v7 global PQ · repetition 2
NYTimes selected 223-probe production repetition CPU RAM disk and cache timeline
NYTimes · selected 223/288 production repetition · 115 MiB peak RSS
GIST code256 selected production repetition CPU RAM disk and cache timeline
GIST · selected 24/96 production repetition · 301 MiB peak RSS
GIST from-scratch adaptive-default rebuild CPU RAM process disk and cache timeline
GIST · no-override adaptive-default rebuild · 328 MiB peak RSS
Deep-Image CPU RAM disk and cache timeline
Deep-Image · median-p95 repetition

Comparable economics

Do not charge only BORSUK for the client

Every system needs application/client compute. BORSUK runs search inside that process; S3 Vectors, turbopuffer, Pinecone, and Chroma receive the request from it and include remote search compute in their service price. The headline product rows therefore exclude common client compute for every system. BORSUK reports its selected S3 index storage and measured GETs, while CPU/RAM/disk stay explicit in the resource evidence above. See the dated cost formulas and current vendor list prices.

Leaf modes

Mode Evaluation

Measured with cargo run --locked --release -p borsuk --example benchmark_report on 10k and 100k synthetic vectors plus the scikit-learn digits CSV, using 100 queries per dataset. Datasets are bulk inserted, compacted into vector-local leaves, then queried. Synthetic datasets use 64 dimensions, max_segments=8, routing_page_overfetch=8, and max_candidates_per_segment=64. Recall charts use tie-aware recall, where a different id at the same exact kth distance is accepted; id recall and termination-reason counts remain visible in the data table. The benchmark CLI also accepts --max-segments, --routing-page-overfetch, and --max-candidates-per-segment for explicit recall/I/O probes. The report fails if pq-scan, vamana-pq, or hybrid falls below 0.95 tie-aware recall@10.

Routing lookahead

Overfetch vs Recall and I/O

The routing-overfetch sweep reruns the high-recall modes with routing_page_overfetch=1,2,4,8,16,32. Higher values allow extra cheap routing metadata when page bounds are tied or close. At each routing layer, overfetch also has a page-level floor, so sibling metadata pages can stay eligible even when the first dense page already contains enough leaf segments for the payload budget. Segment and graph payload budgets stay separate, so the chart shows recall safety without silently raising resident memory.

Filter-first ranking

Metadata sparsity vs work

Here a categorical tier is spread uniformly across every segment, so segment pruning cannot help — the filter simply rejects a growing share of the rows a query would rank. As rejection climbs 0 → 90%, the set of matching rows inside each segment gets small enough that BORSUK switches to filter-first ranking — it ranks the actual matches instead of ranking vector-nearest rows and dropping the non-matches. The result is that id recall@10 stays at 1.00 across the whole sweep while the rows exact-scored per query fall with the match count. Bytes read stay flat because every segment object is still fetched whole — the saving here is scoring work, not I/O.

Why the bump at 80%? It is the crossover, not a bug. Up to ~70% rejection a segment still holds more matches than the candidate budget, so the budgeted scan scores only the budget of vector-nearest rows (a capped subset). At ~80% the matches finally drop below the budget, so filter-first ranking kicks in and scores every match in each segment — briefly more scoring work, but now exact among the matches rather than a capped approximation. Past that point the match count keeps shrinking, so the work falls again. Recall is 1.00 throughout; the bump is BORSUK choosing exactness over a smaller cap once it can afford to. Regenerate with the ignored sparsity_sweep_gate.

More

Reproduce and extend

Each sweep has an ignored gate test that regenerates its CSV — see the section copy for the exact test name. Full-corpus, method, configuration, scale, comparison, and reproduction notes are in the canonical docs/research/ hierarchy. Raw artifacts remain under docs/web/assets/benchmarks/.