First-party data

Serpaix benchmarks: measured, dated, reproducible

This page publishes Serpaix’s own engineering measurements taken on a named, low-end Windows laptop: desktop search latency and where it goes, the bytes moved across the Rust–WebView2 boundary, candidate-recall quality, and local-model inference speed. Every figure is labelled Measured, Derived, Projected or Pending so you can weigh it, and each one traces back to a file in the repository.

Published
2026-10-06
Last reviewed
2026-10-06
Scope
Windows desktop beta

These numbers describe the state of the code at the dates above, not a promise about a future release. Where a result is unflattering it is still published — a benchmark page that reports only wins is marketing, not measurement.

The machine these ran on

Latency without hardware is meaningless, so the target is stated in full. This is a deliberate choice: a four-core laptop with integrated graphics and 7 GB of RAM is a realistic Beta machine, and publishing low-end numbers keeps us honest about what “works on your machine” actually means.

OSWindows 11 Home, build 10.0.26200, 64-bit
CPUAMD Ryzen 3 7320U with Radeon Graphics — 4 cores / 8 threads @ 2.4 GHz
Memory7.31 GiB
GPUIntegrated only. Model inference is CPU-only (`size_vram: 0`)
WebView2 runtime154.0.4258.53
Local runtimeOllama v0.34.4 at 127.0.0.1:11434
Node / RustRelease builds throughout; production executable 39,567,872 bytes

1. Search time scales with candidates

Measured

Median end-to-end wall time for one approximate-nearest-neighbour query, by candidate count and corpus geometry, on the shipped production executable with the natural delivery path. 10,000-row corpus, median of two iterations per cell.

Candidates2505001,0001,5002,0002,500
dense74.3161.0319.7454.0529.1675.2
clustered74.5136.9257.9466.0662.0683.5
sparse65.1111.4244.2349.6481.7593.5
uniform58.4172.2245.5357.0486.0623.9

Milliseconds. A stress leg at 5,000 candidates measured 1,381.2 ms. Growth is close to linear at roughly 0.25–0.28 ms per candidate and is almost geometry-independent — the signature of a transport-bound path, since scoring and parsing together account for about 13 ms of a 675 ms query.

2. Where those milliseconds actually go

Measured

Component timers from the same production run: 6,000-row cleanroom corpus, eight queries, k = 80. End-to-end median was 593.3 ms (p95 603 ms), reproduced to ±1 ms across runs.

StageMedianShareWhat it is
probe SQL6.2 ms1 %Rust SQL, 13–79 buckets
pad SQL21.6 ms4 %Rust SQL, table-wide recency LIMIT 2500
vectorFetch471.2 ms80 %Rust SQL + serialise + IPC transfer + webview parse
Stage-B decode88.9 ms15 %decodeVector over 2,557 rows in the webview
cosine + sort0.15 ms≈ 0 %Exact scoring — not a target

A benchmark-harness stringify of 12.7 ms is excluded above because it is not production cost; the true production figure is therefore nearer 580 ms.

Inside the 471 ms vector fetch

Rust fetch 2,500 rows (TEXT, 400/chunk)34.8 msMeasured
Rust serde_json response serialise (13.63 MB)24.0 msMeasured
Webview JSON.parse (replay, 15.4 MB over 7 chunks)5.9 msMeasured
IPC marshalling + transfer of ~14.3 MB onto the WebView2 boundary≈ 406 msDerived

The dominant cost is moving one ~14.3 MB payload of double-encoded text vectors per query, at roughly 35 MB/s across the process boundary. What is not the problem: the IPC round-trip floor (2.2 ms × 10 calls = 22 ms, 3.7 %), the JSON parse (5.9 ms) and cosine scoring (0.15 ms) together stay under 5 % of end-to-end.

3. The production build was silently on a fallback IPC path

Measured

The Tauri configuration’s connect-src directive did not allow ipc: http://ipc.localhost, so the framework’s fast custom-protocol path was rejected by the content security policy on every call. The runtime then fell back to window.ipc.postMessage plus eval(JSON). Nothing errored visibly; search was simply slower than it needed to be.

It was proven with a timing-independent probe: fetch('http://ipc.localhost/__vt_probe__') was rejected with Failed to fetch, two CSP violations were recorded against ipc.localhost, and the detector reported the delivery path as postMessage-fallback.

This is the class of bug that instrumented benchmarks exist to catch, and it is the reason the numbers in section 1 are what they are. Per-query cost is 14,339,384 bytes received against 36,294 bytes sent, across 10 IPC calls, for a single search.

4. Changing only the encoding shrinks the wire 3.57×

Measured

A Rust-side component split over the same database and the same rows, per 400-row IPC chunk, comparing today’s stored text (serialised as a JSON string inside the IPC body) against a versioned binary f32 wire:

Quantity, per 400-row chunkText (shipped)Binary wireRatio
Bytes crossing IPC2,217,439620,3243.57×
Bytes per row5,5431,5513.6×

With the honest counterweight: storage stays as text, so the binary command must decode it to emit floats — a cost of about 54.9 ms per chunk that the text path never pays. Whether the smaller payload wins overall is decided by the end-to-end comparison, which is in the pending section below rather than asserted here.

5. Scored in-process, the same work takes 3.4× less time

Measured

An isolated Rust plane benchmark doing the identical work — fetch, decode, serialise, score and rank 2,500 rows of 384 dimensions — with vectors stored as text versus binary f32. Release build, bundled SQLite.

MetricTextBinary f32Ratio
insert 2,500 rows290.9 ms40.0 ms7.3×
stored bytes13.62 MB3.84 MB3.55×
fetch (400/chunk)34.8 ms11.9 ms2.9×
decode 2,500 rows27.1 ms1.8 ms14.9×
serialise response24.0 ms11.7 ms2.1×
score + rank (k = 80)0.15 ms0.14 ms—
total plane86.1 ms25.6 ms3.4×

The speed is only interesting because the ranking is unchanged. In the same run, decoded values were identical, the top-80 ordering was identical, and the maximum score delta was 0.000e0. Separately, a Rust f64 scorer was checked against the shipped TypeScript scorer: ranking order and result set matched in 20/20 cases across all eight test blocks with a maximum delta of zero. A f32 binary scorer kept top-1 identical in 160/160 cases; its only deviations were swaps inside a tie band roughly two orders of magnitude tighter than the 1×10⁻⁵ tolerance it is held to.

6. The other problem was recall, and it was not the index

Measured

While measuring speed we also measured quality, on a 50,000-row corpus across four geometries. The truthful summary is that the index is perfect and the candidate set was mostly noise:

MeasurementResult
Index health50,000 / 50,000 rows indexed, 0 orphans, 25/25 spot checks, fidelity identicalMeasured
Radius-1 bucket ball~157 rows (13 buckets of 4,096) — below the k×3 = 240 pad thresholdMeasured
Recency pad contribution~2,491 table-wide “most recently updated” rows per queryMeasured
Candidate set composition~94 % recency noiseMeasured
Pad lift over a random draw1.23× (range 0.75–1.27× across geometries)Measured
Candidate recall (mem-row)5.5–16.0 % on non-dense geometriesMeasured
True neighbours lost to minScore = 0.18true similarity 0.122–0.168Measured
Post-minScore recall0.5–1.3 % on non-dense geometriesMeasured
Dense geometry (for contrast)54.0 % candidate recall, 46.8 % finalMeasured

The mechanism: a radius-1 probe returns roughly 157 rows, which is below the 240-row threshold that triggers a recency-based top-up. That top-up then fills almost the entire candidate set with the most recently updated rows in the database — rows unrelated to the query — and the similarity floor afterwards discards the genuine neighbours that did arrive, because their true similarity sits below it. On dense corpora, where the probe ball is large enough, the same pipeline reaches 54.0 % candidate recall. The conclusion recorded in the report is a structural one: inside the current parameter space, no tuning fixes non-dense recall.

7. Local AI on a four-core laptop with no GPU

Measured

These numbers come from an end-to-end run of the offline product against a realistic 65-file (0.53 MiB) business corpus with 18 labelled questions, with inference on the CPU only.

Embedding call (nomic-embed-text)mean 132 ms · p50 105 ms · p95 341 ms · max 755 ms (50 calls)Measured
Warm tiny generation639 ms wallMeasured
Cold load + tiny generation2,696 ms wallMeasured
Grounded answer, qwen2.5:0.5b2.2–50.0 s · mean 30.6 s · median 28.8 sMeasured
Grounded answer, llama3.2 3B (1,680-token RAG prompt)prompt 63,753 ms (≈ 26 tok/s) + output 13,755 ms (≈ 5.8 tok/s) = 130,199 ms wallMeasured
App-side timeout120,000 ms — so 3B cannot finish a grounded answer on this machineMeasured
Model download (2.02 GB)≈ 4 min 38 s ≈ 7.4 MB/sMeasured

The practical reading: a small local model answers in half a minute on this class of machine, and a 3-billion-parameter model does not answer at all inside the client’s own timeout. That is why the product lets you choose the model rather than picking one for you, and why “runs on your machine” is a claim we only make after measuring the machine in question.

What we got wrong

These were our priors before instrumenting, and what the measurements did to them. They are recorded here because the value of first-party data is mostly in the corrections.

We assumed the ANN candidate cap was a bottleneck.

The cap binds, but the component it belongs to — the probe SQL — costs 6.2 ms, about 1 % of the search path. The 80 % is the vector payload.

We assumed Tauri IPC round-trips were slow.

The round-trip floor is 2.2 ms per call across 10 calls — 3.7 % of end-to-end. The cost is the *bytes*, roughly 406 ms of marshalling and transfer, not the transport.

We assumed a plain `cargo build --release` produced the production binary.

It clobbered the executable with an 11.7 MB development build instead of the 39.6 MB production build, and the benchmark died against a dev URL. Production builds require the project's build shim.

We assumed search quality was healthy because the index was healthy.

The index is perfect — 50,000/50,000 rows, zero orphans — and candidate recall is still 5.5–16.0 % on non-dense corpora. A healthy index says nothing about candidate generation.

We shipped a config omission that silently disabled the fast IPC path.

`connect-src` never allowed `ipc: http://ipc.localhost`, so every fast-path fetch was CSP-rejected and the app quietly used `postMessage` + `eval(JSON)` for every call. Nothing failed loudly; performance was just worse.

What is not measured yet

Pending

Listed so the gaps are visible rather than implied away. Work in this section has a protocol and a harness but no recorded result, and any performance figure derived from it is a projection from measured components, not a measurement.

Open measurementThe question it settles
CSP-fixed TEXT leg (A2)Does the custom-protocol path beat postMessage + eval once connect-src is corrected?
Binary f32 wire, end to end (B2)Payload, marshalling and latency cells for the new transport.
Binary wire over the forced fallback (B1)Does the binary wire survive the fallback delivery path, or does it depend on the CSP fix?
Main-thread long tasks, A vs BUI-responsiveness claim is undecided until the same-exe comparison runs.
Peak RSS, A vs BWhether the wire decoder recreates the payload problem as heap churn.
End-to-end ranking fidelity, new exeTop-k IDs identical and scores within 1e-6 on the real annTopK path.

For the sake of completeness, the projections that follow from the components already measured are a search path around 190–210 ms if the vector payload were the smaller binary encoding, and around 30–50 ms if scoring moved fully into Rust and the vectors stopped crossing the boundary. Both are projected from measured parts and neither should be quoted as a shipped result.

Method

  1. One variable at a time. Within a comparison run, the queries, the candidate rows, the JavaScript cosine function and the database are identical across legs. Only the transport changes, which is what makes the result causal rather than correlational.
  2. Cleanroom isolation. Each run executes against a throwaway profile with fresh application-data directories, and the harness asserts the database path it opened is inside that sandbox. Your own database is never touched by a benchmark.
  3. Production builds only. Latency is measured against the release executable produced by the project build shim, not a development build — after a development binary was mistakenly measured once and produced a meaningless result.
  4. Synthetic corpora, labelled. Recall is measured on generated geometries (dense, clustered, sparse, uniform, near-duplicate) at 6,000, 10,000 and 50,000 rows. No user data is used, and recall percentages are therefore properties of the algorithm on those geometries, not a claim about your library.
  5. Correctness gates before speed. Ranking equivalence is a hard gate: a top-k change halts the benchmark before it scales, and the run fails closed. Unit-level state at the time of writing was 555 Rust tests passed with 0 failures (1 intentionally ignored benchmark), 68/68 local-store tests and 337/337 in the shared brain package.
  6. Verdict labels. Measured is read from an instrumented run. Derived is an arithmetic residual between measured parts. Projected is an extrapolation with no gated implementation behind it. Pending means the protocol exists and no result does.

Evidence artifacts

Every figure above maps to one of these files in the Serpaix monorepo. The repository opens publicly under the MIT License on January 1, 2027, at which point these become independently reproducible rather than merely stated.

Production-path IPC benchapps/desktop/e2e/prod-path-ipc-bench.result.json
IPC path proofapps/desktop/e2e/ipc-path-probe.evidence.json
Vector transport A/B resultapps/desktop/e2e/vector-transport-bench.result.json
Rust plane bench (TEXT vs BLOB)apps/desktop/src-tauri/tests/vector_plane_bench.rs
Ranking equivalence fixture + runnerapps/desktop/src-tauri/scripts/ranking_equivalence.rs
ANN recall diagnostic + gate logspackages/local-store/scripts/ann-recall-diagnostic.mjs
Full recall reportpackages/local-store/scripts/out/ann-recall_full.report.md
Narrative method + decisionsdocs/research/SERPAIX_DESKTOP_ANN_RECALL_AND_IPC_AUDIT.md
Binary transport reportdocs/research/SERPAIX_DESKTOP_BINARY_VECTOR_TRANSPORT_REPORT.md
Real-machine Beta validationdocs/research/REAL_MACHINE_BETA_VALIDATION_REPORT.md

Corrections and contact

If a number here looks wrong, contradictory, or measured in a way that cannot support the claim, say so — that is a bug report, and it is the most useful kind. Write to hello@serpaix.xyz or open an issue once the source is public. Corrections are recorded by updating the review date above rather than by silently editing the text.