First-party data
Serpaix benchmarks: measured, dated, reproducible
This page publishes Serpaix’s own engineering measurements taken on a named, low-end Windows laptop: desktop search latency and where it goes, the bytes moved across the Rust–WebView2 boundary, candidate-recall quality, and local-model inference speed. Every figure is labelled Measured, Derived, Projected or Pending so you can weigh it, and each one traces back to a file in the repository.
- Published
- 2026-10-06
- Last reviewed
- 2026-10-06
- Scope
- Windows desktop beta
These numbers describe the state of the code at the dates above, not a promise about a future release. Where a result is unflattering it is still published — a benchmark page that reports only wins is marketing, not measurement.
The machine these ran on
Latency without hardware is meaningless, so the target is stated in full. This is a deliberate choice: a four-core laptop with integrated graphics and 7 GB of RAM is a realistic Beta machine, and publishing low-end numbers keeps us honest about what “works on your machine” actually means.
| OS | Windows 11 Home, build 10.0.26200, 64-bit |
|---|---|
| CPU | AMD Ryzen 3 7320U with Radeon Graphics — 4 cores / 8 threads @ 2.4 GHz |
| Memory | 7.31 GiB |
| GPU | Integrated only. Model inference is CPU-only (`size_vram: 0`) |
| WebView2 runtime | 154.0.4258.53 |
| Local runtime | Ollama v0.34.4 at 127.0.0.1:11434 |
| Node / Rust | Release builds throughout; production executable 39,567,872 bytes |
1. Search time scales with candidates
MeasuredMedian end-to-end wall time for one approximate-nearest-neighbour query, by candidate count and corpus geometry, on the shipped production executable with the natural delivery path. 10,000-row corpus, median of two iterations per cell.
| Candidates | 250 | 500 | 1,000 | 1,500 | 2,000 | 2,500 |
|---|---|---|---|---|---|---|
| dense | 74.3 | 161.0 | 319.7 | 454.0 | 529.1 | 675.2 |
| clustered | 74.5 | 136.9 | 257.9 | 466.0 | 662.0 | 683.5 |
| sparse | 65.1 | 111.4 | 244.2 | 349.6 | 481.7 | 593.5 |
| uniform | 58.4 | 172.2 | 245.5 | 357.0 | 486.0 | 623.9 |
Milliseconds. A stress leg at 5,000 candidates measured 1,381.2 ms. Growth is close to linear at roughly 0.25–0.28 ms per candidate and is almost geometry-independent — the signature of a transport-bound path, since scoring and parsing together account for about 13 ms of a 675 ms query.
2. Where those milliseconds actually go
MeasuredComponent timers from the same production run: 6,000-row cleanroom corpus, eight queries, k = 80. End-to-end median was 593.3 ms (p95 603 ms), reproduced to ±1 ms across runs.
| Stage | Median | Share | What it is |
|---|---|---|---|
| probe SQL | 6.2 ms | 1 % | Rust SQL, 13–79 buckets |
| pad SQL | 21.6 ms | 4 % | Rust SQL, table-wide recency LIMIT 2500 |
| vectorFetch | 471.2 ms | 80 % | Rust SQL + serialise + IPC transfer + webview parse |
| Stage-B decode | 88.9 ms | 15 % | decodeVector over 2,557 rows in the webview |
| cosine + sort | 0.15 ms | ≈ 0 % | Exact scoring — not a target |
A benchmark-harness stringify of 12.7 ms is excluded above because it is not production cost; the true production figure is therefore nearer 580 ms.
Inside the 471 ms vector fetch
| Rust fetch 2,500 rows (TEXT, 400/chunk) | 34.8 ms | Measured |
|---|---|---|
| Rust serde_json response serialise (13.63 MB) | 24.0 ms | Measured |
| Webview JSON.parse (replay, 15.4 MB over 7 chunks) | 5.9 ms | Measured |
| IPC marshalling + transfer of ~14.3 MB onto the WebView2 boundary | ≈ 406 ms | Derived |
The dominant cost is moving one ~14.3 MB payload of double-encoded text vectors per query, at roughly 35 MB/s across the process boundary. What is not the problem: the IPC round-trip floor (2.2 ms × 10 calls = 22 ms, 3.7 %), the JSON parse (5.9 ms) and cosine scoring (0.15 ms) together stay under 5 % of end-to-end.
3. The production build was silently on a fallback IPC path
MeasuredThe Tauri configuration’s connect-src directive did not allow ipc: http://ipc.localhost, so the framework’s fast custom-protocol path was rejected by the content security policy on every call. The runtime then fell back to window.ipc.postMessage plus eval(JSON). Nothing errored visibly; search was simply slower than it needed to be.
It was proven with a timing-independent probe: fetch('http://ipc.localhost/__vt_probe__') was rejected with Failed to fetch, two CSP violations were recorded against ipc.localhost, and the detector reported the delivery path as postMessage-fallback.
This is the class of bug that instrumented benchmarks exist to catch, and it is the reason the numbers in section 1 are what they are. Per-query cost is 14,339,384 bytes received against 36,294 bytes sent, across 10 IPC calls, for a single search.
4. Changing only the encoding shrinks the wire 3.57×
MeasuredA Rust-side component split over the same database and the same rows, per 400-row IPC chunk, comparing today’s stored text (serialised as a JSON string inside the IPC body) against a versioned binary f32 wire:
| Quantity, per 400-row chunk | Text (shipped) | Binary wire | Ratio |
|---|---|---|---|
| Bytes crossing IPC | 2,217,439 | 620,324 | 3.57× |
| Bytes per row | 5,543 | 1,551 | 3.6× |
With the honest counterweight: storage stays as text, so the binary command must decode it to emit floats — a cost of about 54.9 ms per chunk that the text path never pays. Whether the smaller payload wins overall is decided by the end-to-end comparison, which is in the pending section below rather than asserted here.
5. Scored in-process, the same work takes 3.4× less time
MeasuredAn isolated Rust plane benchmark doing the identical work — fetch, decode, serialise, score and rank 2,500 rows of 384 dimensions — with vectors stored as text versus binary f32. Release build, bundled SQLite.
| Metric | Text | Binary f32 | Ratio |
|---|---|---|---|
| insert 2,500 rows | 290.9 ms | 40.0 ms | 7.3× |
| stored bytes | 13.62 MB | 3.84 MB | 3.55× |
| fetch (400/chunk) | 34.8 ms | 11.9 ms | 2.9× |
| decode 2,500 rows | 27.1 ms | 1.8 ms | 14.9× |
| serialise response | 24.0 ms | 11.7 ms | 2.1× |
| score + rank (k = 80) | 0.15 ms | 0.14 ms | — |
| total plane | 86.1 ms | 25.6 ms | 3.4× |
The speed is only interesting because the ranking is unchanged. In the same run, decoded values were identical, the top-80 ordering was identical, and the maximum score delta was 0.000e0. Separately, a Rust f64 scorer was checked against the shipped TypeScript scorer: ranking order and result set matched in 20/20 cases across all eight test blocks with a maximum delta of zero. A f32 binary scorer kept top-1 identical in 160/160 cases; its only deviations were swaps inside a tie band roughly two orders of magnitude tighter than the 1×10⁻⁵ tolerance it is held to.
6. The other problem was recall, and it was not the index
MeasuredWhile measuring speed we also measured quality, on a 50,000-row corpus across four geometries. The truthful summary is that the index is perfect and the candidate set was mostly noise:
| Measurement | Result | |
|---|---|---|
| Index health | 50,000 / 50,000 rows indexed, 0 orphans, 25/25 spot checks, fidelity identical | Measured |
| Radius-1 bucket ball | ~157 rows (13 buckets of 4,096) — below the k×3 = 240 pad threshold | Measured |
| Recency pad contribution | ~2,491 table-wide “most recently updated” rows per query | Measured |
| Candidate set composition | ~94 % recency noise | Measured |
| Pad lift over a random draw | 1.23× (range 0.75–1.27× across geometries) | Measured |
| Candidate recall (mem-row) | 5.5–16.0 % on non-dense geometries | Measured |
| True neighbours lost to minScore = 0.18 | true similarity 0.122–0.168 | Measured |
| Post-minScore recall | 0.5–1.3 % on non-dense geometries | Measured |
| Dense geometry (for contrast) | 54.0 % candidate recall, 46.8 % final | Measured |
The mechanism: a radius-1 probe returns roughly 157 rows, which is below the 240-row threshold that triggers a recency-based top-up. That top-up then fills almost the entire candidate set with the most recently updated rows in the database — rows unrelated to the query — and the similarity floor afterwards discards the genuine neighbours that did arrive, because their true similarity sits below it. On dense corpora, where the probe ball is large enough, the same pipeline reaches 54.0 % candidate recall. The conclusion recorded in the report is a structural one: inside the current parameter space, no tuning fixes non-dense recall.
7. Local AI on a four-core laptop with no GPU
MeasuredThese numbers come from an end-to-end run of the offline product against a realistic 65-file (0.53 MiB) business corpus with 18 labelled questions, with inference on the CPU only.
| Embedding call (nomic-embed-text) | mean 132 ms · p50 105 ms · p95 341 ms · max 755 ms (50 calls) | Measured |
|---|---|---|
| Warm tiny generation | 639 ms wall | Measured |
| Cold load + tiny generation | 2,696 ms wall | Measured |
| Grounded answer, qwen2.5:0.5b | 2.2–50.0 s · mean 30.6 s · median 28.8 s | Measured |
| Grounded answer, llama3.2 3B (1,680-token RAG prompt) | prompt 63,753 ms (≈ 26 tok/s) + output 13,755 ms (≈ 5.8 tok/s) = 130,199 ms wall | Measured |
| App-side timeout | 120,000 ms — so 3B cannot finish a grounded answer on this machine | Measured |
| Model download (2.02 GB) | ≈ 4 min 38 s ≈ 7.4 MB/s | Measured |
The practical reading: a small local model answers in half a minute on this class of machine, and a 3-billion-parameter model does not answer at all inside the client’s own timeout. That is why the product lets you choose the model rather than picking one for you, and why “runs on your machine” is a claim we only make after measuring the machine in question.
What we got wrong
These were our priors before instrumenting, and what the measurements did to them. They are recorded here because the value of first-party data is mostly in the corrections.
We assumed the ANN candidate cap was a bottleneck.
The cap binds, but the component it belongs to — the probe SQL — costs 6.2 ms, about 1 % of the search path. The 80 % is the vector payload.
We assumed Tauri IPC round-trips were slow.
The round-trip floor is 2.2 ms per call across 10 calls — 3.7 % of end-to-end. The cost is the *bytes*, roughly 406 ms of marshalling and transfer, not the transport.
We assumed a plain `cargo build --release` produced the production binary.
It clobbered the executable with an 11.7 MB development build instead of the 39.6 MB production build, and the benchmark died against a dev URL. Production builds require the project's build shim.
We assumed search quality was healthy because the index was healthy.
The index is perfect — 50,000/50,000 rows, zero orphans — and candidate recall is still 5.5–16.0 % on non-dense corpora. A healthy index says nothing about candidate generation.
We shipped a config omission that silently disabled the fast IPC path.
`connect-src` never allowed `ipc: http://ipc.localhost`, so every fast-path fetch was CSP-rejected and the app quietly used `postMessage` + `eval(JSON)` for every call. Nothing failed loudly; performance was just worse.
What is not measured yet
PendingListed so the gaps are visible rather than implied away. Work in this section has a protocol and a harness but no recorded result, and any performance figure derived from it is a projection from measured components, not a measurement.
| Open measurement | The question it settles |
|---|---|
| CSP-fixed TEXT leg (A2) | Does the custom-protocol path beat postMessage + eval once connect-src is corrected? |
| Binary f32 wire, end to end (B2) | Payload, marshalling and latency cells for the new transport. |
| Binary wire over the forced fallback (B1) | Does the binary wire survive the fallback delivery path, or does it depend on the CSP fix? |
| Main-thread long tasks, A vs B | UI-responsiveness claim is undecided until the same-exe comparison runs. |
| Peak RSS, A vs B | Whether the wire decoder recreates the payload problem as heap churn. |
| End-to-end ranking fidelity, new exe | Top-k IDs identical and scores within 1e-6 on the real annTopK path. |
For the sake of completeness, the projections that follow from the components already measured are a search path around 190–210 ms if the vector payload were the smaller binary encoding, and around 30–50 ms if scoring moved fully into Rust and the vectors stopped crossing the boundary. Both are projected from measured parts and neither should be quoted as a shipped result.
Method
- One variable at a time. Within a comparison run, the queries, the candidate rows, the JavaScript cosine function and the database are identical across legs. Only the transport changes, which is what makes the result causal rather than correlational.
- Cleanroom isolation. Each run executes against a throwaway profile with fresh application-data directories, and the harness asserts the database path it opened is inside that sandbox. Your own database is never touched by a benchmark.
- Production builds only. Latency is measured against the release executable produced by the project build shim, not a development build — after a development binary was mistakenly measured once and produced a meaningless result.
- Synthetic corpora, labelled. Recall is measured on generated geometries (dense, clustered, sparse, uniform, near-duplicate) at 6,000, 10,000 and 50,000 rows. No user data is used, and recall percentages are therefore properties of the algorithm on those geometries, not a claim about your library.
- Correctness gates before speed. Ranking equivalence is a hard gate: a top-k change halts the benchmark before it scales, and the run fails closed. Unit-level state at the time of writing was 555 Rust tests passed with 0 failures (1 intentionally ignored benchmark), 68/68 local-store tests and 337/337 in the shared brain package.
- Verdict labels. Measured is read from an instrumented run. Derived is an arithmetic residual between measured parts. Projected is an extrapolation with no gated implementation behind it. Pending means the protocol exists and no result does.
Evidence artifacts
Every figure above maps to one of these files in the Serpaix monorepo. The repository opens publicly under the MIT License on January 1, 2027, at which point these become independently reproducible rather than merely stated.
| Production-path IPC bench | apps/desktop/e2e/prod-path-ipc-bench.result.json |
|---|---|
| IPC path proof | apps/desktop/e2e/ipc-path-probe.evidence.json |
| Vector transport A/B result | apps/desktop/e2e/vector-transport-bench.result.json |
| Rust plane bench (TEXT vs BLOB) | apps/desktop/src-tauri/tests/vector_plane_bench.rs |
| Ranking equivalence fixture + runner | apps/desktop/src-tauri/scripts/ranking_equivalence.rs |
| ANN recall diagnostic + gate logs | packages/local-store/scripts/ann-recall-diagnostic.mjs |
| Full recall report | packages/local-store/scripts/out/ann-recall_full.report.md |
| Narrative method + decisions | docs/research/SERPAIX_DESKTOP_ANN_RECALL_AND_IPC_AUDIT.md |
| Binary transport report | docs/research/SERPAIX_DESKTOP_BINARY_VECTOR_TRANSPORT_REPORT.md |
| Real-machine Beta validation | docs/research/REAL_MACHINE_BETA_VALIDATION_REPORT.md |
Corrections and contact
If a number here looks wrong, contradictory, or measured in a way that cannot support the claim, say so — that is a bug report, and it is the most useful kind. Write to hello@serpaix.xyz or open an issue once the source is public. Corrections are recorded by updating the review date above rather than by silently editing the text.
