Skip to content

AI · Developer Tools

ContextBench

Stop guessing whether retrieval works — measure it.

  • Recall · MRR · nDCG
  • no LLM judges
  • Apache-2.0
The ContextBench Query Debugger: a question about Raft compaction, a measured five-stage pipeline (vector, BM25, fuse, rerank, assemble), total latency, and vector versus BM25 result columns with per-chunk scores and relevance checkboxes
The Query Debugger — every stage measured, every score on its native scale

01The problem

RAG failures are retrieval failures you can’t see.

When a grounded system answers badly, the model takes the blame — but most of the time the evidence it needed was never retrieved, or was retrieved and buried. Teams tune prompts and swap models while the actual failure stays invisible: nobody can see what the retriever returned, or measure whether a change made it better.

ContextBench is the workbench for exactly that gap: debug and measure your RAG pipeline instead of guessing whether retrieval works. I’d already made the argument in writing — retrieval quality before model quality — and this is the argument as software.

02The measurement stance

Judged by humans, scored by arithmetic.

Every metric in ContextBench — Recall@K, Precision@K, MRR, Hit Rate@K, nDCG@K — is computed from stored, versioned relevance judgments that a person made by marking chunks in the Query Debugger. It does not ask an LLM to grade retrieval, and it refuses to invent a combined “RAG quality” score.

Score scales stay honest too: BM25 and vector similarity remain on their native scales instead of being normalized into a fake common one, and hybrid ranking uses deterministic reciprocal rank fusion:

RRF(d) = sum(1 / (60 + rank_i(d)))

Same rules as the rest of my work: deterministic scoring, honest unknowns — the built-in benchmark marks cross-encoder timing NOT_RUN unless that model was actually tested.

03What the workbench does

What the workbench does

Ingest with provenance

TXT, Markdown, PDF and source files — normalized, never executed. Every chunk keeps its document, heading, page, and character range.

Compare chunking

Fixed-token, paragraph-aware, and heading-aware strategies side by side — where retrieval quality is usually decided.

Frozen indexes

Several index configurations coexist per project, stored in persisted local Qdrant collections with document, type, and tag filters.

Inspect every ranker

Native vector scores, native BM25 scores, deterministic RRF fusion, and an optional local cross-encoder over the full candidate set.

Judge, then measure

Mark relevant chunks in the Query Debugger to build versioned evaluation queries — the only ground truth the metrics ever use.

Keep the receipts

Experiments persist configuration, rankings, metrics, and measured latency, exportable as versioned JSON for CI and agent workflows.

A ContextBench experiment over the demo corpus: grids of recall, precision, nDCG and hit-rate at K equals 1, 3, 5 and 10, each labeled 'computed from stored rankings'
An experiment on the demo corpus — standard IR metrics, computed from stored rankings

04Engineering decisions

Engineering decisions

A zero-download provider is the default
The built-in hash embedding provider is deterministic and needs no model, so the demo, the tests, and CI pipelines run with nothing downloaded. Semantic quality needs a real embedding model — and the docs say so instead of pretending hashes understand meaning.
Model downloads are consent-gated and pinned
Before any request may fetch the recommended bge-small embedder or the MiniLM cross-encoder, the UI shows the model name, approximate size, and cache location. Successful index builds record the resolved model revision, so later query embeddings provably use the same weights.
One Python core, two front doors
The CLI and the SolidJS web app call the same retrieval and evaluation code, and eval --json, exports, and the benchmark emit named, versioned schemas — so the workbench slots into CI gates and agent workflows without screen-scraping.
Local-first, defensively
Documents, vectors, and experiments live in local SQLite and Qdrant files; both servers bind to loopback and the API rejects non-loopback Host headers. Uploaded files are data, never code. Optional Ollama answers label retrieved text as untrusted evidence and reject uncited or spoofed citations.

05Ground truth, versioned

Ground truth, versioned

Judgments are data with a lifecycle: evaluation queries are versioned, each carries its marked-relevant chunks, and experiments pin the dataset version and the frozen index they ran against. Change the chunking, re-run the experiment, and the comparison is apples-to-apples — or the tool tells you it isn’t.

Verification follows the house pattern: hand-computed metric tests, determinism tests across chunking and fusion, fixture-to-evaluation API tests, component tests, and a Playwright flow through the whole product — plus a wheel check proving the built package actually contains the browser app.

ContextBench's evaluation set: versioned, human-marked relevance judgments for three demo queries, with the note that metrics use these judgments only
Versioned judgments — the only ground truth the metrics ever see

06Stack and links

Stack and links

Python 3.12 · SolidJS + TypeScript · Qdrant (local, persisted) · SQLite · BM25 · optional Sentence Transformers and cross-encoder reranking · optional Ollama answers over cited evidence · uv · Playwright · GitHub Actions. Apache-2.0.

Repository · Architecture · Evaluation metrics · Hybrid retrieval · v1.0.0 release