AI · Developer Tools
ContextBench
Stop guessing whether retrieval works — measure it.
- Recall · MRR · nDCG
- no LLM judges
- Apache-2.0
01The problem
RAG failures are retrieval failures you can’t see.
When a grounded system answers badly, the model takes the blame — but most of the time the evidence it needed was never retrieved, or was retrieved and buried. Teams tune prompts and swap models while the actual failure stays invisible: nobody can see what the retriever returned, or measure whether a change made it better.
ContextBench is the workbench for exactly that gap: debug and measure your RAG pipeline instead of guessing whether retrieval works. I’d already made the argument in writing — retrieval quality before model quality — and this is the argument as software.
02The measurement stance
Judged by humans, scored by arithmetic.
Every metric in ContextBench — Recall@K, Precision@K, MRR, Hit Rate@K, nDCG@K — is computed from stored, versioned relevance judgments that a person made by marking chunks in the Query Debugger. It does not ask an LLM to grade retrieval, and it refuses to invent a combined “RAG quality” score.
Score scales stay honest too: BM25 and vector similarity remain on their native scales instead of being normalized into a fake common one, and hybrid ranking uses deterministic reciprocal rank fusion:
RRF(d) = sum(1 / (60 + rank_i(d)))
Same rules as the rest of my work: deterministic scoring, honest unknowns — the built-in benchmark marks cross-encoder timing
NOT_RUN unless that model was actually tested.
03What the workbench does
What the workbench does
Ingest with provenance
TXT, Markdown, PDF and source files — normalized, never executed. Every chunk keeps its document, heading, page, and character range.
Compare chunking
Fixed-token, paragraph-aware, and heading-aware strategies side by side — where retrieval quality is usually decided.
Frozen indexes
Several index configurations coexist per project, stored in persisted local Qdrant collections with document, type, and tag filters.
Inspect every ranker
Native vector scores, native BM25 scores, deterministic RRF fusion, and an optional local cross-encoder over the full candidate set.
Judge, then measure
Mark relevant chunks in the Query Debugger to build versioned evaluation queries — the only ground truth the metrics ever use.
Keep the receipts
Experiments persist configuration, rankings, metrics, and measured latency, exportable as versioned JSON for CI and agent workflows.
04Engineering decisions
Engineering decisions
- A zero-download provider is the default
- The built-in hash embedding provider is deterministic and needs no model, so the demo, the tests, and CI pipelines run with nothing downloaded. Semantic quality needs a real embedding model — and the docs say so instead of pretending hashes understand meaning.
- Model downloads are consent-gated and pinned
-
Before any request may fetch the recommended
bge-smallembedder or the MiniLM cross-encoder, the UI shows the model name, approximate size, and cache location. Successful index builds record the resolved model revision, so later query embeddings provably use the same weights. - One Python core, two front doors
-
The CLI and the SolidJS web app call the same retrieval and evaluation code, and
eval --json, exports, and the benchmark emit named, versioned schemas — so the workbench slots into CI gates and agent workflows without screen-scraping. - Local-first, defensively
- Documents, vectors, and experiments live in local SQLite and Qdrant files; both servers bind to loopback and the API rejects non-loopback Host headers. Uploaded files are data, never code. Optional Ollama answers label retrieved text as untrusted evidence and reject uncited or spoofed citations.
05Ground truth, versioned
Ground truth, versioned
Judgments are data with a lifecycle: evaluation queries are versioned, each carries its marked-relevant chunks, and experiments pin the dataset version and the frozen index they ran against. Change the chunking, re-run the experiment, and the comparison is apples-to-apples — or the tool tells you it isn’t.
Verification follows the house pattern: hand-computed metric tests, determinism tests across chunking and fusion, fixture-to-evaluation API tests, component tests, and a Playwright flow through the whole product — plus a wheel check proving the built package actually contains the browser app.
06Stack and links
Stack and links
Python 3.12 · SolidJS + TypeScript · Qdrant (local, persisted) · SQLite · BM25 · optional Sentence Transformers and cross-encoder reranking · optional Ollama answers over cited evidence · uv · Playwright · GitHub Actions. Apache-2.0.
Repository · Architecture · Evaluation metrics · Hybrid retrieval · v1.0.0 release