Grounded AI
Make the model show its sources.
AI is most useful when it can retrieve trustworthy evidence before it reasons — and when the system around it decides what the model may touch, what it must cite, and what gets validated before anything ships.
01 — The shape of a grounded system
-
Question
a task worth automating, stated precisely
Everything downstream inherits the quality of the question. Vague intent in, plausible nonsense out.
-
Retrieve context
code · docs · data · evidence
Before any model runs, the system gathers what is actually known: the relevant source, the schema, the ticket, the measurement. Retrieval is an engineering problem — indexing, freshness, access control — not a prompt.
-
Rank and select
hybrid retrieval · reranking · budgets
Context windows are budgets. Lexical and semantic retrieval find candidates; reranking decides what deserves the model’s attention, and everything else stays out.
-
Model reasons
over what was found, not what it imagines
The model works from the selected evidence. Its job is synthesis and judgment — not recalling facts the system could have looked up.
-
Evidence-backed output
claims that carry their sources
Useful output points back at what it used. An answer that can cite the retrieved passage, the failing test, or the API response can be checked; one that can’t is a guess with good grammar.
-
Validation
tests · types · review · determinism
Nothing ships on the model’s word. Deterministic checks — compilers, test suites, schema validation, human review — decide what is accepted.
02 — What I work with
Retrieval-augmented generation
Grounding generation in retrieved evidence — and measuring whether retrieval, not the model, is the bottleneck. It usually is.
Hybrid and semantic retrieval
Combining lexical search with embeddings and reranking; knowing when FTS5 at small scale beats a vector database at any scale.
Context engineering
Deciding what enters the window: structure, ordering, compression, and the discipline to leave things out.
AI agents and orchestration
Tool-using agents with narrow jobs, deterministic scaffolding, and multi-agent workflows where the split genuinely helps.
LLM evaluation
Golden sets, failure-mode fixtures, and mocked transports — every AI path in my projects is tested against a transport that never contacts a live model in CI.
Evidence-backed automation
Automation that shows what it observed before what it concluded, and says “unknown” when it cannot know.
03 — Where this shows up in my work
Retrieval over your own material
StudyForge’s “Ask my notes” retrieves the relevant passages from your own documents — and works with no AI configured at all, because retrieval is useful before generation is. At one person’s scale that means FTS5, not a vector database.
Deterministic core, AI at the edges
RepoSignal and Gatehouse compute every score and verdict with pure, tested functions. A model may one day rephrase findings; it will never produce or adjust a number. StudyForge’s scheduler is arithmetic a model cannot reach.
Measuring retrieval itself
ContextBench is the argument as a product: a local-first workbench that inspects what vector search, BM25, and rank fusion actually returned, then scores those rankings with Recall, MRR and nDCG against versioned human judgments — never an LLM grading itself.
Local models, fenced in
ProcessPilot and TraceMark integrate Ollama behind loopback-only, allowlisted transports with explicit user consent — and fail back to deterministic behavior when the model is missing. Every AI failure mode is tested against a mock transport.
What I don’t claim
I’m not a research scientist and I haven’t trained foundation models. My work is the systems engineering around models — retrieval, grounding, orchestration, evaluation, and the deterministic boundaries that make AI features trustworthy enough to ship. On this page, as everywhere else on this site, the claims are limited to what the linked projects demonstrate.