17 Aug 2026 · 9 min read · project: gofetch

What a RAG pipeline actually looks like when you measure it

I built GoFetch to fuse dense search and BM25 with reciprocal rank fusion, then re-rank with a cross-encoder. There's a knowledge graph wired in as a third signal too, though it's never actually fired in anything I've deployed. Every choice that did ship came from a 24-question benchmark, not a hunch, including one number that sat quietly wrong for months before I caught it.

Boundary, stated up front: the retrieval numbers below come from a 14-document, 24-question benchmark I built myself. They tell you which architecture choice won on that corpus. They don't tell you how this behaves at 10x the scale. And they only ever cover two signals, dense and BM25: the knowledge graph is real code, switched on by default, but it only activates if a graph file exists on disk, one that's only produced by running ingestion against live credentials I've never actually done in a deployed environment. So it's built, it's wired, and it has never once fired.
2
Signals actually fused
0.917
Hit@1, hybrid + rerank
14
Documents in corpus
55
Tests, all passing

How one query moves through the system

Two retrievers run concurrently (a third, graph retrieval, is wired in alongside them but has never actually had data to run against), get fused by rank rather than by score, then get cut down to five chunks before generation ever starts:

00
Decompose
off by default
→
01
Embed query
local, no API call
→
02
Dense + BM25 (+ Graph, dormant)
parallel, top 20/20
→
03
RRF fuse
k=60, top 10
→
04
Re-rank
cross-encoder, top 5
→
05
Confidence gate
refuse or warn
→
06
Generate
Gemini, streamed

Key decision

BM25 scores are unbounded term-frequency numbers; pgvector cosine similarity is bounded 0 to 1. Rather than normalize two incompatible scales, fusion works on rank position instead: 1/(k + rank) summed across whichever lists a chunk appears in. That's also why wiring in the knowledge graph as a third signal was cheap: fusion doesn't care what produced a ranked list, so the graph slots in without touching the math, whenever it actually has data to rank.

The numbers, and the one that was wrong

The honest version of this story: my headline ablation table quietly drifted stale after I changed the chunk size, and kept reporting results for a corpus I don't even build anymore. Here's what's actually true for what's shipped today:

ConfigurationHit@1Hit@3Hit@5MRRKW Recall
Dense only0.8751.0001.0000.9380.819
BM25 only0.8750.9580.9580.9170.600
Hybrid (RRF)0.8751.0001.0000.9380.728
Hybrid + Rerank0.9171.0001.0000.9580.788

What actually happened

The results file behind that table was byte-identical to a snapshot from an older, larger chunk size. I'd committed it once and never regenerated it after changing the chunk-size default. For months my README said dense retrieval won outright. It didn't. On the real shipped config, hybrid plus re-ranking wins Hit@1 cleanly. The fix wasn't even a re-run: the correct numbers were already sitting in my repo under a different filename, I'd just never swapped them in. I caught it by re-deriving every number straight from the eval output instead of trusting the table that was already there.

Known limitations

Stated plainly rather than left for someone else to find:

Writing every stage by hand here was mostly a way to actually understand what's happening at each step of a RAG pipeline. Next time I'd probably reach for LangChain, LangGraph, or CrewAI to move faster, now that I know exactly what they're abstracting away underneath.

Python · FastAPI · PostgreSQL (pgvector) · BM25 · cross-encoder · Gemini (Vertex AI) · NetworkX · Gradio