17 Aug 2026 · 9 min read · project: gofetch
What a RAG pipeline actually looks like when you measure it
I built GoFetch to fuse dense search and BM25 with reciprocal rank fusion, then re-rank with a cross-encoder. There's a knowledge graph wired in as a third signal too, though it's never actually fired in anything I've deployed. Every choice that did ship came from a 24-question benchmark, not a hunch, including one number that sat quietly wrong for months before I caught it.
How one query moves through the system
Two retrievers run concurrently (a third, graph retrieval, is wired in alongside them but has never actually had data to run against), get fused by rank rather than by score, then get cut down to five chunks before generation ever starts:
Key decision
1/(k + rank) summed across whichever lists a chunk appears in. That's also why wiring in the knowledge graph as a third signal was cheap: fusion doesn't care what produced a ranked list, so the graph slots in without touching the math, whenever it actually has data to rank.The numbers, and the one that was wrong
The honest version of this story: my headline ablation table quietly drifted stale after I changed the chunk size, and kept reporting results for a corpus I don't even build anymore. Here's what's actually true for what's shipped today:
| Configuration | Hit@1 | Hit@3 | Hit@5 | MRR | KW Recall |
|---|---|---|---|---|---|
| Dense only | 0.875 | 1.000 | 1.000 | 0.938 | 0.819 |
| BM25 only | 0.875 | 0.958 | 0.958 | 0.917 | 0.600 |
| Hybrid (RRF) | 0.875 | 1.000 | 1.000 | 0.938 | 0.728 |
| Hybrid + Rerank | 0.917 | 1.000 | 1.000 | 0.958 | 0.788 |
What actually happened
Known limitations
Stated plainly rather than left for someone else to find:
- A real concurrency bug I haven't fixed yet. My dense retriever is a shared singleton, and setting the query embedding and reading it back are two separate calls with an await boundary between them. Two overlapping requests can interleave so one silently gets results meant for the other.
- The knowledge graph has never actually run. It's real code, wired into fusion and on by default, but it only activates once a graph file exists on disk, and building that file means running ingestion against live credentials I've never done outside local testing. So it ships switched on and has never once fired.
- HyDE and query decomposition are real and wired in, just switched off. Both fully implemented, both default to disabled in every config I ship, neither one benchmarked yet.
- "Built from scratch" has one asterisk. Chunking still uses LangChain's text splitter. Everything else, retrieval, fusion, reranking, generation, the graph, was written from scratch.
Writing every stage by hand here was mostly a way to actually understand what's happening at each step of a RAG pipeline. Next time I'd probably reach for LangChain, LangGraph, or CrewAI to move faster, now that I know exactly what they're abstracting away underneath.
Python · FastAPI · PostgreSQL (pgvector) · BM25 · cross-encoder · Gemini (Vertex AI) · NetworkX · Gradio