Writing
Blog
1 Sept 2026 · 8 min read
What 35 documents got wrong about arithmetic self-correction
A 35-document eval said self-correction was a clear win over plain extraction. At the full 361-document SROIE split, gemini overtook anthropic as the stronger single-call backend, the self-correcting backend's real edge dropped to +2.5 points, and most of its misses turned out to be the tool reconciling an already-correct total into a wrong one.
- Evals
- Sample Size
- DocExtract
22 Aug 2026 · 7 min read
What it takes to make a small model production-shaped
A TF-IDF and LogisticRegression classifier wrapped in CI, a hand-written Helm chart, and real Prometheus and Grafana monitoring, built to prove the platform works, not the model. The image dropped from 1.23 GB to 561 MB by cutting MLflow out of the serving runtime, and a stale confidence score in the README sat wrong until a direct check against the live API caught it.
- Kubernetes
- CI/CD
- RocketML
17 Aug 2026 · 9 min read
What a RAG pipeline actually looks like when you measure it
Hybrid search and cross-encoder re-ranking, built without LangChain's retrieval abstractions, plus a knowledge graph that's wired in but has never once fired. The numbers behind every retrieval decision, including one I got wrong for months without noticing.
- RAG
- Retrieval
- GoFetch