22 Aug 2026 · 7 min read · project: rocketml

What it takes to make a small model production-shaped

RocketML wraps a TF-IDF and LogisticRegression sentiment classifier, deliberately small and deliberately unremarkable, in CI, a hand-written Helm chart, and real Prometheus and Grafana monitoring. The model isn't the point: getting any small model from a training script to a monitored, Kubernetes-deployed API is. That includes a README number that sat wrong until someone actually checked it against the live API.

Boundary, stated up front: the model here is a TF-IDF and LogisticRegression sentiment classifier trained on 5,000 IMDB reviews, and it's deliberately unremarkable: about 0.86 accuracy, an artifact under 1 MB, inference fast enough on CPU that latency was never a question worth asking. That's on purpose. RocketML isn't trying to prove the model is good. It's a self-service platform for getting any small model from a training script to a monitored, containerized, Kubernetes-deployed API, and the sentiment classifier is just the payload that exercises every stage of that pipeline.
561 MB
Serving image, down from 1.23 GB
0.89
predict_proba, corrected from 0.93
2
Live demos on the same joblib artifact
3
ADRs, one per real decision

How a push becomes a running container

Every push and every pull request runs the same two gates: lint, then test. Only a push to main goes further, and it only goes further once the test job passes:

00
Push
any branch or PR
→
01
Lint
ruff check
→
02
Test
pytest
→
03
Train
main only, ~1-2 min
→
04
Build image
bakes joblib artifact
→
05
Push GHCR
:latest and :sha

The last three stages exist because the serving image can't be assembled from anything sitting in the repo: the artifact is gitignored, so CI has to produce a fresh one before it can bake it in. Every image on main carries a model that was just trained, not one checked in months ago.

Key decision

MLflow's registry is real and it stays the source of truth for lineage: train.py logs metrics and registers every run. But serving never talks to it. Training also writes a plain joblib artifact, the image bakes that in, and a config-driven loader reads it straight off disk, lazily and cached, which is also why /health doesn't depend on the model being loaded at all. docker run the container by itself, with no MLflow service anywhere nearby, and it still predicts, because there's nothing left at runtime for it to depend on. The tradeoff: promoting a model means re-running training and rebuilding the image rather than pointing serving at a new registry version. At this scale that's simpler than adding a live registry client to the request path.

The alternative that was actually running first was heavier than it needed to be, and the obvious-looking fix for that turned out not to work:

ApproachRuntime depsImage sizeOutcome
Full mlflow clientFlask, SQLAlchemy, pandas, mlflow1.23 GBWorks, but the registry client is most of the weight
mlflow-skinny--n/aRejected: metapackage in MLflow 3.x, ships no importable module
Joblib, baked inscikit-learn only561 MBChosen

Getting it onto Kubernetes

The chart in deploy/helm/rocketml isn't helm create output. Writing it by hand, a Deployment and a ClusterIP Service with health probes and resource requests and limits, was the point of that phase: this was my first real Kubernetes and Helm work, and the fastest way to actually understand what each piece of a chart does was to write it rather than generate it. In-cluster, RocketML gets scraped through a ServiceMonitor instead of a static Prometheus config, so the Prometheus Operator that ships with kube-prometheus-stack discovers it on its own.

None of it worked on the first try. A ServiceMonitor without the release: monitoring label gets silently ignored by the Operator's Prometheus, no error, nothing scraped. A helm upgrade that requests more than it limits gets rejected outright at admission, which is a safe failure but not an obvious one the first time it happens. And a local image built with docker build is invisible to a kind cluster until it's explicitly kind loaded, since kind keeps its own image store separate from Docker's.

The number that was wrong

The README's curl example claimed a score of 0.93, predict_proba's confidence for one specific negative review, not an accuracy metric. It was wrong.

What actually happened

I loaded the two joblib artifacts that are actually live right now, the one behind the Hugging Face Space and the one behind the Railway demo, and ran that exact input against each independently. Both returned 0.8949749706947533, which is 0.89, not 0.93. A fresh retrain and a direct call against the live production API both confirmed the same number: four independent checks, one consistent answer, and none of them agreed with what the README said. The fix was a one-line edit in two files, README.md and demo-railway/README.md. What's worth remembering isn't the fix, it's that a number this easy to check, one curl command against a model that's already running in production, sat wrong until someone actually ran it.

Known limitations

Stated plainly rather than left for someone else to find:

Building RocketML wasn't really about the sentiment model. It was about everything a model needs around it before someone else can depend on it: a container that doesn't need a live connection to a registry to answer a request, CI that won't build an image until the tests pass, and a chart I wrote by hand so I'd understand what a rolling deploy actually does instead of trusting a template I didn't write. The roadmap has a few more of these left: Terraform for the cluster itself, ArgoCD for GitOps-style deploys, drift monitoring on the predictions this already logs, and rate limiting on the public endpoint before it actually needs it.

Python · FastAPI · scikit-learn · Docker · GitHub Actions · MLflow · Prometheus · Grafana · Kubernetes · Helm