Two and a half thousand models in each cloud runtime. That number is what makes this decision interesting, because past a certain scale the usual fork in the road stops being a trade-off and starts being arithmetic. Bundle everything into one service and every model shares a fate — one lifecycle, one memory limit, one autoscaling policy — so whichever model is having the worst afternoon decides everyone’s afternoon. Split them into a service per model and the isolation problem disappears, replaced by a bill for 2,500 deployments, nearly all of them idle nearly all of the time.
KServe claims you don’t have to choose. Each model becomes
its own InferenceService with its own autoscaling policy, and when nothing
calls it, it scales to zero — no pods, no cost. Isolation with per-model
economics.
That claim is mostly true. This post is about the “mostly”: what the stack is actually made of, how to stand it up without the three mistakes I made, what happened when I pointed real load at it, and why we ended up not using the feature we came for.
The short version, if you only read one line:
Scale-to-zero doesn’t remove the cost. It moves it from your cloud bill to your p99.
What KServe actually is
The first thing worth understanding is that KServe is not a server. It’s a thin, opinionated layer of Kubernetes CRDs on top of a fairly tall stack, and most of the behaviour people attribute to KServe — the autoscaling, the scale-to-zero, the request buffering — is really Knative underneath it.
From the bottom up:
- cert-manager — issues the TLS certificates KServe’s admission webhooks need. It isn’t optional, and its failures look like unrelated timeouts, which is why it’s worth naming here.
- KServe itself — the controller that turns one
InferenceServiceinto a Knative Service, a Deployment, a storage initialiser that pulls your model artifact, and the routing to reach it.
Why does this matter before you type a single kubectl apply? Because when
something misbehaves — and it will — the useful question is which layer. A
request that hangs for two seconds is the activator. A pod that never becomes
ready is the storage initialiser or your model format. A webhook error on apply
is cert-manager. Knowing the stack turns a mystery into a lookup.
What you give up
Worth being explicit, because the docs are not:
- Kubernetes tax. You are now operating Knative, Istio and cert-manager, each with its own version-compatibility matrix. Upgrading one is a project.
- A latency floor. Even warm, requests pass through the ingress gateway and (depending on configuration) the activator. It is not free.
- Cold starts. The headline feature has a price tag, quantified below.
If you have two models and steady traffic, a plain Deployment behind a HorizontalPodAutoscaler is a completely defensible answer and you should probably stop reading. KServe earns its complexity when you have many models with spiky or sparse traffic — which is exactly the situation where scale-to-zero is worth real money.
Standing it up
What follows is the install that worked, in order, with the reasoning. Versions are pinned deliberately: Knative, Istio and KServe are mutually version-sensitive, and mixing releases is the fastest way to a cluster that comes up but doesn’t serve.
| Component | Version |
|---|---|
| Knative Serving | 1.13.1 |
| net-istio | 1.13.1 |
| cert-manager | 1.14.4 |
| KServe | 0.12.0 |
Prerequisites: a Kubernetes cluster, kubectl, Helm, and a GCS bucket (any
object store KServe supports works — S3 and Azure Blob follow the same shape,
with a different secret).
0. A namespace for the models
kubectl create namespace kserve-test
This namespace is for your workloads. The platform components mostly install elsewhere, which turns out to be the single biggest tripwire in the whole process.
1. Knative Serving
CRDs first, then the controller. Applying them together is a race: the controller’s manifests reference types the API server doesn’t know yet.
KN=github.com/knative/serving/releases/download/knative-v1.13.1
kubectl apply -f https://$KN/serving-crds.yaml
kubectl apply -f https://$KN/serving-core.yaml
2. Istio, as Knative’s networking layer
Knative needs a networking layer to program ingress. Istio is the best-trodden option.
NI=github.com/knative/net-istio/releases/download/knative-v1.13.1
# CRDs first, again
kubectl apply -l knative.dev/crd-install=true -f https://$NI/istio.yaml
kubectl apply -f https://$NI/istio.yaml
# then the controller that teaches Knative to drive Istio
kubectl apply -f https://$NI/net-istio.yaml
Verify before moving on. This is a real checkpoint, not a formality — a half-installed networking layer produces symptoms that look like model problems for the next hour.
kubectl get pods -n knative-serving
You want six pods Running: activator, autoscaler, controller,
webhook, net-istio-controller and net-istio-webhook. If the webhook pods
aren’t ready, stop — everything downstream will fail at admission.
3. cert-manager
KServe’s admission webhooks need certificates. cert-manager issues them.
CM=github.com/cert-manager/cert-manager/releases/download/v1.14.4
kubectl apply -f https://$CM/cert-manager.crds.yaml
helm repo add jetstack https://charts.jetstack.io
helm repo update
helm install cert-manager-kserve jetstack/cert-manager \
--namespace cert-manager --create-namespace \
--version v1.14.4
4. KServe
helm install kserve-crd oci://ghcr.io/kserve/charts/kserve-crd \
--version v0.12.0 -n kserve-test
helm install kserve oci://ghcr.io/kserve/charts/kserve \
--version v0.12.0 --values values.yaml -n kserve-test
5. Credentials for the model store
KServe pulls model artifacts at pod startup using a storage initialiser container. It needs credentials, supplied as a secret. For GCS, that’s a service account key with read access to the bucket:
kubectl create secret generic storage-config \
--from-file=gcloud-application-credentials.json=<path-to-key> \
-n kserve-test
Two non-negotiable details: the key inside the secret must be named
gcloud-application-credentials.json, and the secret must live in the same
namespace as the InferenceService. Get either wrong and the pod sits in
Init:0/1 with a permission error buried in the init container’s logs — a good
five minutes of confusion the first time.
Deploying a model
Upload the artifact, then declare an InferenceService that points at it.
gcloud config set project <your-project-id>
gsutil mb gs://<your-bucket>
gsutil cp model.ubj gs://<your-bucket>/<model-name>/
Two conventions that are easy to miss and produce unhelpful errors:
storageUriis a directory, not a file. Point it at the prefix containing the model, with no filename.- The file must be named
model.<ext>. The runtime discovers the artifact by that name —my_forecast_v3.ubjwill not be found.
apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
name: my-model
namespace: kserve-test
annotations:
serving.kserve.io/secretName: storage-config
spec:
predictor:
minReplicas: 0 # <- this is the whole point
scaleMetric: rps
scaleTarget: 10 # target 10 rps per replica
model:
modelFormat:
name: xgboost
protocolVersion: v2
storageUri: gs://<your-bucket>/<model-name>/
resources:
requests:
cpu: 100m
memory: 200Mi
limits:
cpu: 500m
memory: 500Mi
ports:
- name: h2c # Knative requires this exact name for gRPC
protocol: TCP
containerPort: 9000
readinessProbe:
httpGet:
path: /v2/models/my-model/ready
port: 8080
Three of those lines deserve a sentence each:
minReplicas: 0is the feature. Everything else in this post is downstream of it.name: h2con the port is not decoration. Knative uses the port name to decide protocol, and gRPC over HTTP/2 cleartext must be calledh2c. Name itgrpcand requests fail in a way that points nowhere useful.- The
readinessProbematters more than it looks. Without it, the pod is marked ready as soon as the container starts — but the model may still be loading, so traffic arrives before the model can answer. The next section is entirely about what that looks like.
Then:
kubectl apply -f my-model.yaml -n kserve-test
Load testing, and the number that surprised me
I used ghz for gRPC load generation, running it as a Job
inside the cluster so the measurement isn’t dominated by internet latency. It
needs KServe’s grpc_predict_v2.proto, mounted as a ConfigMap:
kubectl create configmap proto-files \
--from-file=grpc_predict_v2.proto -n kserve-test
The interesting part of the ghz config is the load schedule — a step ramp
rather than a constant rate, because a constant rate lets the autoscaler settle
and hides exactly the behaviour I wanted to see:
{
"load-schedule": "step",
"load-start": 10,
"load-end": 20,
"load-step": 5,
"load-step-duration": "5s",
"timeout": "30s",
"max-duration": "180s"
}
The full driver script deploys the model at a given scaleMetric /
scaleTarget, waits for it to become ready, then runs the Job.
The results
Stepping to 20 rps against a model with minReplicas: 0:
| Metric | Value |
|---|---|
| Requests | 1,809 |
| Achieved throughput | 15.07 rps |
| p50 | 1.70 s |
| p90 | 2.09 s |
| p99 | 2.70 s |
| Fastest | 110 ms |
Successful (OK) |
1,030 (57%) |
And the outcome breakdown, which is the actual story:
| Status | Requests | Share |
|---|---|---|
OK |
1,030 | 56.9% |
InvalidArgument — model not ready |
739 | 40.9% |
Unavailable |
21 | 1.2% |
DeadlineExceeded |
19 | 1.1% |
Forty-three percent of requests failed, and almost all of them for one reason: the model wasn’t loaded yet. The autoscaler did its job — it saw traffic and started a pod. But starting a pod means pulling an image, running the storage initialiser to fetch the artifact from GCS, starting the runtime, and loading the model. Until that finishes, the pod is up and answering — with “not ready”.
Note also the gap between the fastest request (110 ms, a warm pod) and p50 (1.70 s). That spread is the cold start. On a steady-state benchmark it disappears entirely, which is precisely why a steady-state benchmark would have told me the wrong thing.
What the failures actually mean
Three separate problems wearing three confusing error codes:
InvalidArgument— “not ready yet”. The model server is reachable before the model is usable. This is what thereadinessProbeabove fixes: gate readiness on/v2/models/<name>/readyrather than on the process starting, and Knative won’t route to the pod until it can actually answer. Adding it was the single highest-value change I made.DeadlineExceeded. The activator held the request while a pod came up, and the 30 s client timeout expired first. The fix isn’t a longer timeout — it’s a shorter cold start.Unavailable. Connections dropped mid-scale-up. Some of this is unavoidable during a topology change; the rest is Istio and the client not agreeing about connection reuse.
Making cold starts smaller
In rough order of leverage:
- Fix readiness first. Until the probe is right, you’re measuring your own misconfiguration rather than the platform.
- Shrink the image. Image pull is often the largest single term. A slimmer runtime image beats almost any tuning.
- Keep the artifact close. Same-region bucket; small model files. An XGBoost model is kilobytes, so this term is negligible here — for anything with weights in the gigabytes it dominates everything else.
- Then choose your poison.
minReplicas: 1eliminates cold starts and the savings with them. Knative’s scale-down delay (scale-to-zero-grace-period) is the middle ground: stay warm through gaps in bursty traffic, still go to zero overnight.
So: worth it?
Here is the honest ending: we kept KServe and turned off the feature we came for.
minReplicas: 0 is not set anywhere in production. The deployment runs a
constant four pods, always warm, and the cold-start problem is solved the
expensive way — by never having one.
Which sounds like a failed experiment, and isn’t. The section above is the result. Forty-three percent of requests failing on a cold start is not a number you tune away next sprint; it’s an answer. Finding it in a load test cost an afternoon. Finding it in production, in front of callers who had been told the model was available, would have cost considerably more — and we would have found it eventually either way, because scale-to-zero doesn’t degrade gracefully. It works perfectly right up until traffic arrives at an empty deployment.
So the value wasn’t the saving. It was learning the price of the saving before committing to it, and then deciding not to pay.
The thing I’d want to know before starting, and didn’t: cold starts are not an implementation detail you tune away later — they’re a product decision you make up front. Either your callers tolerate a multi-second first request, or you keep something warm and hand back part of the saving. KServe makes that trade-off cheap to express, and cheap to measure. It does not make it go away.
We made ours. It just wasn’t the one the feature list suggested.
Comments