← Yang Cao 11 min read 02.04.2024

Serving models on Kubernetes with KServe

Scale-to-zero sounds free until you meet the cold start. What KServe actually buys you, what it quietly costs, and why we turned the headline feature off in the end.

Two and a half thousand models in each cloud runtime. That number is what makes this decision interesting, because past a certain scale the usual fork in the road stops being a trade-off and starts being arithmetic. Bundle everything into one service and every model shares a fate — one lifecycle, one memory limit, one autoscaling policy — so whichever model is having the worst afternoon decides everyone’s afternoon. Split them into a service per model and the isolation problem disappears, replaced by a bill for 2,500 deployments, nearly all of them idle nearly all of the time.

KServe claims you don’t have to choose. Each model becomes its own InferenceService with its own autoscaling policy, and when nothing calls it, it scales to zero — no pods, no cost. Isolation with per-model economics.

That claim is mostly true. This post is about the “mostly”: what the stack is actually made of, how to stand it up without the three mistakes I made, what happened when I pointed real load at it, and why we ended up not using the feature we came for.

The short version, if you only read one line:

Scale-to-zero doesn’t remove the cost. It moves it from your cloud bill to your p99.


What KServe actually is

The first thing worth understanding is that KServe is not a server. It’s a thin, opinionated layer of Kubernetes CRDs on top of a fairly tall stack, and most of the behaviour people attribute to KServe — the autoscaling, the scale-to-zero, the request buffering — is really Knative underneath it.

From the bottom up:

CONTROL PLANE REQUEST PATH Knative autoscaler Client request Istio ingress gateway Knative activator Model server pod KPA · RPS OR CONCURRENCY gRPC h2c · ModelInfer ROUTES INTO THE MESH ONLY WHEN REPLICAS = 0 HOLDS THE REQUEST, WAKES THE POD MLSERVER / TRITON / TORCHSERVE READINESS GATES THE TRAFFIC Response WARM PATH REPLICAS > 0 SKIPS THE ACTIVATOR SCALES
Fig. 1 — one inference request. The activator only appears when replicas are at zero; everything warm skips it.
And sitting beside all of it:

  • cert-manager — issues the TLS certificates KServe’s admission webhooks need. It isn’t optional, and its failures look like unrelated timeouts, which is why it’s worth naming here.
  • KServe itself — the controller that turns one InferenceService into a Knative Service, a Deployment, a storage initialiser that pulls your model artifact, and the routing to reach it.

Why does this matter before you type a single kubectl apply? Because when something misbehaves — and it will — the useful question is which layer. A request that hangs for two seconds is the activator. A pod that never becomes ready is the storage initialiser or your model format. A webhook error on apply is cert-manager. Knowing the stack turns a mystery into a lookup.

What you give up

Worth being explicit, because the docs are not:

  • Kubernetes tax. You are now operating Knative, Istio and cert-manager, each with its own version-compatibility matrix. Upgrading one is a project.
  • A latency floor. Even warm, requests pass through the ingress gateway and (depending on configuration) the activator. It is not free.
  • Cold starts. The headline feature has a price tag, quantified below.

If you have two models and steady traffic, a plain Deployment behind a HorizontalPodAutoscaler is a completely defensible answer and you should probably stop reading. KServe earns its complexity when you have many models with spiky or sparse traffic — which is exactly the situation where scale-to-zero is worth real money.


Standing it up

What follows is the install that worked, in order, with the reasoning. Versions are pinned deliberately: Knative, Istio and KServe are mutually version-sensitive, and mixing releases is the fastest way to a cluster that comes up but doesn’t serve.

Component Version
Knative Serving 1.13.1
net-istio 1.13.1
cert-manager 1.14.4
KServe 0.12.0

Prerequisites: a Kubernetes cluster, kubectl, Helm, and a GCS bucket (any object store KServe supports works — S3 and Azure Blob follow the same shape, with a different secret).

0. A namespace for the models

kubectl create namespace kserve-test

This namespace is for your workloads. The platform components mostly install elsewhere, which turns out to be the single biggest tripwire in the whole process.

1. Knative Serving

CRDs first, then the controller. Applying them together is a race: the controller’s manifests reference types the API server doesn’t know yet.

KN=github.com/knative/serving/releases/download/knative-v1.13.1

kubectl apply -f https://$KN/serving-crds.yaml
kubectl apply -f https://$KN/serving-core.yaml

2. Istio, as Knative’s networking layer

Knative needs a networking layer to program ingress. Istio is the best-trodden option.

NI=github.com/knative/net-istio/releases/download/knative-v1.13.1

# CRDs first, again
kubectl apply -l knative.dev/crd-install=true -f https://$NI/istio.yaml
kubectl apply -f https://$NI/istio.yaml

# then the controller that teaches Knative to drive Istio
kubectl apply -f https://$NI/net-istio.yaml

Verify before moving on. This is a real checkpoint, not a formality — a half-installed networking layer produces symptoms that look like model problems for the next hour.

kubectl get pods -n knative-serving

You want six pods Running: activator, autoscaler, controller, webhook, net-istio-controller and net-istio-webhook. If the webhook pods aren’t ready, stop — everything downstream will fail at admission.

3. cert-manager

KServe’s admission webhooks need certificates. cert-manager issues them.

CM=github.com/cert-manager/cert-manager/releases/download/v1.14.4

kubectl apply -f https://$CM/cert-manager.crds.yaml

helm repo add jetstack https://charts.jetstack.io
helm repo update
helm install cert-manager-kserve jetstack/cert-manager \
  --namespace cert-manager --create-namespace \
  --version v1.14.4

4. KServe

helm install kserve-crd oci://ghcr.io/kserve/charts/kserve-crd \
  --version v0.12.0 -n kserve-test

helm install kserve oci://ghcr.io/kserve/charts/kserve \
  --version v0.12.0 --values values.yaml -n kserve-test

5. Credentials for the model store

KServe pulls model artifacts at pod startup using a storage initialiser container. It needs credentials, supplied as a secret. For GCS, that’s a service account key with read access to the bucket:

kubectl create secret generic storage-config \
  --from-file=gcloud-application-credentials.json=<path-to-key> \
  -n kserve-test

Two non-negotiable details: the key inside the secret must be named gcloud-application-credentials.json, and the secret must live in the same namespace as the InferenceService. Get either wrong and the pod sits in Init:0/1 with a permission error buried in the init container’s logs — a good five minutes of confusion the first time.


Deploying a model

Upload the artifact, then declare an InferenceService that points at it.

gcloud config set project <your-project-id>
gsutil mb gs://<your-bucket>
gsutil cp model.ubj gs://<your-bucket>/<model-name>/

Two conventions that are easy to miss and produce unhelpful errors:

  1. storageUri is a directory, not a file. Point it at the prefix containing the model, with no filename.
  2. The file must be named model.<ext>. The runtime discovers the artifact by that name — my_forecast_v3.ubj will not be found.
apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
  name: my-model
  namespace: kserve-test
  annotations:
    serving.kserve.io/secretName: storage-config
spec:
  predictor:
    minReplicas: 0          # <- this is the whole point
    scaleMetric: rps
    scaleTarget: 10         # target 10 rps per replica
    model:
      modelFormat:
        name: xgboost
      protocolVersion: v2
      storageUri: gs://<your-bucket>/<model-name>/
      resources:
        requests:
          cpu: 100m
          memory: 200Mi
        limits:
          cpu: 500m
          memory: 500Mi
      ports:
        - name: h2c        # Knative requires this exact name for gRPC
          protocol: TCP
          containerPort: 9000
      readinessProbe:
        httpGet:
          path: /v2/models/my-model/ready
          port: 8080

Three of those lines deserve a sentence each:

  • minReplicas: 0 is the feature. Everything else in this post is downstream of it.
  • name: h2c on the port is not decoration. Knative uses the port name to decide protocol, and gRPC over HTTP/2 cleartext must be called h2c. Name it grpc and requests fail in a way that points nowhere useful.
  • The readinessProbe matters more than it looks. Without it, the pod is marked ready as soon as the container starts — but the model may still be loading, so traffic arrives before the model can answer. The next section is entirely about what that looks like.

Then:

kubectl apply -f my-model.yaml -n kserve-test

Load testing, and the number that surprised me

I used ghz for gRPC load generation, running it as a Job inside the cluster so the measurement isn’t dominated by internet latency. It needs KServe’s grpc_predict_v2.proto, mounted as a ConfigMap:

kubectl create configmap proto-files \
  --from-file=grpc_predict_v2.proto -n kserve-test

The interesting part of the ghz config is the load schedule — a step ramp rather than a constant rate, because a constant rate lets the autoscaler settle and hides exactly the behaviour I wanted to see:

{
  "load-schedule": "step",
  "load-start": 10,
  "load-end": 20,
  "load-step": 5,
  "load-step-duration": "5s",
  "timeout": "30s",
  "max-duration": "180s"
}

The full driver script deploys the model at a given scaleMetric / scaleTarget, waits for it to become ready, then runs the Job.

The results

Stepping to 20 rps against a model with minReplicas: 0:

p50
1.70 s
p99
2.70 s
succeeded
57 %
Fig. 2 — ghz, step load 10→20 rps over 180 s, minReplicas 0
Metric Value
Requests 1,809
Achieved throughput 15.07 rps
p50 1.70 s
p90 2.09 s
p99 2.70 s
Fastest 110 ms
Successful (OK) 1,030 (57%)
450300 1500 REQUESTS 453 0.11 0.67 1.23 1.78 2.34 2.90 P50 1.70s P99 2.70s RESPONSE TIME (SECONDS)
Fig. 2 — response time of the 1,030 requests that succeeded. One series, so one colour and no legend; only the peak carries a number.

And the outcome breakdown, which is the actual story:

OK · 1,030 NOT READY · 739 OK InvalidArgument Unavailable DeadlineExceeded
Fig. 3 — outcome by gRPC status code. The two thin segments are too narrow to label, so the table below carries them.
Status Requests Share
OK 1,030 56.9%
InvalidArgument — model not ready 739 40.9%
Unavailable 21 1.2%
DeadlineExceeded 19 1.1%

Forty-three percent of requests failed, and almost all of them for one reason: the model wasn’t loaded yet. The autoscaler did its job — it saw traffic and started a pod. But starting a pod means pulling an image, running the storage initialiser to fetch the artifact from GCS, starting the runtime, and loading the model. Until that finishes, the pod is up and answering — with “not ready”.

Note also the gap between the fastest request (110 ms, a warm pod) and p50 (1.70 s). That spread is the cold start. On a steady-state benchmark it disappears entirely, which is precisely why a steady-state benchmark would have told me the wrong thing.

What the failures actually mean

Three separate problems wearing three confusing error codes:

  • InvalidArgument — “not ready yet”. The model server is reachable before the model is usable. This is what the readinessProbe above fixes: gate readiness on /v2/models/<name>/ready rather than on the process starting, and Knative won’t route to the pod until it can actually answer. Adding it was the single highest-value change I made.
  • DeadlineExceeded. The activator held the request while a pod came up, and the 30 s client timeout expired first. The fix isn’t a longer timeout — it’s a shorter cold start.
  • Unavailable. Connections dropped mid-scale-up. Some of this is unavoidable during a topology change; the rest is Istio and the client not agreeing about connection reuse.

Making cold starts smaller

In rough order of leverage:

  1. Fix readiness first. Until the probe is right, you’re measuring your own misconfiguration rather than the platform.
  2. Shrink the image. Image pull is often the largest single term. A slimmer runtime image beats almost any tuning.
  3. Keep the artifact close. Same-region bucket; small model files. An XGBoost model is kilobytes, so this term is negligible here — for anything with weights in the gigabytes it dominates everything else.
  4. Then choose your poison. minReplicas: 1 eliminates cold starts and the savings with them. Knative’s scale-down delay (scale-to-zero-grace-period) is the middle ground: stay warm through gaps in bursty traffic, still go to zero overnight.

So: worth it?

Here is the honest ending: we kept KServe and turned off the feature we came for.

minReplicas: 0 is not set anywhere in production. The deployment runs a constant four pods, always warm, and the cold-start problem is solved the expensive way — by never having one.

Which sounds like a failed experiment, and isn’t. The section above is the result. Forty-three percent of requests failing on a cold start is not a number you tune away next sprint; it’s an answer. Finding it in a load test cost an afternoon. Finding it in production, in front of callers who had been told the model was available, would have cost considerably more — and we would have found it eventually either way, because scale-to-zero doesn’t degrade gracefully. It works perfectly right up until traffic arrives at an empty deployment.

So the value wasn’t the saving. It was learning the price of the saving before committing to it, and then deciding not to pay.

The thing I’d want to know before starting, and didn’t: cold starts are not an implementation detail you tune away later — they’re a product decision you make up front. Either your callers tolerate a multi-second first request, or you keep something warm and hand back part of the saving. KServe makes that trade-off cheap to express, and cheap to measure. It does not make it go away.

We made ours. It just wasn’t the one the feature list suggested.

Comments