All advisories

Out-of-Bounds Read Through Numeric Lark Token IDs

mlc-ai/xgrammar / GHSA-3m52-c897-q6qj

Affected packages

xgrammar pip
Affected versions>= 0.2.5
Patched versionsNot specified

Description

Out-of-Bounds Read Through Numeric Lark Token IDs

Affected repository: mlc-ai/xgrammar
Observed HEAD: f07ca3cb03affaca98809c4e9dad41deed2f7730

Summary

An application that accepts Lark grammar text and compiles it with XGrammar can be crashed by a numeric special-token ID that is representable as int32_t but is not present in the active tokenizer. On the assessed main revision f07ca3cb03affaca98809c4e9dad41deed2f7730, compiling start: <[2147483647]> with a one-token vocabulary carries 2147483647 into token-mask preprocessing and uses it as a vector index. I reproduced the resulting invalid read and process abort under AddressSanitizer. The demonstrated impact is process-level denial of service; this test does not establish disclosure or code execution.

The Lark entry path first appears in the v0.2.5 release, the earliest affected release I verified. The same source-to-sink path is present in the assessed main revision and the inspected v0.2.6 release candidates. v0.2.4 predates the Lark frontend.

Detail

The root cause is a mismatch between the validation performed by the Lark frontend and the invariant assumed by the grammar compiler. Numeric special tokens are tokenizer-relative: for a tokenizer whose reverse map has N entries, every token ID used as an index must satisfy 0 <= id < N. LarkConverter::ResolveSpecialToken in cpp/lark_converter.cc instead checks only that the parsed range is non-negative, ordered, no larger than INT32_MAX, and not excessively wide:

if (first < 0 || last < first || last > std::numeric_limits<int32_t>::max()) {
  RaiseLarkError(
      source_, location, "invalid numeric special-token range '" + range + "'"
  );
}

Those checks make 2147483647 a valid parser result even when tokenizer_info_ is available and reports a vocabulary size of one. There is no comparison with tokenizer_info_->GetVocabSize(). ResolveSpecialToken returns the value in SpecialTokenSet::token_ids, and CompileNode passes it to GrammarBuilder::AddTokenSet. The grammar normalizer then preserves the same ID in an FSM token edge; no intervening stage establishes the tokenizer-relative bound.

During compilation, GrammarMatcherForTokenMaskCache::GetTokenEdgeAcceptedIndices in cpp/grammar_compiler.cc obtains tid_to_sorted, the vocabulary-sized reverse map from token ID to sorted-vocabulary index. It then treats the untrusted FSM value as an index:

int32_t tid = info.TokenIds()[i];
XGRAMMAR_DCHECK(tid >= 0 && tid < static_cast<int32_t>(tid_to_sorted.size()));
if (tid_to_sorted[tid] >= 0) {
  tmp_token_edge_accepted_.push_back(tid_to_sorted[tid]);
}

XGRAMMAR_DCHECK is not a runtime guard in the shipped default configuration: XGRAMMAR_ENABLE_INTERNAL_CHECK defaults to OFF, so the macro expands to an unreachable while (false) statement. The following operator[] therefore performs an unchecked read. With the PoC's one-entry map, index 2147483647 is far outside the allocation and AddressSanitizer reports the fault at this access. The equivalent kExcludeToken branch has the same missing release-mode check, so the missing release-mode range check affects both branches even though the minimal PoC reaches the ordinary token branch.

This path requires the surrounding application to let an untrusted party influence Lark grammar text and then compile that grammar in the affected process. XGrammar itself does not define that trust boundary. The primitive demonstrated here is an invalid read that aborts the tested process; stronger memory-safety impact is not established.

Reproduce

Fresh validation at current default-branch commit f07ca3cb03affaca98809c4e9dad41deed2f7730 reached the reported sink under AddressSanitizer.

The following disposable Docker recipe checks out the assessed revision, builds it with AddressSanitizer and compiles the malformed grammar. It requires network access to fetch the repository, Debian packages and Python wheels. The final Python command intentionally crashes the container process; no host files are mounted or modified.

docker run --rm -i python:3.12-slim-bookworm bash <<'DOCKER'
set -eux
export DEBIAN_FRONTEND=noninteractive CMAKE_BUILD_PARALLEL_LEVEL=2
apt-get update -qq
apt-get install -y -qq --no-install-recommends ca-certificates git cmake ninja-build g++
python -m pip install -q --index-url https://download.pytorch.org/whl/cpu torch
python -m pip install -q scikit-build-core apache-tvm-ffi pydantic transformers numpy typing-extensions
git clone --depth 1 --quiet https://github.com/mlc-ai/xgrammar.git xgrammar
cd xgrammar
git rev-parse HEAD
git submodule update --init --recursive --depth 1
ASAN_OPTIONS=verify_asan_link_order=0:detect_leaks=0 \
CXXFLAGS='-D_GLIBCXX_NO_ASSERTIONS -fno-lto -fsanitize=address -fno-omit-frame-pointer -g' \
LDFLAGS='-fno-lto -fsanitize=address' \
  python -m pip install -q --config-settings=cmake.build-type=Debug --no-build-isolation --no-deps .
cat > repro.py <<'PY'
import xgrammar as xgr

tokenizer = xgr.TokenizerInfo(["a"])
grammar = xgr.Grammar.from_lark("start: <[2147483647]>", tokenizer_info=tokenizer)
xgr.GrammarCompiler(tokenizer, max_threads=1, cache_enabled=False).compile_grammar(grammar)
PY
ASAN_OPTIONS=abort_on_error=1:detect_leaks=0:symbolize=1:handle_segv=2:use_sigaltstack=0:disable_coredump=1 \
LD_PRELOAD="$(gcc -print-file-name=libasan.so)" \
python repro.py
DOCKER

I ran the ASan build and trigger at the assessed revision. The process aborted with an AddressSanitizer SEGV caused by a read in GrammarMatcherForTokenMaskCache::GetTokenEdgeAcceptedIndices, reached from GetAdaptiveTokenMask during compile_grammar.

Credit

Zheng Yu @ Depthfirst