All advisories

Out-of-Bounds Write Through an Undersized Tokenizer Vocabulary

mlc-ai/xgrammar / GHSA-87rx-hx49-g8h6

Affected packages

xgrammar pip
Affected versions>= 0.1.33
Patched versionsNot specified

Description

Out-of-Bounds Write Through an Undersized Tokenizer Vocabulary

Affected repository: mlc-ai/xgrammar
Observed HEAD: f07ca3cb03affaca98809c4e9dad41deed2f7730

Summary

The public C++ and Python TokenizerInfo constructors accept a declared vocab_size smaller than the supplied vocabulary. In the assessed main-branch snapshot f07ca3c, XGrammar sizes a token-ID reverse map from that declaration but populates it with IDs derived from every supplied token. Supplying two ordinary tokens with vocab_size=1 therefore writes one int32_t immediately beyond a one-element heap allocation. AddressSanitizer confirms the out-of-bounds write during construction.

The reverse map and this unsafe population loop first appear in v0.1.33; v0.1.32 does not contain them. I confirmed the behavior in the v0.2.5 release, both v0.2.6 release candidates, and the assessed main snapshot. No fixed release was present in the inspected history. The demonstrated impact is memory corruption and process termination for a caller that can supply inconsistent tokenizer metadata. The ordinary Python API makes the bad state reachable, but this report does not establish that an unauthenticated remote request can control tokenizer construction in a particular deployment or that the four-byte write is exploitable beyond denial of service.

Detail

TokenizerInfo intentionally permits vocab_size to be larger than encoded_vocab.size(): model vocabularies may contain padding IDs that are absent from the tokenizer vocabulary. The constructor therefore keeps two related sizes. encoded_vocab.size() determines how many real tokens are decoded and assigns token ID i to encoded_vocab[i], while the optional vocab_size becomes vocab_size_ and determines the advertised model vocabulary size.

The missing invariant is the lower bound on that declaration. Every real token ID must satisfy 0 <= id < vocab_size_, which for the constructor's contiguous IDs is equivalent to requiring vocab_size_ >= encoded_vocab.size(). Neither the Python wrapper nor TokenizerInfo::Impl checks this relationship. The Python wrapper passes both values directly to the native binding, and the C++ constructor starts processing them independently.

In cpp/tokenizer_info.cc, TokenizerInfo::Impl::Impl enumerates every supplied token and preserves its position as the token ID. For {"a", "b"}, neither token is empty or a recognized stop token, so sorted_decoded_vocab_ contains entries with IDs 0 and 1 even when the caller declares vocab_size=1:

sorted_decoded_vocab_.push_back({i, token});

After sorting by token bytes, the constructor allocates token_id_to_sorted_vocab_index_ from the independently controlled vocab_size_. It then uses each preserved token ID as an unchecked std::vector::operator[] index:

token_id_to_sorted_vocab_index_.assign(vocab_size_, -1);
for (int32_t i = 0; i < static_cast<int32_t>(sorted_decoded_vocab_.size()); ++i) {
  token_id_to_sorted_vocab_index_[sorted_decoded_vocab_[i].first] = i;
}

With the reproducer's values, assign(1, -1) allocates four bytes for index 0. The second iteration obtains token ID 1 from sorted_decoded_vocab_ and writes another four-byte integer at index 1, exactly one element past that allocation. Sorting cannot make the ID safe because it changes only entry order, not the stored ID. The padding loop also cannot help: it runs only when vocab_size_ exceeds the supplied vocabulary size. There is no later guard that can contain the corruption because the write occurs synchronously inside construction, before a TokenizerInfo object is returned.

The vulnerable reverse map was added with token-level grammar support in v0.1.33. Later development on the v0.2.6 release-candidate branch moved the same derived-map population into BuildTokenCharData, but retained the missing size invariant and remains affected. The proposed check belongs before either implementation allocates or populates derived vectors, so it applies to the assessed main snapshot and also applies by context to v0.2.6rc2.

Reproduce

Fresh validation at current default-branch commit f07ca3cb03affaca98809c4e9dad41deed2f7730 reached the reported sink under AddressSanitizer.

The following command checks out the exact assessed main snapshot, builds only the public C++ library with AddressSanitizer, and invokes the public constructor with the same two-token/size-1 input accepted by the Python API. The container is disposable; the command does not require Torch, Python bindings, or unpinned Python build dependencies.

docker run --rm -i debian:bookworm-slim bash <<'DOCKER'
set -eux
export DEBIAN_FRONTEND=noninteractive CMAKE_BUILD_PARALLEL_LEVEL=2
apt-get update -qq
apt-get install -y -qq --no-install-recommends ca-certificates git cmake ninja-build g++
git clone --depth 1 --quiet https://github.com/mlc-ai/xgrammar.git xgrammar
cd xgrammar
git rev-parse HEAD
git submodule update --init --depth 1 --quiet 3rdparty/dlpack
mkdir build
cat > build/config.cmake <<'CMAKE'
set(XGRAMMAR_BUILD_PYTHON_BINDINGS OFF)
set(XGRAMMAR_BUILD_CXX_TESTS OFF)
CMAKE
cmake -S . -B build -G Ninja \
  -DCMAKE_BUILD_TYPE=Debug \
  -DCMAKE_CXX_FLAGS='-D_GLIBCXX_NO_ASSERTIONS -fno-lto -fsanitize=address -fno-omit-frame-pointer -g' \
  -DCMAKE_EXE_LINKER_FLAGS='-fno-lto -fsanitize=address'
cmake --build build
cat > repro.cc <<'CPP'
#include <xgrammar/tokenizer_info.h>

int main() {
  xgrammar::TokenizerInfo tokenizer({"a", "b"}, xgrammar::VocabType::RAW, 1);
}
CPP
g++ -std=c++17 -fsanitize=address -fno-omit-frame-pointer -g \
  -Iinclude -I3rdparty/dlpack/include repro.cc build/libxgrammar.a -o repro
ASAN_OPTIONS=abort_on_error=1:detect_leaks=0:symbolize=1:disable_coredump=1 ./repro
DOCKER

I ran this constructor path in an isolated AddressSanitizer build. The relevant output is reproduced below with volatile addresses and absolute build paths omitted:

ERROR: AddressSanitizer: heap-buffer-overflow
WRITE of size 4
0 bytes to the right of 4-byte region
SUMMARY: AddressSanitizer: heap-buffer-overflow cpp/tokenizer_info.cc:304
ABORTING

As a negative control, changing the final constructor argument from 1 to 2 constructs the exact-size vocabulary without an AddressSanitizer finding. A value of 3 also succeeds and confirms that the supported padded-vocabulary case is distinct from the invalid undersized declaration.

Credit

Zheng Yu @ Depthfirst