Skip to content

Take the model input as NHWC - #26

Merged
MatthiasHertelArm merged 1 commit into
mainfrom
nhwc-input
Sep 30, 2026
Merged

MatthiasHertelArm merged 1 commit into
mainfrom
nhwc-input

Conversation

@MatthiasHertelArm

@MatthiasHertelArm MatthiasHertelArm commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Problem

TinyCNN takes its image as NCHW. The Ethos-U computes on NHWC feature maps, so the command stream of the exported program begins with an operation that does nothing but reorder the input:

POOL.SUM   ifm 3x16x16 NHWC  ->  ofm NHCWB16   (layout conversion)
CONV       ...

On a small model this is cheap. On an image model of realistic size it is not: for a 224x224 SqueezeNet 1.1 the same operation is 15% of the inference cycles Vela estimates, and it was 7.8% of the NPU cycles measured on the Corstone-320 FVP. The example should show the layout that avoids it.

Change

  • model/model.py: INPUT_SHAPE is (1, 16, 16, 3). TinyCNN.forward() permutes the input to the NCHW its convolutions take; Vela folds that into the first convolution.
  • src/app_main.cpp: the input tensor is 1x16x16x3.
  • README: a paragraph in "Adapting the example" on why an image model should take NHWC.

Effect

NCHW NHWC
NPU operations 6 5
Program 8,784 bytes 8,720 bytes
Ethos-U scratch 8,704 bytes 5,376 bytes

The runner's ramp now fills the image in NHWC order, so the logits change; the README shows the new ones.

Checked

  • The regenerated ai_layer/ is committed; the first NPU operation is the convolution, reading the NHWC input directly.
  • CI builds this branch with AC6, GCC and CLANG and runs each image on the FVP: all three pass and print the same logits, which are the ones now in the README.

Note for merging

This PR regenerates ai_layer/model_pte.c. Any other open PR that regenerates the layer needs it regenerated again after this one merges.

TinyCNN took its image as NCHW. The Ethos-U computes on NHWC feature
maps, so the command stream began with an operation that only reorders
the input. The model now takes NHWC, the layout of a camera frame, and
permutes it to NCHW itself; Vela folds that into the first convolution.

The layout conversion is gone from the command stream, the program is
8720 bytes instead of 8784, and the Ethos-U scratch 5376 bytes instead
of 8704. The runner declares its input as 1x16x16x3; its ramp now fills
the image in that order, so the logits differ from the ones before.
@MatthiasHertelArm
MatthiasHertelArm merged commit 7c1d239 into main Sep 30, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants