Skip to content

Fit the example's memory to the Corstone-320 SRAM - #24

Open
MatthiasHertelArm wants to merge 1 commit into
mainfrom
fit-memory-to-sram
Open

MatthiasHertelArm wants to merge 1 commit into
mainfrom
fit-memory-to-sram

Conversation

@MatthiasHertelArm

@MatthiasHertelArm MatthiasHertelArm commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Problem

The image reserves about 9 MB of RAM in DDR4 for a program that needs 12 KB:

Object Reserved Needed by TinyCNN
Method pool 4 MiB 3,840 bytes of planned tensors plus the loaded method
Temp pool 4 MiB 8,704 bytes of Ethos-U scratch
Ethos-U cache buffer 384 KB nothing: only the Dedicated_Sram memory modes use it

The Vela memory mode does not match that layout either. Shared_Sram assumes weights in slow memory and the scratch in SRAM. Here the weights are in the FPGA SRAM and the scratch is in DDR4, so the command stream has the NPU copy a weight stream from SRAM into DDR4 on every inference (DMA_START before the first convolution).

Change

  • Pools follow the program. create_ai_layer.py writes MODEL_PTE_PLANNED_SIZE and MODEL_PTE_SCRATCH_SIZE into model_pte.h; app_main.cpp sizes the two pools from them, plus 16 KiB each. APP_METHOD_POOL_SIZE, APP_TEMP_POOL_SIZE and APP_POOL_SECTION still override.
  • All RAM in SRAM. __RAM0 is the two adjacent SRAM banks VM0 and VM1 (4 MB) instead of DDR4. The linker scripts are unchanged; a model that no longer fits fails at link time.
  • memory: Sram_Only. It describes this layout and is ExecuTorch's own default for an Ethos-U85. The NPU reads the weights in place; the DMA copy is gone.
  • No cache buffer. ETHOS_CACHE_BUF_SIZE: 0 in the board layer; ethos_setup.c then passes no buffer to the driver. Removing the define restores the 384 KB default for Dedicated_Sram_384KB.
  • Two cleanups. ARM_MODEL_USE_PMU_COUNTERS is removed from the board layer (nothing reads it), and the heap comment in regions_SSE-320.h no longer says the planned buffers are on the heap.

Trade-off

Sram_Only encodes the weights for reading in place, so the program grows from 8,832 to 10,960 bytes. The scratch shrinks from 8,704 to 8,192 bytes.

Checked

  • The regenerated ai_layer/ is committed; the command stream is CONV/POOL only, without DMA.
  • The same changes build with AC6 and pass on the Corstone-320 FVP in a derived project with a larger model (SqueezeNet 1.1, 435 KB scratch).
  • CI builds this branch with AC6, GCC and CLANG and runs each image on the FVP: all three pass and print the logits the README already shows, with the 10,960-byte program.

Note for merging

This PR regenerates ai_layer/model_pte.c. Any other open PR that regenerates the layer needs it regenerated again after this one merges.

The image kept 9 MB of RAM in DDR4 for a program that needs 12 KB: two
4 MiB pools and a 384 KB Ethos-U cache buffer nothing uses. The Vela
memory mode Shared_Sram described the opposite of that layout, weights in
slow memory and the scratch in SRAM, so the NPU copied weight streams
from the FPGA SRAM into DDR4 on every inference.

create_ai_layer.py now writes the bytes of the program's memory-planned
tensors and of its Ethos-U scratch into model_pte.h, and app_main.cpp
sizes the method and temp pools from them. With pools of that size all
RAM fits the two SRAM banks VM0 and VM1, so __RAM0 moves there and DDR4
is no longer used. The memory mode becomes Sram_Only, ExecuTorch's own
default for an Ethos-U85: the NPU reads the weights in place. The cache
buffer, which only the Dedicated_Sram modes use, is switched off with
ETHOS_CACHE_BUF_SIZE: 0 in the board layer.

Also removes ARM_MODEL_USE_PMU_COUNTERS from the board layer, which
nothing reads, and corrects the heap comment in regions_SSE-320.h: the
planned buffers are in the method pool, not on the heap.

Sram_Only encodes the weights for reading in place: the program grows
from 8832 to 10960 bytes, the scratch shrinks from 8704 to 8192 bytes.
@MatthiasHertelArm
MatthiasHertelArm marked this pull request as ready for review September 30, 2026 13:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants