Fit the example's memory to the Corstone-320 SRAM - #24
Open
MatthiasHertelArm wants to merge 1 commit into
Open
MatthiasHertelArm wants to merge 1 commit into
MatthiasHertelArm wants to merge 1 commit into
Conversation
The image kept 9 MB of RAM in DDR4 for a program that needs 12 KB: two 4 MiB pools and a 384 KB Ethos-U cache buffer nothing uses. The Vela memory mode Shared_Sram described the opposite of that layout, weights in slow memory and the scratch in SRAM, so the NPU copied weight streams from the FPGA SRAM into DDR4 on every inference. create_ai_layer.py now writes the bytes of the program's memory-planned tensors and of its Ethos-U scratch into model_pte.h, and app_main.cpp sizes the method and temp pools from them. With pools of that size all RAM fits the two SRAM banks VM0 and VM1, so __RAM0 moves there and DDR4 is no longer used. The memory mode becomes Sram_Only, ExecuTorch's own default for an Ethos-U85: the NPU reads the weights in place. The cache buffer, which only the Dedicated_Sram modes use, is switched off with ETHOS_CACHE_BUF_SIZE: 0 in the board layer. Also removes ARM_MODEL_USE_PMU_COUNTERS from the board layer, which nothing reads, and corrects the heap comment in regions_SSE-320.h: the planned buffers are in the method pool, not on the heap. Sram_Only encodes the weights for reading in place: the program grows from 8832 to 10960 bytes, the scratch shrinks from 8704 to 8192 bytes.
MatthiasHertelArm
marked this pull request as ready for review
September 30, 2026 13:38
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
The image reserves about 9 MB of RAM in DDR4 for a program that needs 12 KB:
Dedicated_Srammemory modes use itThe Vela memory mode does not match that layout either.
Shared_Sramassumes weights in slow memory and the scratch in SRAM. Here the weights are in the FPGA SRAM and the scratch is in DDR4, so the command stream has the NPU copy a weight stream from SRAM into DDR4 on every inference (DMA_STARTbefore the first convolution).Change
create_ai_layer.pywritesMODEL_PTE_PLANNED_SIZEandMODEL_PTE_SCRATCH_SIZEintomodel_pte.h;app_main.cppsizes the two pools from them, plus 16 KiB each.APP_METHOD_POOL_SIZE,APP_TEMP_POOL_SIZEandAPP_POOL_SECTIONstill override.__RAM0is the two adjacent SRAM banks VM0 and VM1 (4 MB) instead of DDR4. The linker scripts are unchanged; a model that no longer fits fails at link time.memory: Sram_Only. It describes this layout and is ExecuTorch's own default for an Ethos-U85. The NPU reads the weights in place; the DMA copy is gone.ETHOS_CACHE_BUF_SIZE: 0in the board layer;ethos_setup.cthen passes no buffer to the driver. Removing the define restores the 384 KB default forDedicated_Sram_384KB.ARM_MODEL_USE_PMU_COUNTERSis removed from the board layer (nothing reads it), and the heap comment inregions_SSE-320.hno longer says the planned buffers are on the heap.Trade-off
Sram_Onlyencodes the weights for reading in place, so the program grows from 8,832 to 10,960 bytes. The scratch shrinks from 8,704 to 8,192 bytes.Checked
ai_layer/is committed; the command stream isCONV/POOLonly, without DMA.Note for merging
This PR regenerates
ai_layer/model_pte.c. Any other open PR that regenerates the layer needs it regenerated again after this one merges.