Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 26 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -73,7 +73,7 @@ Ethos-U version info:
Arch: v2.0.0
MACs/cc: 256
Cmd stream: v1
ExecuTorch Ethos-U85 example: 8832 byte model
ExecuTorch Ethos-U85 example: 10960 byte model
Output: 10 element(s): 0.0079 0.0459 0.0475 -0.0475 0.0791 0.0411 -0.0285 -0.0744 -0.2246 -0.0016
Test_result: PASS
```
Expand Down Expand Up @@ -172,7 +172,7 @@ mlops:
type: Ethos-U85
vela:
system: Ethos_U85_SYS_DRAM_Mid
memory: Shared_Sram
memory: Sram_Only
model:
clayer: $AI-Layer$
name: TinyCNN
Expand All @@ -185,11 +185,29 @@ ExecuTorch's `EthosUCompileSpec`, so the target configuration is never
duplicated in Python. The script then writes:

- `ai_layer/ai_layer.clayer.yml`: the CMSIS components required by the model.
- `ai_layer/model_pte.c` and `model_pte.h`: the ExecuTorch program embedded as a C array.
- `ai_layer/model_pte.c` and `model_pte.h`: the ExecuTorch program embedded as
a C array, and the run-time memory it needs (its memory-planned tensors and
the Ethos-U scratch), from which the application sizes its two pools.
- `ai_layer/model.pte`: the program itself, for inspection.

See [the MLOps flow](documentation/mlops-flow.md) for a detailed walkthrough.

### Fitting the model to the target

The Vela memory mode and the memory the application reserves follow from where
the image puts things on the Corstone-320 FVP
([`regions_SSE-320.h`](board/Corstone-320/regions_SSE-320.h)):

- **`memory: Sram_Only`:** code and the program with its weights are in the
2 MB FPGA SRAM, all RAM in the two 2 MB SRAM banks. The NPU reads the weights
where they are. `Shared_Sram` assumes weights in slower memory and lets the
NPU copy weight streams into the scratch first, on every inference.
- **Pools from the program:** `model_pte.h` names the bytes of the program's
memory-planned tensors and of its Ethos-U scratch, and `src/app_main.cpp`
sizes the method and temp pools from them, so they follow the model.
- **No Ethos-U cache buffer:** only the `Dedicated_Sram` memory modes use the
driver's 384 KB fast scratch; the board layer sets `ETHOS_CACHE_BUF_SIZE` to 0.

## Component selection

The [ExecuTorch CMSIS Pack](https://www.keil.arm.com/packs/executorch-pytorch/)
Expand All @@ -213,7 +231,10 @@ and build.

To target another Ethos-U configuration, update the target and `mlops:`
settings in the CMSIS solution and re-run all three steps. The generated Vela
options then follow that configuration automatically. Moving to a different
options then follow that configuration automatically. Choose the memory mode
for where the target keeps the weights and the scratch, and give a
`Dedicated_Sram` mode its cache buffer (`ETHOS_CACHE_BUF_SIZE` in the board
layer). Moving to a different
board or reference platform also requires the corresponding device pack, board
support, memory layout, and FVP configuration.

Expand All @@ -234,7 +255,7 @@ together. More information is available in
| `.vscode.d/tasks.json` | The VS Code tasks (venv setup, Create AI layer) merged by the CMSIS Solution extension |
| `.vscode/fvp.sh`, `.vscode/fvp.Dockerfile` | The FVP model command used by Run and Debug; runs the model in Docker on macOS |
| `board/Corstone-320/` | Corstone-320 platform support and FVP configuration |
| `src/app_main.cpp` | Loads the model, runs inference, and prints the result; pool sizes overridable with `APP_METHOD_POOL_SIZE`, `APP_TEMP_POOL_SIZE`, `APP_POOL_SECTION` |
| `src/app_main.cpp` | Loads the model, runs inference, and prints the result; the pools are sized from `model_pte.h`, overridable with `APP_METHOD_POOL_SIZE`, `APP_TEMP_POOL_SIZE`, `APP_POOL_SECTION` |
| `src/arm_embedded_module.*` | `EmbeddedModule`: ExecuTorch's `Module` class without the POSIX file loading (BSD-3-Clause, `src/LICENSE-ExecuTorch`) |
| `documentation/mlops-flow.md` | The MLOps flow in detail |
| `documentation/pack-provenance.md` | Where the ExecuTorch pack comes from, how to update it |
Expand Down
825 changes: 479 additions & 346 deletions ai_layer/model_pte.c

Large diffs are not rendered by default.

6 changes: 6 additions & 0 deletions ai_layer/model_pte.h
Original file line number Diff line number Diff line change
@@ -1,5 +1,11 @@
// Generated by create_ai_layer.py -- do not edit.
#pragma once

// Run-time memory of the program: its memory-planned tensors, and the scratch
// the Ethos-U backend takes from the temp allocator for an inference.
#define MODEL_PTE_PLANNED_SIZE 3840
#define MODEL_PTE_SCRATCH_SIZE 8192

#ifdef __cplusplus
extern "C" {
#endif
Expand Down
5 changes: 4 additions & 1 deletion board/Corstone-320/Board-U85.clayer.yml
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,10 @@ layer:
define:
- CMSIS_target_header: \"Corstone-320.h\"
- ETHOSU85
- ARM_MODEL_USE_PMU_COUNTERS
# No Ethos-U cache buffer: only the Dedicated_Sram memory modes of Vela use
# one, and the csolution's mlops: node selects Sram_Only. Remove this
# define for the 384 KB default that Dedicated_Sram_384KB needs.
- ETHOS_CACHE_BUF_SIZE: 0

packs:
- pack: ARM::CMSIS
Expand Down
7 changes: 4 additions & 3 deletions board/Corstone-320/ethos_setup.c
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,8 @@
#include "main.h"

#if defined(ETHOSU65) || defined(ETHOSU85)
/* Define Ethos-U NPU cache buffer size */
/* Define Ethos-U NPU cache buffer size: the fast scratch (base address 2) of a
model compiled for a Dedicated_Sram memory mode, 0 for no buffer. */
#ifndef ETHOS_CACHE_BUF_SIZE
#define ETHOS_CACHE_BUF_SIZE 393216
#endif
Expand Down Expand Up @@ -54,7 +55,7 @@
/* Ethos NPU driver instance. */
static struct ethosu_driver EthosDriver;

#if defined(ETHOSU65) || defined(ETHOSU85)
#if (defined(ETHOSU65) || defined(ETHOSU85)) && (ETHOS_CACHE_BUF_SIZE > 0)
static uint8_t ethos_cache[ETHOS_CACHE_BUF_SIZE] ETHOS_CACHE_BUF_ATTRIBUTES;
#endif

Expand All @@ -76,7 +77,7 @@ void ethos_setup (void) {
/* Initialize Ethos-U NPU driver. */
rval = ethosu_init(&EthosDriver, /* Ethos-U device driver */
ethos_base_addr, /* Ethos-U base address */
#if defined(ETHOSU65) || defined(ETHOSU85)
#if (defined(ETHOSU65) || defined(ETHOSU85)) && (ETHOS_CACHE_BUF_SIZE > 0)
ethos_cache, /* Cache memory pointer */
sizeof(ethos_cache), /* Cache memory size */
#else
Expand Down
24 changes: 14 additions & 10 deletions board/Corstone-320/regions_SSE-320.h
Original file line number Diff line number Diff line change
Expand Up @@ -69,35 +69,38 @@
// <h> __RAM0
// <y> Base address
// <i> Defines base address of memory region.
// <i> Pack default: 0x10000000 (ITCM). 256 MB DDR4 here holds .data/.bss, the two
// <i> 4 MB inference pools, the Ethos-U cache buffer, heap and stack
#define __RAM0_BASE DDR4_1_S_BASE
// <i> Pack default: 0x10000000 (ITCM). The two adjacent 2 MB SRAM banks VM0 and
// <i> VM1 here hold .data/.bss, the inference pools (with the Ethos-U scratch),
// <i> heap and stack: all of the example's RAM. The Vela memory mode Sram_Only
// <i> in the csolution's mlops: node relies on the scratch being in SRAM.
#define __RAM0_BASE SRAM_VM0_S_BASE
// <y> Region size [bytes]
// <i> Defines size of memory region.
// <i> Pack default: 0x00008000
#define __RAM0_SIZE DDR4_1_S_SIZE
#define __RAM0_SIZE (SRAM_VM0_S_SIZE + SRAM_VM1_S_SIZE)
// </h>

// <h> __RAM1
// <y> Base address
// <i> Defines base address of memory region.
// <i> Pack default: 0x12000000 (FPGA SRAM)
#define __RAM1_BASE SRAM_VM0_S_BASE
// <i> Pack default: 0x12000000 (FPGA SRAM). 256 MB DDR4 here, unused: the
// <i> linker scripts place everything in __RAM0 and fail when it is full.
#define __RAM1_BASE DDR4_1_S_BASE
// <y> Region size [bytes]
// <i> Defines size of memory region.
// <i> Pack default: 0x00200000
#define __RAM1_SIZE SRAM_VM0_S_SIZE
#define __RAM1_SIZE DDR4_1_S_SIZE
// </h>

// <h> __RAM2
// <y> Base address
// <i> Defines base address of memory region.
// <i> Pack default: 0x30000000 (DTCM)
// <i> Pack default: 0x30000000 (DTCM). Unused: SRAM VM1 is part of __RAM0.
#define __RAM2_BASE SRAM_VM1_S_BASE
// <y> Region size [bytes]
// <i> Defines size of memory region.
// <i> Pack default: 0x00008000
#define __RAM2_SIZE SRAM_VM1_S_SIZE
#define __RAM2_SIZE 0
// </h>

// <h> __RAM3
Expand All @@ -115,7 +118,8 @@

// <h> Stack / Heap Configuration
// <i> Pack defaults: 0x600 stack, 0xC00 heap. The runner's EmbeddedModule
// <i> keeps its method table and planned buffers on the heap.
// <i> keeps the program and its method table on the heap; the planned
// <i> buffers are in the method pool of app_main.cpp.
// <o0> Stack Size (in Bytes) <0x0-0xFFFFFFFF:8>
// <o1> Heap Size (in Bytes) <0x0-0xFFFFFFFF:8>
#define __STACK_SIZE 0x00001000
Expand Down
5 changes: 4 additions & 1 deletion cmsis-executorch.csolution.yml
Original file line number Diff line number Diff line change
Expand Up @@ -79,7 +79,10 @@ solution:
type: Ethos-U85
vela:
system: Ethos_U85_SYS_DRAM_Mid # system-config from the Vela config
memory: Shared_Sram # memory-mode from the Vela config
# memory-mode from the Vela config. Sram_Only: weights and scratch are
# both in SRAM on this target (regions_SSE-320.h), so the NPU reads the
# weights in place. Shared_Sram would copy them to the scratch first.
memory: Sram_Only
model:
clayer: $AI-Layer$
name: TinyCNN
Expand Down
43 changes: 40 additions & 3 deletions create_ai_layer.py
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,8 @@
ai_layer.clayer.yml runtime, kernel utilities and registration, Ethos-U
backend and the operator components the exported
program actually uses
model_pte.c / .h the ExecuTorch program as a C array
model_pte.c / .h the ExecuTorch program as a C array, and the
run-time memory it needs
model.pte the program itself, for inspection

The script runs itself in the solution's .venv (see setup_venv.py) when it is
Expand Down Expand Up @@ -202,8 +203,42 @@ def c_array(pte: bytes) -> str:
)


HEADER = f"""// Generated by create_ai_layer.py -- do not edit.
def memory_needs(pte: bytes) -> tuple[int, int]:
"""Bytes of the program's memory-planned buffers, and of the largest Ethos-U scratch among its delegates.

The application sizes its method and temp allocator pools from them. The
scratch size is a block of the delegate blob (16-byte name, 32-bit size,
data from offset 32, padded to 16 bytes; see VelaBinStream.cpp in the pack).
"""
from executorch.exir._serialize._program import deserialize_pte_binary

result = deserialize_pte_binary(pte)
plan = getattr(result, "program", result).execution_plan[0]
scratch = 0
for stream in re.finditer(rb"vela_bin_stream\0", pte):
pos = stream.start()
while True:
name = pte[pos : pos + 16].split(b"\0")[0]
size = int.from_bytes(pte[pos + 16 : pos + 20], "little")
if name == b"scratch_size":
scratch = max(scratch, int.from_bytes(pte[pos + 32 : pos + 36], "little"))
if name == b"vela_end_stream":
break
pos += 32 + ((size + 15) & ~15)
return sum(plan.non_const_buffer_sizes), scratch


def header(pte: bytes) -> str:
planned, scratch = memory_needs(pte)
macro = SYMBOL.upper()
return f"""// Generated by create_ai_layer.py -- do not edit.
#pragma once

// Run-time memory of the program: its memory-planned tensors, and the scratch
// the Ethos-U backend takes from the temp allocator for an inference.
#define {macro}_PLANNED_SIZE {planned}
#define {macro}_SCRATCH_SIZE {scratch}

#ifdef __cplusplus
extern "C" {{
#endif
Expand Down Expand Up @@ -266,10 +301,12 @@ def main() -> None:
layer_dir.mkdir(parents=True, exist_ok=True)
(layer_dir / "model.pte").write_bytes(pte)
(layer_dir / f"{SYMBOL}.c").write_text(c_array(pte), newline="\n")
(layer_dir / f"{SYMBOL}.h").write_text(HEADER, newline="\n")
(layer_dir / f"{SYMBOL}.h").write_text(header(pte), newline="\n")
layer_file.write_text(clayer(mlops, runtime, operators, mlops_file, version), newline="\n")

planned, scratch = memory_needs(pte)
print(f"[ai_layer] {len(pte)} byte program, operators: {operators}")
print(f"[ai_layer] run-time memory: {planned} bytes planned tensors, {scratch} bytes Ethos-U scratch")
print(f"[ai_layer] wrote {layer_file}")


Expand Down
6 changes: 3 additions & 3 deletions documentation/mlops-flow.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,7 @@ solution:
type: Ethos-U85
vela:
system: Ethos_U85_SYS_DRAM_Mid # system-config from the Vela config
memory: Shared_Sram # memory-mode from the Vela config
memory: Sram_Only # memory-mode from the Vela config
model:
clayer: $AI-Layer$
name: TinyCNN
Expand All @@ -51,7 +51,7 @@ cbuild-mlops:
npu:
type: Ethos-U85
vela:
options: --system-config Ethos_U85_SYS_DRAM_Mid --memory-mode Shared_Sram
options: --system-config Ethos_U85_SYS_DRAM_Mid --memory-mode Sram_Only
model:
clayer: ai_layer/ai_layer.clayer.yml
name: TinyCNN
Expand Down Expand Up @@ -97,7 +97,7 @@ for an MLOps system. It reads the file and:
| File | Content |
|------|---------|
| `ai_layer/ai_layer.clayer.yml` | runtime, kernel utilities and registration, Ethos-U backend and the operator components the program uses |
| `ai_layer/model_pte.c` / `.h` | the program as a 16-byte-aligned C array, `model_pte` / `model_pte_size` |
| `ai_layer/model_pte.c` / `.h` | the program as a 16-byte-aligned C array, `model_pte` / `model_pte_size`, and its run-time memory, `MODEL_PTE_PLANNED_SIZE` / `MODEL_PTE_SCRATCH_SIZE`, for the application's pools |
| `ai_layer/model.pte` | the program itself, for inspection (not committed) |

The clayer and the C array are committed, so a checkout builds without
Expand Down
14 changes: 8 additions & 6 deletions src/app_main.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -36,15 +36,17 @@ using executorch::runtime::MemoryAllocator;

namespace {

// Pool sizes and placement are board-overridable: a board layer that cannot
// fit 8 MB of .bss in its default RAM sets APP_*_POOL_SIZE and, via
// APP_POOL_SECTION, the linker section its linker script routes to a larger
// (NPU-accessible) memory.
// The pools follow the program: model_pte.h names the bytes of its
// memory-planned tensors, which the method pool holds next to the loaded
// method itself, and of the Ethos-U scratch, which is drawn from the temp pool
// for each inference. Sizes and placement are board-overridable: a board layer
// sets APP_*_POOL_SIZE and, via APP_POOL_SECTION, the linker section its
// linker script routes to another (NPU-accessible) memory.
#ifndef APP_METHOD_POOL_SIZE
#define APP_METHOD_POOL_SIZE (4 * 1024 * 1024)
#define APP_METHOD_POOL_SIZE (MODEL_PTE_PLANNED_SIZE + 16 * 1024)
#endif
#ifndef APP_TEMP_POOL_SIZE
#define APP_TEMP_POOL_SIZE (4 * 1024 * 1024) // Ethos-U scratch is drawn from here.
#define APP_TEMP_POOL_SIZE (MODEL_PTE_SCRATCH_SIZE + 16 * 1024)
#endif
#ifdef APP_POOL_SECTION
#define APP_POOL_ATTRIBUTES __attribute__((section(APP_POOL_SECTION)))
Expand Down
Loading