Skip to content

Replace TinyCNN with pico-faces on the hackathon branch - #22

Open
MatthiasHertelArm wants to merge 5 commits into
hackathonfrom
hackathon-pico-faces
Open

MatthiasHertelArm wants to merge 5 commits into
hackathonfrom
hackathon-pico-faces

Conversation

@MatthiasHertelArm

@MatthiasHertelArm MatthiasHertelArm commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

The hackathon example becomes pico-faces (MIT): a latent diffusion transformer plus a small decoder that generate 128x128 faces. Both methods of the exported program, dit_step and decode, run entirely on the Ethos-U85, and the Cortex-M runs the sampler. On the DevKit-E8 every face appears on the LCD, the SW2 joystick generates new ones, and pico-faces' serial viewer works against the console.

The branch also moves to the current tools and adds a fix for the SETOOLS debug-stub task.

Changes

Model and export

  • model/model.py: the pico-faces modules, re-implemented with operators the Ethos-U delegate supports. They load the upstream checkpoints, which setup_venv.py downloads from a pinned upstream commit and checks by SHA-256.
  • create_ai_layer.py: exports several methods, each with its own calibration data and quantization (dit_step a16w8, decode a8w8), and writes model_params.h from get_params(). It now reads the ExecuTorch pack version from the csolution's own pack entry, not from whichever version is first in cbuild-pack.yml.
  • ai_layer/model_pte.c (a 14 MB C source) is generated, not committed. A fresh checkout therefore needs the venv and Create AI layer once before its first build; the README says so.
  • model/verify_export.py: host checks (float, fake-quant, TOSA reference model, comparison with the upstream code, PSNR). The host noise equals the firmware's, so a given seed gives the same face on the host, the FVP and the board.

Firmware

  • src/app_main.cpp: the sampler, built on EmbeddedModule, reads the Ethos-U PMU per method.
  • Corstone-320: the model goes to DDR in the AC6, GCC and Clang linker scripts. At the end the app exits the FVP through semihosting SYS_EXIT and writes out/fvp_image.bin for the image check.
  • DevKit-E8: new display and joystick drivers. The MRAM code region is extended to fit the 2.8 MB program, and the pools start above the A32 boot stub in SRAM0.

Tools

  • CMSIS-Toolbox 2.15.0, GCC 15.3.1, LLVM 23.1.0, CMake 4.3.3.
  • PyTorch::ExecuTorch@1.5.1 with executorch==1.5.1 and torch==2.14.0.
  • The Clang -mfpu workaround is removed; 2.15.0 no longer needs it.

Debug-stub task

  • The task now selects the DevKit-E8 part (AE822FA0E5597LS0, rev A0) before app-gen-toc. The Secure Enclave skips a boot table built for another part.
  • It also repairs SETOOLS' stored part/revision first. A tools-config -p without -r from another project (E7, AppKit-E8) otherwise leaves SETOOLS failing every command with Revision is invalid!.

CI

  • One job exports the AI layer.
  • Both targets are built with AC6, GCC and CLANG.
  • The FVP run's image is compared with the host's fake-quant rendering and must reach at least 35 dB PSNR.

Verification

  • AC6, GCC and CLANG builds of both targets, all passing locally.

  • Corstone-320 FVP (AC6): Test_result: PASS, CRC 6b938c66, 48.4 dB PSNR against the host fake-quant image.

  • DevKit-E8 (AC6):

    • the boot demo's image, read back with the debugger, is bit-identical to the FVP's (CRC 6b938c66);
    • 78 ms per guided 4-step face;
    • the display initializes;
    • the joystick's continuous mode ran for several hundred faces.
  • TOSA reference model vs the fake-quant graphs: 0.8% worst relative deviation.

  • CI on this PR: the export and all six builds pass; the FVP runs (AC6, GCC, CLANG) print Test_result: PASS, stop by themselves, and match the host reference at 48.8 dB.

The CRC depends on the exported program. CI's Linux export differs from the macOS export in a few quantization parameters, so its image has CRC 3f8a2b1e instead of 6b938c66. All three toolchains, the FVP and the board give the same CRC for the same program.

Not verified yet

  • The DevKit-E8's UART console and pico-faces' viewer over serial (SW4 was on SEUART during the board run).
  • The Windows (PowerShell) variant of the debug-stub task.

…ch 1.5.1

The hackathon example now runs pico-faces (github.com/cpldcpu/pico-faces,
MIT): a latent diffusion transformer with a convolutional decoder that
generates 128x128 faces. Both methods of the exported program, dit_step and
decode, run entirely on the Ethos-U85; the Cortex-M runs the Euler sampler
and the guidance blend.

Model and export:
- model/model.py re-implements the pico-faces modules with operators the
  Ethos-U delegate supports and describes the program through
  get_methods() and get_params(). Its noise is the firmware's PCG32/CLT-12,
  so a seed gives the same face on the host, the FVP and the board.
- create_ai_layer.py exports several methods, each with its own
  calibration data and quantization (dit_step a16w8, decode a8w8), reports
  the delegates and CPU operators per method, writes model_params.h from
  get_params(), and puts the C array in section .rodata.model. The
  ExecuTorch pack version now comes from the csolution's own pack entry:
  the cbuild-pack.yml lock still selected the old version through the
  previous AI layer. Runtime logs are limited to errors.
- model/verify_export.py: float, fake-quant and TOSA reference-model
  sampling, the upstream comparison, and PSNR between images (also the raw
  frame the FVP writes).
- setup_venv.py downloads the checkpoints from a pinned upstream commit and
  checks their SHA-256. ai_layer/model_pte.c (a 14 MB source) is generated,
  not committed.

Firmware:
- src/app_main.cpp runs the sampler on EmbeddedModule, reads the Ethos-U
  PMU per method and prints timings, a CRC and an ASCII preview. With the
  DevKit-E8 layer it shows every face on the LCD, generates new ones on the
  SW2 joystick and serves pico-faces' serial protocol for its viewer.
- Corstone-320: all three linker scripts put the program in DDR (ROM2);
  the AC6 scatter file gives RAM its own load region with RW_RAM0
  NOCOMPRESS; 32 kB stack; the application ends the simulation through
  semihosting SYS_EXIT and writes out/fvp_image.bin and out/fvp_result.txt.
- DevKit-E8: display (CDC200, MIPI DSI, ILI9806E) and joystick drivers,
  MIPI DPHY power-up, cache invalidation at startup, the HP MRAM region
  extended to 3.4375 MB for the 2.8 MB program, and the pools and frame
  buffer in SRAM0 above the A32 boot stub.

Tools: CMSIS-Toolbox 2.15.0, GCC 15.3.1, LLVM 23.1.0, CMake 4.3.3, and
PyTorch::ExecuTorch 1.5.1 with executorch 1.5.1 and torch 2.14.0. The
Clang -mfpu workaround goes: 2.15.0 derives the FPU from -mcpu.

CI exports the AI layer once, builds both targets with AC6, GCC and CLANG,
runs the FVP, and compares its image with the host's fake-quant rendering
(at least 35 dB PSNR).

Verified: AC6, GCC and CLANG builds of both targets. On the Corstone-320
FVP (AC6) the demo prints Test_result: PASS with CRC 6b938c66, and its
image matches the host fake-quant rendering at 48.4 dB PSNR; the TOSA
reference model agrees with the fake-quant graphs to 0.8 %. On the
DevKit-E8 (AC6) the boot demo's image is bit-identical to the FVP's
(CRC 6b938c66, read back with the debugger), a guided 4-step face takes
78 ms, the display comes up, and the joystick's continuous mode generated
several hundred faces.
app-gen-toc writes the part selected in SETOOLS into the ATOC, and the
Secure Enclave does not boot an ATOC built for another part. The task now
runs tools-config -p 'E8 (AE822FA0E5597LS0) ...' -r A0 first, so a SETOOLS
that another project left on an E7 or on the AppKit-E8 (AE822FA0E5597BS0)
no longer produces stubs the DevKit ignores.

tools-config -p without -r keeps the previous revision. Switching an E8
selection (revision A0) to the E7, whose only revision is B4, that way
leaves a pair that does not exist, and from then on every SETOOLS tool,
tools-config included, stops with "Revision is invalid!". So the task first
writes the part and revision into SETOOLS' utils/global-cfg.db (sed on
Linux and macOS, -replace in PowerShell) and then lets tools-config set
them. The README names the current SETOOLS release (V1.112) and the fix for
SETOOLS commands run by hand.

The stub and its ATOC configuration are the same as in Ensemble 2.2.1.
Checked offline against a copy of SETOOLS V1.110 (bash variant): from an
E7/A0, an AppKit-E8/B4 and an E7/B4 selection the task selects the DevKit
part and app-gen-toc builds the package; app-write-mram was not run. The
PowerShell variant is untested.
The first CI run stopped at cbuild setup: the committed AI layer lists
ai_layer/model_pte.c, which only create_ai_layer.py generates, and
csolution refuses a layer whose files are missing. The export job now
touches an empty model_pte.c before cbuild setup; create_ai_layer.py
overwrites it. The command-line flow in example.md and the MLOps page say
the same for a fresh checkout.
cbuild setup of CMSIS-Toolbox 2.15 configures the CMake project, which
probes the solution's compiler: without the license the AC6 check fails
and cbuild setup exits before the AI layer can be created.
CI's export on Linux gives a program that differs from the macOS export in
a few quantization parameters (the calibration runs in floating point),
and with it the CRC: 3f8a2b1e instead of 6b938c66. AC6, GCC and Clang, the
FVP and the DevKit-E8 all give the same CRC for the same program.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants