Replace TinyCNN with pico-faces on the hackathon branch - #22
Open
MatthiasHertelArm wants to merge 5 commits into
Open
MatthiasHertelArm wants to merge 5 commits into
MatthiasHertelArm wants to merge 5 commits into
Conversation
…ch 1.5.1 The hackathon example now runs pico-faces (github.com/cpldcpu/pico-faces, MIT): a latent diffusion transformer with a convolutional decoder that generates 128x128 faces. Both methods of the exported program, dit_step and decode, run entirely on the Ethos-U85; the Cortex-M runs the Euler sampler and the guidance blend. Model and export: - model/model.py re-implements the pico-faces modules with operators the Ethos-U delegate supports and describes the program through get_methods() and get_params(). Its noise is the firmware's PCG32/CLT-12, so a seed gives the same face on the host, the FVP and the board. - create_ai_layer.py exports several methods, each with its own calibration data and quantization (dit_step a16w8, decode a8w8), reports the delegates and CPU operators per method, writes model_params.h from get_params(), and puts the C array in section .rodata.model. The ExecuTorch pack version now comes from the csolution's own pack entry: the cbuild-pack.yml lock still selected the old version through the previous AI layer. Runtime logs are limited to errors. - model/verify_export.py: float, fake-quant and TOSA reference-model sampling, the upstream comparison, and PSNR between images (also the raw frame the FVP writes). - setup_venv.py downloads the checkpoints from a pinned upstream commit and checks their SHA-256. ai_layer/model_pte.c (a 14 MB source) is generated, not committed. Firmware: - src/app_main.cpp runs the sampler on EmbeddedModule, reads the Ethos-U PMU per method and prints timings, a CRC and an ASCII preview. With the DevKit-E8 layer it shows every face on the LCD, generates new ones on the SW2 joystick and serves pico-faces' serial protocol for its viewer. - Corstone-320: all three linker scripts put the program in DDR (ROM2); the AC6 scatter file gives RAM its own load region with RW_RAM0 NOCOMPRESS; 32 kB stack; the application ends the simulation through semihosting SYS_EXIT and writes out/fvp_image.bin and out/fvp_result.txt. - DevKit-E8: display (CDC200, MIPI DSI, ILI9806E) and joystick drivers, MIPI DPHY power-up, cache invalidation at startup, the HP MRAM region extended to 3.4375 MB for the 2.8 MB program, and the pools and frame buffer in SRAM0 above the A32 boot stub. Tools: CMSIS-Toolbox 2.15.0, GCC 15.3.1, LLVM 23.1.0, CMake 4.3.3, and PyTorch::ExecuTorch 1.5.1 with executorch 1.5.1 and torch 2.14.0. The Clang -mfpu workaround goes: 2.15.0 derives the FPU from -mcpu. CI exports the AI layer once, builds both targets with AC6, GCC and CLANG, runs the FVP, and compares its image with the host's fake-quant rendering (at least 35 dB PSNR). Verified: AC6, GCC and CLANG builds of both targets. On the Corstone-320 FVP (AC6) the demo prints Test_result: PASS with CRC 6b938c66, and its image matches the host fake-quant rendering at 48.4 dB PSNR; the TOSA reference model agrees with the fake-quant graphs to 0.8 %. On the DevKit-E8 (AC6) the boot demo's image is bit-identical to the FVP's (CRC 6b938c66, read back with the debugger), a guided 4-step face takes 78 ms, the display comes up, and the joystick's continuous mode generated several hundred faces.
app-gen-toc writes the part selected in SETOOLS into the ATOC, and the Secure Enclave does not boot an ATOC built for another part. The task now runs tools-config -p 'E8 (AE822FA0E5597LS0) ...' -r A0 first, so a SETOOLS that another project left on an E7 or on the AppKit-E8 (AE822FA0E5597BS0) no longer produces stubs the DevKit ignores. tools-config -p without -r keeps the previous revision. Switching an E8 selection (revision A0) to the E7, whose only revision is B4, that way leaves a pair that does not exist, and from then on every SETOOLS tool, tools-config included, stops with "Revision is invalid!". So the task first writes the part and revision into SETOOLS' utils/global-cfg.db (sed on Linux and macOS, -replace in PowerShell) and then lets tools-config set them. The README names the current SETOOLS release (V1.112) and the fix for SETOOLS commands run by hand. The stub and its ATOC configuration are the same as in Ensemble 2.2.1. Checked offline against a copy of SETOOLS V1.110 (bash variant): from an E7/A0, an AppKit-E8/B4 and an E7/B4 selection the task selects the DevKit part and app-gen-toc builds the package; app-write-mram was not run. The PowerShell variant is untested.
The first CI run stopped at cbuild setup: the committed AI layer lists ai_layer/model_pte.c, which only create_ai_layer.py generates, and csolution refuses a layer whose files are missing. The export job now touches an empty model_pte.c before cbuild setup; create_ai_layer.py overwrites it. The command-line flow in example.md and the MLOps page say the same for a fresh checkout.
cbuild setup of CMSIS-Toolbox 2.15 configures the CMake project, which probes the solution's compiler: without the license the AC6 check fails and cbuild setup exits before the AI layer can be created.
CI's export on Linux gives a program that differs from the macOS export in a few quantization parameters (the calibration runs in floating point), and with it the CRC: 3f8a2b1e instead of 6b938c66. AC6, GCC and Clang, the FVP and the DevKit-E8 all give the same CRC for the same program.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The hackathon example becomes pico-faces (MIT): a latent diffusion transformer plus a small decoder that generate 128x128 faces. Both methods of the exported program,
dit_stepanddecode, run entirely on the Ethos-U85, and the Cortex-M runs the sampler. On the DevKit-E8 every face appears on the LCD, the SW2 joystick generates new ones, and pico-faces' serial viewer works against the console.The branch also moves to the current tools and adds a fix for the SETOOLS debug-stub task.
Changes
Model and export
model/model.py: the pico-faces modules, re-implemented with operators the Ethos-U delegate supports. They load the upstream checkpoints, whichsetup_venv.pydownloads from a pinned upstream commit and checks by SHA-256.create_ai_layer.py: exports several methods, each with its own calibration data and quantization (dit_stepa16w8,decodea8w8), and writesmodel_params.hfromget_params(). It now reads the ExecuTorch pack version from the csolution's own pack entry, not from whichever version is first incbuild-pack.yml.ai_layer/model_pte.c(a 14 MB C source) is generated, not committed. A fresh checkout therefore needs the venv and Create AI layer once before its first build; the README says so.model/verify_export.py: host checks (float, fake-quant, TOSA reference model, comparison with the upstream code, PSNR). The host noise equals the firmware's, so a given seed gives the same face on the host, the FVP and the board.Firmware
src/app_main.cpp: the sampler, built onEmbeddedModule, reads the Ethos-U PMU per method.SYS_EXITand writesout/fvp_image.binfor the image check.Tools
PyTorch::ExecuTorch@1.5.1withexecutorch==1.5.1andtorch==2.14.0.-mfpuworkaround is removed; 2.15.0 no longer needs it.Debug-stub task
AE822FA0E5597LS0, rev A0) beforeapp-gen-toc. The Secure Enclave skips a boot table built for another part.tools-config -pwithout-rfrom another project (E7, AppKit-E8) otherwise leaves SETOOLS failing every command withRevision is invalid!.CI
Verification
AC6, GCC and CLANG builds of both targets, all passing locally.
Corstone-320 FVP (AC6):
Test_result: PASS, CRC6b938c66, 48.4 dB PSNR against the host fake-quant image.DevKit-E8 (AC6):
6b938c66);TOSA reference model vs the fake-quant graphs: 0.8% worst relative deviation.
CI on this PR: the export and all six builds pass; the FVP runs (AC6, GCC, CLANG) print
Test_result: PASS, stop by themselves, and match the host reference at 48.8 dB.The CRC depends on the exported program. CI's Linux export differs from the macOS export in a few quantization parameters, so its image has CRC
3f8a2b1einstead of6b938c66. All three toolchains, the FVP and the board give the same CRC for the same program.Not verified yet