Repository navigation
zephyr-cp: release builds lose USB during BLE activity — the UDC thread stack fix in #5 is debug-only #18
Description
Activity
What gdb shows about the "post-scan stall": nothing is stalled — the scan timeout is ignored
Reproduced on
10.3.0-51-g8031983fc9(ELF verified against flash withcompare-sections), then halted the core 9 s into astart_scan(timeout=3)loop and dumped every thread:main thread common_hal_bleio_scanresults_next()atshared-module/_bleio/ScanResults.c:27— the wait loop,done = false, ring bufferused = 0bt_dev.flags0x35= ENABLE | READY | HAS_PUB_KEY | SCANNINGbt_dev.ncmd_semcount = 1;sent_cmd = 0x0— no HCI command outstandingcyw43_bus_mutexowner = 0x0, lock_count = 0— gSPI bus freebt_poll_threadstate 0x04, PC in z_tick_sleep— its normal 4 ms pollall work_queue_main,usbd_thread,udc_rpi_pico_thread_0,tc_rx_handler,airoc_event_task,bt_tx_processorpended (0x02) Neither candidate holds: the CDC TX path is idle, and the HCI driver is idle.
stop_scan()called explicitly takes 5–6 ms.Cause: Zephyr's
start_le_scan_legacy()never readsbt_le_scan_param.timeout— only the extended-scan path passes a duration to the controller — andCONFIG_BT_EXT_ADV=n(required: the CYW43439 has no extended advertising) selects the legacy path. Sotimeout=was silently dropped; atimeout=3scan was still yielding at 57 s / 1600 reports.scanresults_next()returnsNonewhen a Ctrl-C is pending, so the loop exits cleanly and theKeyboardInterruptlands on the next statement — thestop_scan()/print()in the two runs here. The "12 s scan, 925 reports" was ~33 s at the observed ~28 reports/s.Fix: #20 enforces the deadline from the main thread's background task. Measured after:
timeout=3→ 3.0 s,1→ 1.0 s,0.5→ 0.53 s,2→ 2.05 s.Two other things seen while at it (not the stall)
- Any SWD halt kills USB CDC on this board — even a ~0.5 s
gdbattach +bt. Windows keeps a stale COM entry; the fix ismonitor reset run. gdb evidence therefore has to be one scripted batch per reset. udc_rpi_picologs<err> Endpoint 0x83 busyin bursts of >1400 (dropped) lines whenever CDC transmits (the class enqueues while a transfer is in flight — the driver'sXFER_NEWhandler treats that as an error). Harmless with deferred logging; withLOG_MODE_IMMEDIATEit would block the UDC thread for seconds — plausibly the original debug-build "USB loss" amplifier.
🤖 Generated with Claude Code
- Any SWD halt kills USB CDC on this board — even a ~0.5 s
Resolved. Both halves of this issue are now accounted for.
The USB loss is fixed by #19, which moves
CONFIG_UDC_RPI_PICO_STACK_SIZE=2048out ofdebug.confand intoprj.confso release images stop running Zephyr's 512-byte default. Verified on hardware: a sustained active scan completes with CDC and the CIRCUITPY volume both intact, where the same workload previously took the board off the USB bus entirely.The "residual, cause not established" post-scan stall was not a stall. Zephyr's
start_le_scan_legacy()never readsbt_le_scan_param.timeout, andCONFIG_BT_EXT_ADV=n(forced — the CYW43439 has no extended advertising) selects that path, so scans never ended.scanresults_next()returnsNonewhen a Ctrl-C is pending, so the loop exited cleanly and theKeyboardInterruptlanded on whatever statement followed — which is exactly the two tracebacks recorded here. Fixed by #20; the full gdb evidence is in #25.That also corrects the scan figures quoted in this thread: the run described as "12 s" was ~33 s at ~28 reports/s, because it was never being stopped by the timeout.
Closing. The remaining
udc_rpi_picoannoyance found alongside this —<err> Endpoint 0x83 busybursts during normal CDC transmission — is tracked separately for upstreaming.🤖 Generated with Claude Code
Symptom
On a release build, the board disappears from USB partway through a
_bleioscan: the CDC port (COM51) and theCIRCUITPYmass-storage volume both vanish mid-run. On reset Windows reports a USB device/stack error.Recovered with an SWD
reset run— no power cycle needed (unlike the flash-XIP hang in #15).Observed 2026-09-09 on a Pico 2 W,
raspberrypi_rpi_pico2_w_zephyr,10.3.0-51-g8031983fc9, running a 15 sadapter.start_scan().Cause
CONFIG_UDC_RPI_PICO_STACK_SIZEis 512 in release builds — Zephyr's driver default (drivers/usb/udc/Kconfig.rpi_pico).The 2048-byte fix from #5 was added to
debug.confonly — that PR's own title scopes it to debug builds — so every release image still ran the 512-byte stack. The stack overflow we originally diagnosed asusbd@50110000was therefore never fixed for the images we actually flash, which also retroactively explains the "USB came up badly" resets during the original Pico 2 W bring-up.Fix
Move it out of
debug.confand intoprj.confso all builds get it:Cost: +1,536 B RAM (with
CONFIG_BT_MAX_CONN=4alongside it, 231,112 → 236,864 B, 43.40% → 44.48%).Verified on hardware: with the fix, a full 12 s active scan producing 925 advertisement reports from 20 distinct devices completed with USB still enumerated — CDC and the CIRCUITPY volume both intact. The same workload took the board off the bus before.
Residual, separate, cause not established
With USB surviving, a second symptom remains: execution stalls on the first statement after the scan window closes, for minutes, until Ctrl-C.
a.stop_scan()print()immediately after the scan loop; its output never arrivedBoth times Ctrl-C was delivered and honoured, and the REPL was immediately healthy afterwards (
print('ALIVE', 6*7)→ALIVE 42). So the core is not faulted and BLE RX works throughout — only the post-scan transition stalls.Candidates not ruled out: CDC writes blocking when the host stops draining the port, or scan teardown blocking in the HCI driver. Filing separately rather than guessing, since it does not share the stack-size cause and is not fixed by it.
🤖 Generated with Claude Code