Skip to content

Make the updater path work end to end, and stop losing the reason a boot failed - #22

Merged
Pummelchen merged 16 commits into
mainfrom
work/ci-coverage-and-fuzzing
Aug 23, 2026
Merged

Pummelchen merged 16 commits into
mainfrom
work/ci-coverage-and-fuzzing

Conversation

@Pummelchen

@Pummelchen Pummelchen commented Aug 23, 2026 •

Copy link
Copy Markdown
Owner

Follow-on to #21. Sixteen commits. It began as CI coverage and console
parity, and grew as each fix exposed the next thing standing between the
updater and a working end-to-end path.

CI coverage and the aggregate gate

qemu-persistence-reboot and parser-fuzz were never in CI. The fuzz runner
also reset its corpus on every invocation, so each campaign relearned the same
shallow coverage from one seed; it now merges and caches. Seed corpora come
from a sustained local campaign (600s per target, three workers, ASan+UBSan):
ssh-packet 1→19, dns 1→202, sftp 1→260. No crashes, leaks or undefined
behaviour in the SSH packet layer, the DNS/DNSSEC response path, or the SFTP
decoder — each reachable before authentication.

Three aggregate-gate defects, none in the code under test: an NVMe launcher
asking for a QEMU machine property only newer releases have, an SMMU gate
asserting a cumulative fault counter was exactly 1, and a marker string that
only exists in the non-test build while the gate boots the test image.

Console parity

The local console and SSH ran different code, so applications behaved
differently depending on how you connected. They now share one session layer:
htop and less render identically, terminal applications launch with the
same options, and the local console gained the pager it never had.

Boot display

The loader took whatever mode firmware left set — 1024x768 on Fusion — and
never enumerated alternatives. It now selects the largest 32-bit mode (capped
2560x1600); QueryMode/SetMode were void * and had to be typed first.
Glyph scale follows display width. The boot screen also shows a public IPv6
address when SLAAC provides one.

The updater path

The updater moved to xaios.91.99.176.243.nip.io, and each layer failed in
turn:

  • TLS. xapt pinned the exact leaf RSA key, which cannot survive a
    publicly issued certificate that is reissued on every renewal, and the edge
    serves ECDSA, which its single cipher suite could not verify at all. It now
    validates the chain against compiled-in ISRG roots with name and validity
    checks. Pinning is kept for a private origin reached by address.
  • The clock. Chain validation checks expiry, and the wall clock came from
    an RTC that reports epoch zero under QEMU. ntp_sync existed but nothing
    called it. Boot now synchronises once, bounded and non-fatal, against a bare
    address so it needs no DNS.
  • Name resolution, twice. net_resolve is authorised by XAIOS_CAP_NET,
    which xapt never held, so no query was ever built and every hostname failed
    identically whether signed or not. Then child_zone_name scanned back for N
    dots to return N labels, which is off by one when the zone is the whole
    name, killing every apex name such as example.com.
  • Insecure delegations. DNSSEC has three outcomes and the resolver had
    two. nip.io is unsigned, and proving an absent DS under the opt-out NSEC3
    its parent serves was not implemented — nor was there an insecure state to
    move into. Both now exist; iteration counts above 150 are refused rather
    than computed, and insecure answers are counted apart from authenticated
    ones.

The shipped configuration fetches over plain HTTP for now. Authenticity does
not depend on it — catalogs are signed and every artifact carries a sha256 —
but transport confidentiality is forfeited until TLS is restored, which now
needs only a configuration change.

Durability and diagnosis

A torn metadata write was unrecoverable. MutableFS rewrote its metadata in
place, so an interrupted write left it neither valid nor blank; mount then
refused to continue, correctly, because formatting would destroy the volume.
It now keeps two copies and alternates writes. The mirror sits past the data
region and the write sequence in existing slack, so no offset moves and older
volumes keep mounting; a volume with no room stays single-copy. With both
copies damaged the mount still refuses. A host test damages each copy the way
a torn write does — and caught a real defect in the first implementation,
which probed the mirror only when the volume reported v5, skipping it exactly
when the primary was torn.

A boot panic never said why. The reason was logged and then erased by the
progress display, and the log ring only started after MutableFS mounted, so an
early panic had nothing to replay. Capture now begins at the top of boot and
the panic screen replays it. Boot-image reads retry instead of ending the boot
on one transient sector error.

SSH no longer strands the machine. It refused to start when the host key
could not be persisted, which is precisely when an operator needs in to repair
storage. It keeps the key and reports it as ephemeral.

Audit

A full pass — mechanical sweeps, a clang-analyzer run over every production C
file, and a re-read of everything this branch changed. Eight defects fixed.
The largest: /bin was invisible to ls because the userspace listing goes
through the VFS and the VFS had never heard of the initramfs, so the shell
could run /bin/htop while the directory looked empty. The image is now
mounted read-only and mount points appear in their parent's listing.

panic.h carried two copies of its body from a bad merge and panic_at was
not noreturn — which mattered beyond codegen, because kassert(p != 0)
only guards the dereference after it if tools know a failed assertion cannot
fall through. One attribute cleared analyzer warnings in fifteen files.

Also: an IPv6 renderer that emitted invalid text whenever the zero run reached
the last group (proven by exhaustive test, 0/256 patterns failing after), two
seqlock readers comparing uninitialised values, a null-pointer inconsistency
at the control-protocol entry, and two bounds issues in the xapt config and
header parsers.

Verification

Hosted tests including the new recovery cases, compile-check, docs-check, the
local console and xapt gates on both architectures, the network suite,
persistence across reboot, and the storage crash test — which reports all
metadata kill points recovered. All three images build and boot to a login
prompt with working password authentication.

André Borchert and others added 16 commits August 23, 2026 16:44
Three separate gaps, all of which let real problems sit unnoticed.

qemu-persistence-reboot was never in CI. It boots twice against one volume and
requires the second boot to reload what the first wrote, and it is the only
thing covering durability across a restart -- which is exactly how a
persistence subsystem writing to a snapshot-backed device went unnoticed. It
now runs on every push.

parser-fuzz was never in CI either, and its runner reset the corpus on every
invocation, so each campaign relearned the same shallow coverage from a single
seed. The corpus is merged rather than reset now, crashes land where CI
collects them, and a bounded campaign runs per push with the corpus cached
between runs.

The seed corpora are replaced with the coverage-unique inputs from a sustained
local campaign: 600 seconds per target across three workers under ASan and
UBSan, then minimised with -merge=1. ssh-packet grew from 1 seed to 19, dns to
202, sftp to 260. The campaign found no crashes, no leaks and no undefined
behaviour in the SSH binary packet layer, the DNS/DNSSEC response path or the
SFTP request decoder -- each of them reachable before authentication.

Core OS Aggregate RC had three defects of its own, none in the code it tests:

  - The NVMe gate asked QEMU for a virt machine "msi" property that only newer
    releases have. On the CI QEMU the aarch64 guest refused to start with
    "Property 'virt-8.2-machine.msi' not found", so the run produced no NVMe
    evidence at all. The launcher now asks the binary what it supports and
    falls back to its=on alone, which routes MSI through the ITS regardless.
  - The SMMU gate required "faults=1". That counter is cumulative across the
    whole SMMU, and unrelated streams raise C_BAD_STE before the test device
    runs, so the total is whatever the boot happened to reach -- it was 9
    locally. The assertion now checks the outcome and that at least one fault
    was recorded, without pinning an exact running total.
  - The aggregate required "userspace DNS resolve/cache path passed", wording
    that only exists in the non-test build. The aggregate boots the
    XAIOS_BOOT_TEST_APPS image, where nettest emits the fixture wording, so
    that marker could never appear. It also expected "AArch64/x86_64" where
    the gate prints the architecture names lowercase.

Verified: the SMMU and NVMe gates pass locally, and the ABI and documentation
contracts pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three things made the guest console look like a DOS box and tell the operator
less than it knows.

The UEFI loader took whatever display mode the firmware happened to leave set,
which on VMware Fusion is 1024x768. It never enumerated the alternatives. The
loader now queries every mode the firmware offers and selects the largest one
in a directly addressable 32-bit format, bounded at 2560x1600 so an unusually
large mode cannot produce a framebuffer the kernel will not map. QueryMode and
SetMode were declared as void pointers and are now typed, since they could not
be called otherwise.

The 8x8 bitmap font was drawn at a fixed 1x2 scale. That reads correctly at
1024x768 and turns into specks at 1920x1200, so selecting a better mode alone
would have made things worse. Glyph scale now follows the display width,
keeping roughly 100-160 columns at any supported resolution.

The boot screen showed only IPv4. The kernel already knew the public IPv6
address -- the old graphical status panel drew it -- but nothing exposed it to
userspace, so the console had no way to print it, and the framebuffer terminal
that replaced that panel inherited the gap. Add a net_local_ipv6 syscall
mirroring net_local_ipv4, and print the address in RFC 5952 form with the
longest zero run collapsed. Nothing is printed when no global unicast address
is configured, so an IPv4-only network still shows an IPv4-only screen: QEMU
user networking offers only a site-local address and correctly prints nothing.

The dual-colour XAI OS brand is unchanged and still renders, on the serial
console through the escape sequences it always used and on the framebuffer
through the terminal's SGR handling.

Verified: boots clean on AArch64, the boot screen renders, and the ABI and
documentation contracts pass with the new syscall recorded in the frozen
contract.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
htop rendered as a flat text dump on the local console and as the familiar
full-screen monitor over SSH. The application was never the difference: the
SSH channel rewrites the command to "--color --interactive --columns N
--rows N" before running it, but that promotion lived inside
prepare_terminal_command, gated on a PTY request, and the local console had no
equivalent. Same binary, different invocation, so the two surfaces disagreed
about how an application should look.

Extract the rules into ssh_terminal_promote_command and call it from both,
rather than keeping a second copy in sshd.c that would drift. The local
console has no window-size protocol, so it passes a conservative 80x24 that
renders correctly on a serial line and inside the framebuffer terminal alike.

The local console now shows the CPU meters, load average, uptime, memory and
swap bars, the aligned PID/CPU%/MEM%/TIME+/RES_KIB columns and the function
key bar, exactly as an SSH session does.

The syscall self-test marker moves from 50 to 51 entries to match the
net_local_ipv6 addition.

Known remaining difference, not addressed here: over SSH the channel keeps an
interactive session alive after the first frame, holding the sort key, filter,
selection and refresh timer in ssh_channel_t and feeding keystrokes back in.
The local console renders one frame and returns to the prompt. Closing that
gap means lifting the interactive session state out of ssh_channel_t into a
surface-independent session both consoles drive, which is a refactor of its
own rather than an option change.

Verified: qemu-smoke, qemu-local-console-gate, the ABI contract and the
documentation contract all pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Making htop launch the same way on both surfaces got the first frame right and
stopped there: over SSH the channel kept a live session, and locally the frame
was printed once and the shell came back. The two surfaces could not behave the
same because only one of them had a session at all -- the state lived in
ssh_channel_t, so the console had nothing to drive.

Lift it out. xaios_htop_session_t holds what the monitor is doing: sort key and
direction, filter text and mode, selection, scroll offsets, refresh interval,
help state and the terminal it was launched for. xaios_htop_sink_t says where
output goes and how a sampling command runs. The logic -- start, build command,
frame pacing, keystroke handling, help, filter prompt, teardown -- now takes
those two and nothing else, so both consoles run the same state machine rather
than two implementations that drift.

The SSH channel embeds the session and supplies a sink over its transport and
connection. The local console holds one and supplies a sink that writes to the
console and samples through its own remote-login session, wired into command
dispatch, input routing and the service tick the way nano and pong already are.

Ending the session is signalled by clearing session->active rather than a
return code, because each surface finishes differently: the channel returns to
its shell or closes, the console reprints its prompt. An earlier version keyed
on the return value and silently skipped the console teardown, leaving no
prompt after quitting.

Two conventions worth recording, both of which cost a debug cycle here:
console_write_bytes reports success as 0 rather than a byte count, and the
page-geometry helpers now take their dimensions from the session instead of
reading the channel's terminal size.

Verified: on the local console htop now refreshes without input, accepts sort
keys with the change persisting across frames, opens and closes help, quits
back to a working shell. Over SSH native_htop_pty_ansi, native_htop_non_pty_plain,
native_htop_shell_restore and native_htop_invalid_option_rejected all still
pass, along with qemu-smoke, qemu-local-console-gate, qemu-keyboard-input-gate,
the ABI contract and the documentation contract.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
nano and pong were already shared: both surfaces drive the same nano_editor_t
and pong_game_t, so they behave identically without anything further. less was
not. The pager existed only on the SSH channel, so "less FILE" on the local
console fell through to the generic command path and dumped the file instead
of paging it -- the same split htop had.

less_pager_t is surface independent already, with open, render, input, resize
and close, so this is wiring rather than a refactor. The console holds a pager,
starts it from the same command form the channel accepts, routes keystrokes to
less_pager_input, redraws from less_pager_render, and restores the screen and
prompt on quit, exactly as it does for nano.

That leaves every interactive application on one implementation across both
consoles: nano and pong through their shared components, htop through the
session extracted previously, and now less.

Verified: on the local console "less /tmp/pager.txt" enters the alternate
screen, renders the file with tilde filler, quits on q and returns a working
shell. qemu-smoke, qemu-local-console-gate and the documentation contract pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both failures came from putting parser-fuzz and the libc contract in front of
the new net_local_ipv6 syscall, and both are worth having found.

The libc contract pins the syscall count so the C99 profile cannot quietly
grow the ABI. Recording net_local_ipv6 moves the budget from 50 to 51, matching
the frozen release-candidate contract. The success line printed the count as
literal text, so it would have kept claiming 50 forever; it now reads the
budget it just checked.

The fuzz stubs declared xaios_net_recv, xaios_net_send and xaios_clock_nanos
with uint64_t where xaios_user.h declares u64. Those agree on macOS, where
uint64_t is unsigned long long, and conflict on Linux, where it is unsigned
long. The stubs had therefore only ever compiled on a development host --
nothing noticed, because parser-fuzz had never run in CI. Adding the job
surfaced it on its first run.

Verified: check-libc-contract, run-parser-fuzz, the ABI contract and the
documentation contract all pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The FreeBSD gate kept failing an SFTP close with fs-close-denied even after
process resources moved off the reusable pid. The layer above was the problem:
vfs.c had no mutual exclusion at all. vfs_open scans the handle table for a
free slot and fills it in separate steps, while vfs_close and
vfs_release_owner clear entries, all reachable from every CPU through the
filesystem syscalls. MutableFS was serialised earlier; the table in front of it
was not.

Each public entry point now takes one lock around its body, renamed *_locked.
The internal cross-calls use the locked forms so a non-recursive lock cannot
re-enter: vfs_read and vfs_write through vfs_pread and vfs_pwrite, and
vfs_delete through vfs_stat, vfs_rmdir and vfs_unlink. Lock order runs one way
only -- a VFS entry point may call a backend that takes the MutableFS lock, and
MutableFS never calls back into the VFS -- so the two cannot invert.

parser-fuzz also could not build on Linux. The DNS target compiles BearSSL,
whose sysrng.c calls getentropy(), which glibc hides under -std=c99 without
_DEFAULT_SOURCE, and openbsd-compat uses __nonstring__, unknown to clang before
21. Both are already handled in scripts/build-image.sh and the hosted test
build; the fuzz build now passes the same two flags. As with the stub type
mismatch, this only ever built on a development host, and adding the job to CI
is what surfaced it.

Verified: qemu-smoke, qemu-local-console-gate, qemu-storage-crash-test,
compile-check, run-parser-fuzz and the documentation contract pass, and the
FreeBSD bidirectional suite passes end to end locally.

Noted but not addressed: one suite run failed earlier with "persistent mount
skipped status=-4" after formatting a fresh volume, and passed on re-run. That
is an intermittent failure in the persistent mount path, separate from this
change, and it now has kernel diagnostics captured when it recurs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Serialising vfs.c made the hosted test-vfs binary fail to link: the kernel
ticket lock's single-CPU fast path calls smp_online_count, and that test links
only the filesystem translation unit. Provide the one symbol it needs, which
reports a single CPU because the hosted test is single threaded.

Verified: test-vfs links and passes, and make hosted-test runs to completion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The updater moved to https://xaios.91.99.176.243.nip.io, served through the
master Caddy edge with a publicly issued certificate. xapt pinned the exact
leaf RSA key, which cannot work against a certificate that is reissued with a
fresh key every renewal, and the edge now presents ECDSA, which the single
offered cipher suite could not verify at all.

xapt now validates the presented chain against compiled-in ISRG roots, checking
the server name and validity window, and offers the ECDSA suite alongside the
RSA one. The roots are generated from a trusted local CA store, never from the
server being validated. Pinning is kept for a private origin reached by
address, where no public chain exists, so the existing gate is unchanged and
still passes on both architectures.

An unset realtime clock is refused outright rather than validated against 1970,
so the failure names itself instead of surfacing as a confusing expiry error.

Also repair two hazards in the publisher, which predate this change but became
dangerous once a shared edge existed: it reloaded the `caddy` service, now the
master fronting unrelated projects, and its --delete-delay rsync would have
removed the updater's own binary and config, which live inside the directory it
serves.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The clock was whatever the RTC reported, and QEMU's PL031 commonly reports
epoch zero, so a booted system sat in 1970. Nothing that checks a certificate
validity window can work from there: xapt's chain validation against the
updater refused to run at all, because validating an expiry against 1970 makes
every certificate look not yet valid.

ntp_sync already existed but was only reachable through an operator control
operation, so nothing set the clock unless someone asked. Boot now runs one
synchronization once the network is up and before any service starts. The
network poll path already dispatches NTP frames and drives the retry and
timeout, so this only starts the exchange and waits for it.

Bounded and non-fatal. The default server is a bare address, so no DNS is
involved, and a network that filters UDP/123 costs a pause of at most six
seconds before boot continues on the RTC reading. An offset this large steps
rather than slews, so the clock is correct immediately rather than converging
over hours.

Verified on QEMU aarch64: synced from the first attempt in 178ms, leaving
`date` reporting source=ntp, and xapt now reaches the TLS handshake instead of
refusing on an unset clock. The xapt gate passes on both architectures.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A booted guest could not resolve any name at all, signed or unsigned. Two
independent defects, both invisible to the gates because every gate configures
an IP literal and so never resolves anything.

The first is a missing capability. net_resolve is authorized by XAIOS_CAP_NET,
which xapt was never granted; it held only XAIOS_CAP_NET_SOCKET. Sockets
therefore worked and name resolution was rejected before a query was ever
built, which is why a literal address reached TLS while every hostname failed
identically whether or not it was DNSSEC-signed.

The second is an off-by-one in the zone walk. child_zone_name returns the last
N labels by scanning back for the Nth dot, but the last N labels are preceded
by N-1 dots whenever the zone is the whole name. Asking for a zone equal to
the hostname therefore always failed, and the guard that would have caught it
tests zone_labels against hostname_labels before the increment that makes them
equal. Every name whose apex is the name being resolved died there, which is
most of them: example.com, cloudflare.com.

Verified on QEMU aarch64 against live servers. cloudflare.com and example.com
now complete the chain, root DNSKEY through DS(com) and DS(example.com) to the
address, and xapt reaches the TLS handshake instead of failing to resolve.

An unsigned delegation is still refused: nip.io has no DS, and proving that
under an opt-out NSEC3 parent is not implemented, so the insecure outcome
remains unavailable. That is the remaining blocker for the updater hostname.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The updater hostname still does not resolve: nip.io is an unsigned delegation
and the resolver has no insecure outcome, so the name is refused as bogus. The
:8090 fallback that field devices used is closed, leaving no address-based path
to the origin.

The edge does serve the origin over plain HTTP on :80, so a client that dials
the address while still naming the origin reaches it today. xapt used the
configured host for both the dial and the Host header and so could not express
that; `address` now overrides the dial target while `host` remains the origin
identity, used for the Host header and for any certificate check. Without
`address` the client resolves `host` exactly as before.

The shipped configuration uses that path with tls=off. Update authenticity is
unaffected: catalogs and manifests are signed and every artifact carries a
sha256, so a tampered payload is rejected whatever the transport. What is given
up is confidentiality and transport integrity, and an observer can see which
artifacts a host fetches. Restoring tls=required with port=443 needs only the
resolver to reach the name.

Verified on QEMU aarch64 against the live origin: `xapt update` fetches and
activates the signed catalog. The xapt gate passes on both architectures.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The active administration configuration lives in persistent storage and a
stored record takes precedence over the compiled default, but the persistent
disk is not recreated by a rebuild. An image rebuilt over a disk written by an
earlier key-only build therefore inherits password=disabled from that disk and
boots with the console locked, whatever the new build was configured for. The
credentials are packaged correctly and the boot log reports the real cause, but
nothing about `make image` suggests the disk is what decided it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A full sweep, mechanical scans plus a clang-analyzer pass over every
production C file plus an adversarial re-read of everything this branch
changed. Fixes, in rough order of consequence:

/bin is now the boot image mounted read-only into the VFS. The shell could
run /bin/htop while ls showed an empty directory, because the userspace ls
lists through the VFS and the VFS had never heard of the initramfs. The
image's flat path table gains derived directory semantics in initramfs.c,
a small read-only adapter exposes it, and vfs_list also names mount points
in their parent's listing, so /bin and /models appear in ls /. The console
builtin cd accepts image-backed directories, and the test-image kernel ls
merges image entries for parity.

panic.h carried two copies of its body from a bad merge, and panic_at was
not declared noreturn. The attribute is load-bearing: kassert(p != 0)
guards the dereference after it only if tools know a failed assertion
cannot fall through — its absence made 15 files' worth of analyzer noise
out of correctly guarded code. Also gained printf format checking, which
found zero mismatched call sites.

The boot screen's IPv6 renderer emitted invalid text whenever the zero run
reached the last group ("2001:db8:1:2:3:4:" — one colon short) and ":::1"
for a leading run. The marker now carries both its colons and the group
after it adds none. Proven by a host transcription: 3 of 7 curated cases
failed before, 0 of 256 exhaustive patterns fail after.

Two seqlock readers in user.c compared an uninitialized sequence value
when a writer was mid-update; the retry happened by luck. The retry is now
forced through a defined value.

control_protocol_dispatch tolerated a null response_bytes at entry while
every handler's success path dereferences it; nulls are now rejected once
at dispatch.

xapt's pinned TLS mode offered the ECDSA suite first, which would steer a
dual-certificate origin into a suite an RSA pin can never satisfy; each
validation mode now offers only what it can verify. Its config parser
could read past the buffer for a 4 KiB config and capped header bytes
against the wrong buffer's size; both bounded.

One dead store removed from the DNS stage machine.

Verified: compile-check, hosted tests, the local console gate and the
xapt gate on both architectures all pass; the analyzer reports zero
warnings across kernel, engine, boot and the SSH stack, with two
documented false positives remaining elsewhere.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
scripts/ holds runtime scripts against an explicit allowlist, and the layout
check enforces it. The trust-anchor generator is build-time tooling and belongs
in tools/ with the other generators, so docs-check failed the moment it landed,
which also failed the aggregate Core OS RC that runs docs-check inside it.

Regenerating from the new location reproduces the anchors byte for byte; only
the provenance comment changes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three areas, all reachable from the same complaint: a system that fails
without saying why, or fails in a way it cannot come back from.

DNSSEC has three outcomes and the resolver implemented two. An unsigned
delegation was refused as bogus, because proving an absent DS under the
opt-out NSEC3 its parent serves was not implemented and the stage machine had
no insecure state to move into: even a successful proof led to a stage that
demanded an RRSIG the zone would never have. NSEC3 denial now exists,
including opt-out, with iteration counts above 150 refused rather than
computed, and a proven-insecure zone accepts its unsigned answer. Those
answers are counted apart from authenticated ones, which carry a stronger
guarantee and should not be reported as the same thing. Verified against the
live origin: the updater hostname resolves through root, io, and an insecure
nip.io, and fetches its signed catalog.

A boot panic printed registers and a backtrace but never said what failed.
The reason was logged, and then erased: a normal boot redraws the progress
display over the serial console moments before the panic replaces it. Worse,
the log ring only started once MutableFS was mounted, because capture and
persistence were the same switch, so an early panic had nothing to replay
even in principle. Capture now begins at the top of boot and depends on
nothing; persistence is enabled separately once storage exists; and the panic
screen replays the tail through a lock-free read, since taking a lock there
could hang instead of printing. Boot-image reads also retry now rather than
ending the boot on one transient sector error, which is safe because every
file is hash-checked and a retry that was needed is logged.

MutableFS rewrote its metadata in place, so a write interrupted by power loss
left the region neither valid nor blank. Mount then refused to continue,
correctly, because formatting would have destroyed the volume; the result was
a filesystem that could be neither mounted nor repaired. It now keeps two
copies and alternates writes, so a tear only ever damages the copy that is
not authoritative. The mirror sits past the data region and the write
sequence in the slack at the end, so no existing offset moves and volumes
written before this keep mounting; a volume with no room stays single-copy.
Alternating rather than writing both keeps the write cost unchanged. With
both copies damaged the mount still refuses, because falling back is a
recovery and not a licence to discard data.

The host test that covers this damaged each copy the way a torn write does
and caught a real defect: the first implementation probed the mirror only
when the volume reported v5, but a torn primary has no readable version, so
the fallback was skipped in exactly the case it exists for.

Also stop refusing SSH startup when the host key cannot be persisted. The key
in hand is good; only its durability is in question, and an unwritable
filesystem is precisely when an operator needs to get in and repair storage.
The service keeps the key and reports it as ephemeral, which clients surface
as a changed key, where an unreachable machine offers nothing to act on.

Verified: hosted tests including the new recovery cases, compile-check,
docs-check, the local console and xapt gates, the network suite, persistence
across reboot, and the storage crash test, which reports all metadata kill
points recovered.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@Pummelchen Pummelchen changed the title Close the CI coverage gaps, repair the aggregate gate, and unify the consoles Make the updater path work end to end, and stop losing the reason a boot failed Aug 23, 2026
@Pummelchen
Pummelchen merged commit 6aa6f33 into main Aug 23, 2026
18 of 19 checks passed
@Pummelchen
Pummelchen deleted the work/ci-coverage-and-fuzzing branch August 25, 2026 14:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant