Make the updater path work end to end, and stop losing the reason a boot failed - #22
Merged
Merged
Conversation
Three separate gaps, all of which let real problems sit unnoticed.
qemu-persistence-reboot was never in CI. It boots twice against one volume and
requires the second boot to reload what the first wrote, and it is the only
thing covering durability across a restart -- which is exactly how a
persistence subsystem writing to a snapshot-backed device went unnoticed. It
now runs on every push.
parser-fuzz was never in CI either, and its runner reset the corpus on every
invocation, so each campaign relearned the same shallow coverage from a single
seed. The corpus is merged rather than reset now, crashes land where CI
collects them, and a bounded campaign runs per push with the corpus cached
between runs.
The seed corpora are replaced with the coverage-unique inputs from a sustained
local campaign: 600 seconds per target across three workers under ASan and
UBSan, then minimised with -merge=1. ssh-packet grew from 1 seed to 19, dns to
202, sftp to 260. The campaign found no crashes, no leaks and no undefined
behaviour in the SSH binary packet layer, the DNS/DNSSEC response path or the
SFTP request decoder -- each of them reachable before authentication.
Core OS Aggregate RC had three defects of its own, none in the code it tests:
- The NVMe gate asked QEMU for a virt machine "msi" property that only newer
releases have. On the CI QEMU the aarch64 guest refused to start with
"Property 'virt-8.2-machine.msi' not found", so the run produced no NVMe
evidence at all. The launcher now asks the binary what it supports and
falls back to its=on alone, which routes MSI through the ITS regardless.
- The SMMU gate required "faults=1". That counter is cumulative across the
whole SMMU, and unrelated streams raise C_BAD_STE before the test device
runs, so the total is whatever the boot happened to reach -- it was 9
locally. The assertion now checks the outcome and that at least one fault
was recorded, without pinning an exact running total.
- The aggregate required "userspace DNS resolve/cache path passed", wording
that only exists in the non-test build. The aggregate boots the
XAIOS_BOOT_TEST_APPS image, where nettest emits the fixture wording, so
that marker could never appear. It also expected "AArch64/x86_64" where
the gate prints the architecture names lowercase.
Verified: the SMMU and NVMe gates pass locally, and the ABI and documentation
contracts pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three things made the guest console look like a DOS box and tell the operator less than it knows. The UEFI loader took whatever display mode the firmware happened to leave set, which on VMware Fusion is 1024x768. It never enumerated the alternatives. The loader now queries every mode the firmware offers and selects the largest one in a directly addressable 32-bit format, bounded at 2560x1600 so an unusually large mode cannot produce a framebuffer the kernel will not map. QueryMode and SetMode were declared as void pointers and are now typed, since they could not be called otherwise. The 8x8 bitmap font was drawn at a fixed 1x2 scale. That reads correctly at 1024x768 and turns into specks at 1920x1200, so selecting a better mode alone would have made things worse. Glyph scale now follows the display width, keeping roughly 100-160 columns at any supported resolution. The boot screen showed only IPv4. The kernel already knew the public IPv6 address -- the old graphical status panel drew it -- but nothing exposed it to userspace, so the console had no way to print it, and the framebuffer terminal that replaced that panel inherited the gap. Add a net_local_ipv6 syscall mirroring net_local_ipv4, and print the address in RFC 5952 form with the longest zero run collapsed. Nothing is printed when no global unicast address is configured, so an IPv4-only network still shows an IPv4-only screen: QEMU user networking offers only a site-local address and correctly prints nothing. The dual-colour XAI OS brand is unchanged and still renders, on the serial console through the escape sequences it always used and on the framebuffer through the terminal's SGR handling. Verified: boots clean on AArch64, the boot screen renders, and the ABI and documentation contracts pass with the new syscall recorded in the frozen contract. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
htop rendered as a flat text dump on the local console and as the familiar full-screen monitor over SSH. The application was never the difference: the SSH channel rewrites the command to "--color --interactive --columns N --rows N" before running it, but that promotion lived inside prepare_terminal_command, gated on a PTY request, and the local console had no equivalent. Same binary, different invocation, so the two surfaces disagreed about how an application should look. Extract the rules into ssh_terminal_promote_command and call it from both, rather than keeping a second copy in sshd.c that would drift. The local console has no window-size protocol, so it passes a conservative 80x24 that renders correctly on a serial line and inside the framebuffer terminal alike. The local console now shows the CPU meters, load average, uptime, memory and swap bars, the aligned PID/CPU%/MEM%/TIME+/RES_KIB columns and the function key bar, exactly as an SSH session does. The syscall self-test marker moves from 50 to 51 entries to match the net_local_ipv6 addition. Known remaining difference, not addressed here: over SSH the channel keeps an interactive session alive after the first frame, holding the sort key, filter, selection and refresh timer in ssh_channel_t and feeding keystrokes back in. The local console renders one frame and returns to the prompt. Closing that gap means lifting the interactive session state out of ssh_channel_t into a surface-independent session both consoles drive, which is a refactor of its own rather than an option change. Verified: qemu-smoke, qemu-local-console-gate, the ABI contract and the documentation contract all pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Making htop launch the same way on both surfaces got the first frame right and stopped there: over SSH the channel kept a live session, and locally the frame was printed once and the shell came back. The two surfaces could not behave the same because only one of them had a session at all -- the state lived in ssh_channel_t, so the console had nothing to drive. Lift it out. xaios_htop_session_t holds what the monitor is doing: sort key and direction, filter text and mode, selection, scroll offsets, refresh interval, help state and the terminal it was launched for. xaios_htop_sink_t says where output goes and how a sampling command runs. The logic -- start, build command, frame pacing, keystroke handling, help, filter prompt, teardown -- now takes those two and nothing else, so both consoles run the same state machine rather than two implementations that drift. The SSH channel embeds the session and supplies a sink over its transport and connection. The local console holds one and supplies a sink that writes to the console and samples through its own remote-login session, wired into command dispatch, input routing and the service tick the way nano and pong already are. Ending the session is signalled by clearing session->active rather than a return code, because each surface finishes differently: the channel returns to its shell or closes, the console reprints its prompt. An earlier version keyed on the return value and silently skipped the console teardown, leaving no prompt after quitting. Two conventions worth recording, both of which cost a debug cycle here: console_write_bytes reports success as 0 rather than a byte count, and the page-geometry helpers now take their dimensions from the session instead of reading the channel's terminal size. Verified: on the local console htop now refreshes without input, accepts sort keys with the change persisting across frames, opens and closes help, quits back to a working shell. Over SSH native_htop_pty_ansi, native_htop_non_pty_plain, native_htop_shell_restore and native_htop_invalid_option_rejected all still pass, along with qemu-smoke, qemu-local-console-gate, qemu-keyboard-input-gate, the ABI contract and the documentation contract. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
nano and pong were already shared: both surfaces drive the same nano_editor_t and pong_game_t, so they behave identically without anything further. less was not. The pager existed only on the SSH channel, so "less FILE" on the local console fell through to the generic command path and dumped the file instead of paging it -- the same split htop had. less_pager_t is surface independent already, with open, render, input, resize and close, so this is wiring rather than a refactor. The console holds a pager, starts it from the same command form the channel accepts, routes keystrokes to less_pager_input, redraws from less_pager_render, and restores the screen and prompt on quit, exactly as it does for nano. That leaves every interactive application on one implementation across both consoles: nano and pong through their shared components, htop through the session extracted previously, and now less. Verified: on the local console "less /tmp/pager.txt" enters the alternate screen, renders the file with tilde filler, quits on q and returns a working shell. qemu-smoke, qemu-local-console-gate and the documentation contract pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both failures came from putting parser-fuzz and the libc contract in front of the new net_local_ipv6 syscall, and both are worth having found. The libc contract pins the syscall count so the C99 profile cannot quietly grow the ABI. Recording net_local_ipv6 moves the budget from 50 to 51, matching the frozen release-candidate contract. The success line printed the count as literal text, so it would have kept claiming 50 forever; it now reads the budget it just checked. The fuzz stubs declared xaios_net_recv, xaios_net_send and xaios_clock_nanos with uint64_t where xaios_user.h declares u64. Those agree on macOS, where uint64_t is unsigned long long, and conflict on Linux, where it is unsigned long. The stubs had therefore only ever compiled on a development host -- nothing noticed, because parser-fuzz had never run in CI. Adding the job surfaced it on its first run. Verified: check-libc-contract, run-parser-fuzz, the ABI contract and the documentation contract all pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The FreeBSD gate kept failing an SFTP close with fs-close-denied even after process resources moved off the reusable pid. The layer above was the problem: vfs.c had no mutual exclusion at all. vfs_open scans the handle table for a free slot and fills it in separate steps, while vfs_close and vfs_release_owner clear entries, all reachable from every CPU through the filesystem syscalls. MutableFS was serialised earlier; the table in front of it was not. Each public entry point now takes one lock around its body, renamed *_locked. The internal cross-calls use the locked forms so a non-recursive lock cannot re-enter: vfs_read and vfs_write through vfs_pread and vfs_pwrite, and vfs_delete through vfs_stat, vfs_rmdir and vfs_unlink. Lock order runs one way only -- a VFS entry point may call a backend that takes the MutableFS lock, and MutableFS never calls back into the VFS -- so the two cannot invert. parser-fuzz also could not build on Linux. The DNS target compiles BearSSL, whose sysrng.c calls getentropy(), which glibc hides under -std=c99 without _DEFAULT_SOURCE, and openbsd-compat uses __nonstring__, unknown to clang before 21. Both are already handled in scripts/build-image.sh and the hosted test build; the fuzz build now passes the same two flags. As with the stub type mismatch, this only ever built on a development host, and adding the job to CI is what surfaced it. Verified: qemu-smoke, qemu-local-console-gate, qemu-storage-crash-test, compile-check, run-parser-fuzz and the documentation contract pass, and the FreeBSD bidirectional suite passes end to end locally. Noted but not addressed: one suite run failed earlier with "persistent mount skipped status=-4" after formatting a fresh volume, and passed on re-run. That is an intermittent failure in the persistent mount path, separate from this change, and it now has kernel diagnostics captured when it recurs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Serialising vfs.c made the hosted test-vfs binary fail to link: the kernel ticket lock's single-CPU fast path calls smp_online_count, and that test links only the filesystem translation unit. Provide the one symbol it needs, which reports a single CPU because the hosted test is single threaded. Verified: test-vfs links and passes, and make hosted-test runs to completion. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The updater moved to https://xaios.91.99.176.243.nip.io, served through the master Caddy edge with a publicly issued certificate. xapt pinned the exact leaf RSA key, which cannot work against a certificate that is reissued with a fresh key every renewal, and the edge now presents ECDSA, which the single offered cipher suite could not verify at all. xapt now validates the presented chain against compiled-in ISRG roots, checking the server name and validity window, and offers the ECDSA suite alongside the RSA one. The roots are generated from a trusted local CA store, never from the server being validated. Pinning is kept for a private origin reached by address, where no public chain exists, so the existing gate is unchanged and still passes on both architectures. An unset realtime clock is refused outright rather than validated against 1970, so the failure names itself instead of surfacing as a confusing expiry error. Also repair two hazards in the publisher, which predate this change but became dangerous once a shared edge existed: it reloaded the `caddy` service, now the master fronting unrelated projects, and its --delete-delay rsync would have removed the updater's own binary and config, which live inside the directory it serves. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The clock was whatever the RTC reported, and QEMU's PL031 commonly reports epoch zero, so a booted system sat in 1970. Nothing that checks a certificate validity window can work from there: xapt's chain validation against the updater refused to run at all, because validating an expiry against 1970 makes every certificate look not yet valid. ntp_sync already existed but was only reachable through an operator control operation, so nothing set the clock unless someone asked. Boot now runs one synchronization once the network is up and before any service starts. The network poll path already dispatches NTP frames and drives the retry and timeout, so this only starts the exchange and waits for it. Bounded and non-fatal. The default server is a bare address, so no DNS is involved, and a network that filters UDP/123 costs a pause of at most six seconds before boot continues on the RTC reading. An offset this large steps rather than slews, so the clock is correct immediately rather than converging over hours. Verified on QEMU aarch64: synced from the first attempt in 178ms, leaving `date` reporting source=ntp, and xapt now reaches the TLS handshake instead of refusing on an unset clock. The xapt gate passes on both architectures. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A booted guest could not resolve any name at all, signed or unsigned. Two independent defects, both invisible to the gates because every gate configures an IP literal and so never resolves anything. The first is a missing capability. net_resolve is authorized by XAIOS_CAP_NET, which xapt was never granted; it held only XAIOS_CAP_NET_SOCKET. Sockets therefore worked and name resolution was rejected before a query was ever built, which is why a literal address reached TLS while every hostname failed identically whether or not it was DNSSEC-signed. The second is an off-by-one in the zone walk. child_zone_name returns the last N labels by scanning back for the Nth dot, but the last N labels are preceded by N-1 dots whenever the zone is the whole name. Asking for a zone equal to the hostname therefore always failed, and the guard that would have caught it tests zone_labels against hostname_labels before the increment that makes them equal. Every name whose apex is the name being resolved died there, which is most of them: example.com, cloudflare.com. Verified on QEMU aarch64 against live servers. cloudflare.com and example.com now complete the chain, root DNSKEY through DS(com) and DS(example.com) to the address, and xapt reaches the TLS handshake instead of failing to resolve. An unsigned delegation is still refused: nip.io has no DS, and proving that under an opt-out NSEC3 parent is not implemented, so the insecure outcome remains unavailable. That is the remaining blocker for the updater hostname. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The updater hostname still does not resolve: nip.io is an unsigned delegation and the resolver has no insecure outcome, so the name is refused as bogus. The :8090 fallback that field devices used is closed, leaving no address-based path to the origin. The edge does serve the origin over plain HTTP on :80, so a client that dials the address while still naming the origin reaches it today. xapt used the configured host for both the dial and the Host header and so could not express that; `address` now overrides the dial target while `host` remains the origin identity, used for the Host header and for any certificate check. Without `address` the client resolves `host` exactly as before. The shipped configuration uses that path with tls=off. Update authenticity is unaffected: catalogs and manifests are signed and every artifact carries a sha256, so a tampered payload is rejected whatever the transport. What is given up is confidentiality and transport integrity, and an observer can see which artifacts a host fetches. Restoring tls=required with port=443 needs only the resolver to reach the name. Verified on QEMU aarch64 against the live origin: `xapt update` fetches and activates the signed catalog. The xapt gate passes on both architectures. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The active administration configuration lives in persistent storage and a stored record takes precedence over the compiled default, but the persistent disk is not recreated by a rebuild. An image rebuilt over a disk written by an earlier key-only build therefore inherits password=disabled from that disk and boots with the console locked, whatever the new build was configured for. The credentials are packaged correctly and the boot log reports the real cause, but nothing about `make image` suggests the disk is what decided it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A full sweep, mechanical scans plus a clang-analyzer pass over every
production C file plus an adversarial re-read of everything this branch
changed. Fixes, in rough order of consequence:
/bin is now the boot image mounted read-only into the VFS. The shell could
run /bin/htop while ls showed an empty directory, because the userspace ls
lists through the VFS and the VFS had never heard of the initramfs. The
image's flat path table gains derived directory semantics in initramfs.c,
a small read-only adapter exposes it, and vfs_list also names mount points
in their parent's listing, so /bin and /models appear in ls /. The console
builtin cd accepts image-backed directories, and the test-image kernel ls
merges image entries for parity.
panic.h carried two copies of its body from a bad merge, and panic_at was
not declared noreturn. The attribute is load-bearing: kassert(p != 0)
guards the dereference after it only if tools know a failed assertion
cannot fall through — its absence made 15 files' worth of analyzer noise
out of correctly guarded code. Also gained printf format checking, which
found zero mismatched call sites.
The boot screen's IPv6 renderer emitted invalid text whenever the zero run
reached the last group ("2001:db8:1:2:3:4:" — one colon short) and ":::1"
for a leading run. The marker now carries both its colons and the group
after it adds none. Proven by a host transcription: 3 of 7 curated cases
failed before, 0 of 256 exhaustive patterns fail after.
Two seqlock readers in user.c compared an uninitialized sequence value
when a writer was mid-update; the retry happened by luck. The retry is now
forced through a defined value.
control_protocol_dispatch tolerated a null response_bytes at entry while
every handler's success path dereferences it; nulls are now rejected once
at dispatch.
xapt's pinned TLS mode offered the ECDSA suite first, which would steer a
dual-certificate origin into a suite an RSA pin can never satisfy; each
validation mode now offers only what it can verify. Its config parser
could read past the buffer for a 4 KiB config and capped header bytes
against the wrong buffer's size; both bounded.
One dead store removed from the DNS stage machine.
Verified: compile-check, hosted tests, the local console gate and the
xapt gate on both architectures all pass; the analyzer reports zero
warnings across kernel, engine, boot and the SSH stack, with two
documented false positives remaining elsewhere.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
scripts/ holds runtime scripts against an explicit allowlist, and the layout check enforces it. The trust-anchor generator is build-time tooling and belongs in tools/ with the other generators, so docs-check failed the moment it landed, which also failed the aggregate Core OS RC that runs docs-check inside it. Regenerating from the new location reproduces the anchors byte for byte; only the provenance comment changes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three areas, all reachable from the same complaint: a system that fails without saying why, or fails in a way it cannot come back from. DNSSEC has three outcomes and the resolver implemented two. An unsigned delegation was refused as bogus, because proving an absent DS under the opt-out NSEC3 its parent serves was not implemented and the stage machine had no insecure state to move into: even a successful proof led to a stage that demanded an RRSIG the zone would never have. NSEC3 denial now exists, including opt-out, with iteration counts above 150 refused rather than computed, and a proven-insecure zone accepts its unsigned answer. Those answers are counted apart from authenticated ones, which carry a stronger guarantee and should not be reported as the same thing. Verified against the live origin: the updater hostname resolves through root, io, and an insecure nip.io, and fetches its signed catalog. A boot panic printed registers and a backtrace but never said what failed. The reason was logged, and then erased: a normal boot redraws the progress display over the serial console moments before the panic replaces it. Worse, the log ring only started once MutableFS was mounted, because capture and persistence were the same switch, so an early panic had nothing to replay even in principle. Capture now begins at the top of boot and depends on nothing; persistence is enabled separately once storage exists; and the panic screen replays the tail through a lock-free read, since taking a lock there could hang instead of printing. Boot-image reads also retry now rather than ending the boot on one transient sector error, which is safe because every file is hash-checked and a retry that was needed is logged. MutableFS rewrote its metadata in place, so a write interrupted by power loss left the region neither valid nor blank. Mount then refused to continue, correctly, because formatting would have destroyed the volume; the result was a filesystem that could be neither mounted nor repaired. It now keeps two copies and alternates writes, so a tear only ever damages the copy that is not authoritative. The mirror sits past the data region and the write sequence in the slack at the end, so no existing offset moves and volumes written before this keep mounting; a volume with no room stays single-copy. Alternating rather than writing both keeps the write cost unchanged. With both copies damaged the mount still refuses, because falling back is a recovery and not a licence to discard data. The host test that covers this damaged each copy the way a torn write does and caught a real defect: the first implementation probed the mirror only when the volume reported v5, but a torn primary has no readable version, so the fallback was skipped in exactly the case it exists for. Also stop refusing SSH startup when the host key cannot be persisted. The key in hand is good; only its durability is in question, and an unwritable filesystem is precisely when an operator needs to get in and repair storage. The service keeps the key and reports it as ephemeral, which clients surface as a changed key, where an unreachable machine offers nothing to act on. Verified: hosted tests including the new recovery cases, compile-check, docs-check, the local console and xapt gates, the network suite, persistence across reboot, and the storage crash test, which reports all metadata kill points recovered. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-on to #21. Sixteen commits. It began as CI coverage and console
parity, and grew as each fix exposed the next thing standing between the
updater and a working end-to-end path.
CI coverage and the aggregate gate
qemu-persistence-rebootandparser-fuzzwere never in CI. The fuzz runneralso reset its corpus on every invocation, so each campaign relearned the same
shallow coverage from one seed; it now merges and caches. Seed corpora come
from a sustained local campaign (600s per target, three workers, ASan+UBSan):
ssh-packet 1→19, dns 1→202, sftp 1→260. No crashes, leaks or undefined
behaviour in the SSH packet layer, the DNS/DNSSEC response path, or the SFTP
decoder — each reachable before authentication.
Three aggregate-gate defects, none in the code under test: an NVMe launcher
asking for a QEMU machine property only newer releases have, an SMMU gate
asserting a cumulative fault counter was exactly 1, and a marker string that
only exists in the non-test build while the gate boots the test image.
Console parity
The local console and SSH ran different code, so applications behaved
differently depending on how you connected. They now share one session layer:
htopandlessrender identically, terminal applications launch with thesame options, and the local console gained the pager it never had.
Boot display
The loader took whatever mode firmware left set — 1024x768 on Fusion — and
never enumerated alternatives. It now selects the largest 32-bit mode (capped
2560x1600);
QueryMode/SetModewerevoid *and had to be typed first.Glyph scale follows display width. The boot screen also shows a public IPv6
address when SLAAC provides one.
The updater path
The updater moved to
xaios.91.99.176.243.nip.io, and each layer failed inturn:
publicly issued certificate that is reissued on every renewal, and the edge
serves ECDSA, which its single cipher suite could not verify at all. It now
validates the chain against compiled-in ISRG roots with name and validity
checks. Pinning is kept for a private origin reached by address.
an RTC that reports epoch zero under QEMU.
ntp_syncexisted but nothingcalled it. Boot now synchronises once, bounded and non-fatal, against a bare
address so it needs no DNS.
net_resolveis authorised byXAIOS_CAP_NET,which xapt never held, so no query was ever built and every hostname failed
identically whether signed or not. Then
child_zone_namescanned back for Ndots to return N labels, which is off by one when the zone is the whole
name, killing every apex name such as
example.com.two.
nip.iois unsigned, and proving an absent DS under the opt-out NSEC3its parent serves was not implemented — nor was there an insecure state to
move into. Both now exist; iteration counts above 150 are refused rather
than computed, and insecure answers are counted apart from authenticated
ones.
The shipped configuration fetches over plain HTTP for now. Authenticity does
not depend on it — catalogs are signed and every artifact carries a sha256 —
but transport confidentiality is forfeited until TLS is restored, which now
needs only a configuration change.
Durability and diagnosis
A torn metadata write was unrecoverable. MutableFS rewrote its metadata in
place, so an interrupted write left it neither valid nor blank; mount then
refused to continue, correctly, because formatting would destroy the volume.
It now keeps two copies and alternates writes. The mirror sits past the data
region and the write sequence in existing slack, so no offset moves and older
volumes keep mounting; a volume with no room stays single-copy. With both
copies damaged the mount still refuses. A host test damages each copy the way
a torn write does — and caught a real defect in the first implementation,
which probed the mirror only when the volume reported v5, skipping it exactly
when the primary was torn.
A boot panic never said why. The reason was logged and then erased by the
progress display, and the log ring only started after MutableFS mounted, so an
early panic had nothing to replay. Capture now begins at the top of boot and
the panic screen replays it. Boot-image reads retry instead of ending the boot
on one transient sector error.
SSH no longer strands the machine. It refused to start when the host key
could not be persisted, which is precisely when an operator needs in to repair
storage. It keeps the key and reports it as ephemeral.
Audit
A full pass — mechanical sweeps, a clang-analyzer run over every production C
file, and a re-read of everything this branch changed. Eight defects fixed.
The largest:
/binwas invisible tolsbecause the userspace listing goesthrough the VFS and the VFS had never heard of the initramfs, so the shell
could run
/bin/htopwhile the directory looked empty. The image is nowmounted read-only and mount points appear in their parent's listing.
panic.hcarried two copies of its body from a bad merge andpanic_atwasnot
noreturn— which mattered beyond codegen, becausekassert(p != 0)only guards the dereference after it if tools know a failed assertion cannot
fall through. One attribute cleared analyzer warnings in fifteen files.
Also: an IPv6 renderer that emitted invalid text whenever the zero run reached
the last group (proven by exhaustive test, 0/256 patterns failing after), two
seqlock readers comparing uninitialised values, a null-pointer inconsistency
at the control-protocol entry, and two bounds issues in the xapt config and
header parsers.
Verification
Hosted tests including the new recovery cases, compile-check, docs-check, the
local console and xapt gates on both architectures, the network suite,
persistence across reboot, and the storage crash test — which reports all
metadata kill points recovered. All three images build and boot to a login
prompt with working password authentication.