fix(evidence): publish tropical validation with explicit cohort and error scope - #76
Merged
Merged
Conversation
…ed the Belgian figure The accuracy panel said 0.90 cm, from 65 isolated temperate trees. That number is true and it is not the one this product is about. The tropical cohort had been measured three days earlier -- 61 trees scanned, felled and weighed in Cameroon, the only cohort here that reaches the allometric stage at all -- and none of it could reach a reader, because sync_truth.py never emitted the block. render_typescript wrote wanHeldOut, demol65 and pointnetIndependent and stopped; render_truth_block wrote the first two. So core-demo-evidence.ts had no tropical figure to render, and the truth block in three controlled documents was silent about the strongest evidence in the repository. The gate that exists to stop documents drifting from evidence could not help. The evidence never arrived. What the panel now says, from the manifest rather than from prose: temperate 0.90 cm 65 trees, Belgium tropical 1.37 cm 60 measurable, Cameroon -- from the 27 that pass the gate refused 33 of 60 stems the circle fit will not describe, mostly buttressed ceiling 11.25 cm the same stage forced to answer for every tree The refusal count is not a caveat, it is half the headline: an average over the trees that passed says nothing about the trees that did not, and 27 of 60 reads as 60 of 60 without it. The ceiling is there so 1.37 cm cannot be read as the error on an arbitrary tropical tree. Two gaps in the evidence chain closed with it: - validate_cameroon re-hashes docs/evidence/cameroon_61/result.json and compares all 25 published figures against it, and load_manifest now requires the block. It was optional: the cohort could be dropped from the manifest and every gate would still pass while the site reverted to temperate figures. validate_demol has done this for the 65-tree cohort since June, so the newest and strongest evidence had been the least guarded. One field needed a mapping rather than an assumed equality -- the manifest calls the cohort size `trees` and the artefact calls it `cohort_size` -- and without it only 24 of the 25 would have been compared. - test_cameroon_evidence_is_current.py asks the question the hash cannot: does the artefact still reproduce from the trees. It re-derives the evaluation and compares measurements within tolerance, counts exactly. Verified against the real 1.29 GB archive: 14 passed, gate still 27/33. It skips where the archive is absent, like its Demol counterpart -- see WHAT_CI_DOES_NOT_CHECK.md. The NSC 2026 framing is retired from the surfaces that speak in the present tense. The competition was not won; the badge, the hero eyebrow, the page metadata and the footer went on claiming it for weeks because nothing checked. A test checks now, and a second test requires DOCUMENT_STATUS.md to keep the reference, because erasing it from the historical record would be the worse fault. Verified: scripts/tests 168 passed; web 214 passed, tsc clean, next build ok; test_cameroon_evidence_is_current 14 passed against the real archive; ruff clean; sync_truth --check and judge_demo_manifest check both ok. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The landing page and generated truth blocks did not expose the recorded Cameroon cohort alongside the temperate Demol result. This change publishes the existing tropical evidence with its denominator and refusal coverage, and makes the manifest-to-artifact checks mandatory.
libEGL.so.1in the ML runner, thenlibusb-1.0.so.0during the Docker API's PLY request. Installlibegl1andlibusb-1.0-0explicitly, probe Open3D import before the ML suite and during image build, and retain the real API request smoke test. EGL is also listed in the Open3D Docker dependencies; no failing tests are disabled.Validation of the follow-up patch: truth synchronizer tests 92 passed; landing SSR tests 15 passed; TypeScript, ESLint, Ruff, generated truth check and diff whitespace check passed. Full ML tests and a new empirical cohort evaluation were not run locally. The existing raw-cohort reproducibility tests skip when their external archives are unavailable; CI success must not be read as new scientific validation.
PR #75 has been merged and this PR now targets main. Fresh GitHub Actions checks must pass for the final head before merging.