Skip to content

fix(verify): make era advisory rule cores-aware - #111

Merged
Seungpyo1007 merged 1 commit into
mainfrom
Seungpyo1007/era-rule-cores-aware
Sep 29, 2026
Merged

Seungpyo1007 merged 1 commit into
mainfrom
Seungpyo1007/era-rule-cores-aware

Conversation

@Seungpyo1007

Copy link
Copy Markdown
Member

Summary

Makes the era-vs-score advisory rule in integrity_check.py cores/threads-aware, fixing a class of false positives surfaced during today's CPU advisory re-check (Refs #98).

Why

The era check flagged an "old chip carrying a modern score" using a flat per-chip ceiling — passmark_cpu_mark > 1500 before 2006 and cinebench_r23_multi > 3000 before 2011 — and ignored core/thread count. But Cinebench R23 and PassMark run on the physical silicon regardless of when the part launched, so legitimate pre-2011 high-core-for-their-era chips clear a flat gate:

  • i7-980X / i7-990X (2010, 6c/12t): R23 ~6,300–6,500
  • i7-970 (2010, 6c/12t): R23 5,800
  • i7-920 / i7-965 EE (2008, 4c/8t): R23 3,800 / 4,500
  • i7-870 (2009, 4c/8t): R23 4,200
  • Phenom II X6 1090T / 1100T (2010, 6c/6t): R23 3,400 / 3,500

Today's advisory re-check found that all 8 era findings were exactly this population — false positives, not data errors.

What changed

  • Scale the ceiling by thread count instead of a flat value: ERA_R23_PER_THREAD = 1000, ERA_PASSMARK_PER_THREAD = 900 (threads default to cores, then 1). Pre-2011 microarchitectures (Nehalem/Westmere/K10) top out around ~600 R23 and well under 900 PassMark per thread, so the known-good chips stop flagging — while a genuinely implausible combo still trips it (e.g. a 2c/2009 part claiming 20,000 R23 = 10,000/thread).
  • Extract the rule into an importable, side-effect-free era_score_outliers(rec) helper and guard the scan body under if __name__ == "__main__":, so the rule can be unit-tested without running the full filesystem scan.
  • Add tests/unit/test_integrity_era.py: the 8 known false positives (must not flag), a genuine implausible case (must flag), the per-thread boundary, the thread→cores→1 fallback, the pre-2006 PassMark path, missing-score no-ops, and an import-has-no-scan check.

Testing

  • ruff check app tests — all checks passed.
  • mypy app — success, no issues in 111 source files.
  • pytest --cov=app --cov-fail-under=60 — 586 passed, total coverage 77.69%.
  • New era tests: 8 passed.
  • Ran integrity_check.py against a live TechAPI develop checkout: the CPU era-vs-score section now reports 0 findings (was 8), with the cross-source ratio and structural sections unchanged (183 ratio lines, 0 hard anomalies — identical to before).

Scope

TechEngine only; no TechAPI data touched. The other advisory tiers (cross-source ratio outliers) are intentionally unchanged — those are the heterogeneous-catalog / benchmark-methodology outliers documented in the #98 re-check, not addressed here.

Refs #98

The era-vs-score check in integrity_check.py used a flat per-chip ceiling
(PassMark>1500 before 2006, R23>3000 before 2011) and ignored core/thread
count. Cinebench R23 and PassMark run on the physical silicon regardless of
launch date, so legitimate pre-2011 high-core enthusiast parts sail past a flat
gate: a 6c/12t Gulftown (i7-980X/990X) posts ~6,000-6,500 R23 today and a 4c/8t
Bloomfield (i7-920) ~3,000-3,800. Today's CPU advisory re-check (Refs #98) found
all 8 era findings were exactly this population of false positives
(i7-920/965/870/970/980X/990X, Phenom II X6 1090T/1100T).

Scale the ceiling by thread count instead (R23 1000/thread, PassMark 900/thread;
threads default to cores, then 1). Pre-2011 microarchitectures top out around
~600 R23 and well under 900 PassMark per thread, so the known-good chips stop
flagging, while a genuinely implausible old-chip/modern-score combo still trips
it (e.g. a 2c/2009 part claiming 20,000 R23 = 10,000/thread).

Extract the rule into an importable, side-effect-free era_score_outliers()
helper and guard the scan under __main__ so it can be unit-tested. Add
tests/unit/test_integrity_era.py covering the 8 known false positives, the
thread-scaling boundary, the thread-count fallback, the pre-2006 PassMark path,
and a genuine implausible case that must still flag.

Verified against a live TechAPI develop checkout: the era-vs-score section now
reports 0 findings (was 8) with the ratio/structural sections unchanged.

Refs #98
@Seungpyo1007
Seungpyo1007 merged commit 2bbcaac into main Sep 29, 2026
1 check passed
Seungpyo1007 added a commit that referenced this pull request Sep 29, 2026
Audit follow-up to #111 (cores-aware era rule): swept every rule in
integrity_check.py for the same shape of bug — a value compared against a
fixed reference that ignores a variable which legitimately shifts what
"normal" looks like — and found one more instance.

The CPU cross-source ratio detector ran a single global median±MAD over the
whole catalog. But ratios like cinebench_r23_multi/geekbench_multi are
confounded by core count: R23 multi scales near-linearly with cores while
Geekbench multicore compresses, so the ratio climbs monotonically with thread
count (measured on live data: ~1.05 at 1-4T up to ~1.48 at 65T+, Pearson
corr(threads, log-ratio) +0.52; PassMark/R23 falls, corr -0.59). A global
median therefore flagged entire legitimate core-count strata — 90 of 739
R23/GB pairs, median 56 threads vs 16 overall, the whole EPYC/Threadripper/
Xeon many-core cluster with genuine scores — as contamination. Same shape as
the flat era ceiling.

Fix mirrors #111 (which divided score by threads): regress the confounder out.
mad_outliers now accepts an optional per-part covariate, fits a robust
Theil-Sen line of log-ratio vs log(covariate), and runs the median±MAD test on
the residuals, so each part is judged against the ratio expected for its own
core count. CPU pairs pass thread count; the systematic core-count gradient no
longer flags while a part anomalous for its own class still does (live R23/GB
flags 90 -> 10, survivors are genuine per-class outliers). Callers without a
covariate (GPUs) are unchanged. Coarse thread-banding was rejected: it removes
the between-band trend but shrinks the within-band envelope, netting more false
positives.

The GPU cross-source ratios were checked and left as a single population on
purpose: they mix a theoretical spec (fp32_tflops) with empirical benchmarks
across gaming vs. compute cards and many hardware eras, with no single clean
stratifying variable, and stay advisory-only.

Adds tests/unit/test_integrity_cross_source.py.

Refs #98
@Seungpyo1007
Seungpyo1007 deleted the Seungpyo1007/era-rule-cores-aware branch September 30, 2026 08:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant