Skip to content

Commit c313e00

Browse files
committed
fix(verify): make cross-source ratio outliers cores-aware
Audit follow-up to #111 (cores-aware era rule): swept every rule in integrity_check.py for the same shape of bug — a value compared against a fixed reference that ignores a variable which legitimately shifts what "normal" looks like — and found one more instance. The CPU cross-source ratio detector ran a single global median±MAD over the whole catalog. But ratios like cinebench_r23_multi/geekbench_multi are confounded by core count: R23 multi scales near-linearly with cores while Geekbench multicore compresses, so the ratio climbs monotonically with thread count (measured on live data: ~1.05 at 1-4T up to ~1.48 at 65T+, Pearson corr(threads, log-ratio) +0.52; PassMark/R23 falls, corr -0.59). A global median therefore flagged entire legitimate core-count strata — 90 of 739 R23/GB pairs, median 56 threads vs 16 overall, the whole EPYC/Threadripper/ Xeon many-core cluster with genuine scores — as contamination. Same shape as the flat era ceiling. Fix mirrors #111 (which divided score by threads): regress the confounder out. mad_outliers now accepts an optional per-part covariate, fits a robust Theil-Sen line of log-ratio vs log(covariate), and runs the median±MAD test on the residuals, so each part is judged against the ratio expected for its own core count. CPU pairs pass thread count; the systematic core-count gradient no longer flags while a part anomalous for its own class still does (live R23/GB flags 90 -> 10, survivors are genuine per-class outliers). Callers without a covariate (GPUs) are unchanged. Coarse thread-banding was rejected: it removes the between-band trend but shrinks the within-band envelope, netting more false positives. The GPU cross-source ratios were checked and left as a single population on purpose: they mix a theoretical spec (fp32_tflops) with empirical benchmarks across gaming vs. compute cards and many hardware eras, with no single clean stratifying variable, and stay advisory-only. Adds tests/unit/test_integrity_cross_source.py. Refs #98
1 parent de153f0 commit c313e00

2 files changed

Lines changed: 212 additions & 9 deletions

File tree

‎integrity_check.py‎

Lines changed: 89 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -15,7 +15,10 @@
1515
slug/filename mismatches, and physically-impossible single>multi benchmarks.
1616
The statistical cross-source/era outliers stay advisory (a heterogeneous catalog
1717
of server + desktop + mobile parts legitimately produces many ratio outliers), so
18-
they are printed for review but never fail the gate.
18+
they are printed for review but never fail the gate. The CPU cross-source ratio
19+
check regresses out the core-count trend so it compares each part against the
20+
ratio expected for its own core count instead of a desktop-dominated global
21+
median (see mad_outliers).
1922
``--hard-report`` writes a JSON list of hard anomalies for baseline comparison.
2023
"""
2124
from __future__ import annotations
@@ -91,20 +94,88 @@ def load(comp):
9194
recs.append((p, fn[:-5], json.load(open(p, encoding="utf-8"))))
9295
return recs
9396

94-
def mad_outliers(pairs, lo=0.34, hi=3.0):
95-
"""pairs: list of (label, a, b); flag log(a/b) outliers via median±3*MAD."""
96-
rs = [(l, math.log(a / b)) for l, a, b in pairs if a and b]
97-
if len(rs) < 8:
97+
# Cross-source ratio outliers: same wrong-variant detector as the era rule, but
98+
# statistical. The original computed ONE global median±MAD over the whole CPU
99+
# catalog for ratios like cinebench_r23_multi/geekbench_multi. That ratio is not
100+
# scale-free across the catalog — it is confounded by core/thread count, because
101+
# the two benchmarks scale differently with parallelism: Cinebench R23 multi
102+
# scales near-linearly with cores while Geekbench multicore compresses at high
103+
# core counts, so the R23/GB ratio climbs monotonically with thread count
104+
# (measured on live data: ~1.05 at 1-4T rising to ~1.48 at 65T+, Pearson
105+
# corr(threads, log-ratio) ≈ +0.52; the PassMark/R23 ratio falls with threads,
106+
# corr ≈ -0.59). A single global median therefore flags entire legitimate
107+
# core-count strata — every many-core EPYC/Threadripper/Xeon and, at the other
108+
# end, the low-core parts — as "contamination" (90 of 739 R23/GB pairs, median
109+
# 56 threads vs 16 overall, all with genuine scores). That is the same shape of
110+
# bug as the flat era ceiling: a fixed reference blind to a variable that
111+
# legitimately shifts what "normal" looks like.
112+
#
113+
# Fix (mirrors PR #111's cores-aware era rule, which divided the score by thread
114+
# count): regress the confounder out. When a per-part covariate is supplied we
115+
# fit a robust (Theil–Sen) line of log-ratio against log(covariate) and run the
116+
# median±MAD test on the *residuals*, so a part is measured against the ratio
117+
# expected for its own core count rather than a desktop-dominated global median.
118+
# The systematic core-count gradient no longer flags; a part whose ratio is
119+
# anomalous for its own class — the real wrong-variant signal — still does
120+
# (live R23/GB flags drop 90 → 10, and the survivors are genuine per-class
121+
# outliers). Coarse banding was rejected: it removes the between-band trend but
122+
# shrinks the within-band MAD envelope, netting *more* false positives.
123+
# Callers with no meaningful covariate (e.g. GPUs) omit it and get the original
124+
# single-population behaviour unchanged.
125+
def _theil_sen(xs: list[float], ys: list[float]) -> tuple[float, float]:
126+
"""Robust slope/intercept via median of pairwise slopes (Theil–Sen)."""
127+
slopes = [
128+
(ys[j] - ys[i]) / (xs[j] - xs[i])
129+
for i in range(len(xs)) for j in range(i + 1, len(xs))
130+
if xs[j] != xs[i]
131+
]
132+
slope = statistics.median(slopes) if slopes else 0.0
133+
intercept = statistics.median(y - slope * x for x, y in zip(xs, ys, strict=True))
134+
return slope, intercept
135+
136+
def mad_outliers(pairs):
137+
"""Flag log(a/b) outliers via median±4*MAD.
138+
139+
``pairs`` is a list of ``(label, a, b)`` or ``(label, a, b, covariate)``.
140+
With a positive per-part ``covariate`` (e.g. thread count) the log-ratio is
141+
first detrended against ``log(covariate)`` with a robust Theil–Sen fit and
142+
the outlier test runs on the residuals, so a variable that legitimately
143+
shifts the ratio does not turn a whole stratum into false positives. Fewer
144+
than 8 usable points returns nothing — too few to estimate a robust
145+
median/MAD. Omitting the covariate reproduces the original global test.
146+
"""
147+
rows: list[tuple[str, float, float | None]] = []
148+
for item in pairs:
149+
label, a, b = item[0], item[1], item[2]
150+
cov = item[3] if len(item) > 3 else None
151+
if a and b:
152+
rows.append((label, math.log(a / b), cov))
153+
if len(rows) < 8:
98154
return []
99-
med = statistics.median(r for _, r in rs)
100-
mad = statistics.median(abs(r - med) for _, r in rs) or 1e-9
101-
return [(l, round(math.exp(r), 2)) for l, r in rs if abs(r - med) > 4 * mad]
155+
ys = [r for _, r, _ in rows]
156+
xs = [math.log(c) for _, _, c in rows if c and c > 0]
157+
if len(xs) == len(rows) and len({round(x, 9) for x in xs}) > 1:
158+
slope, intercept = _theil_sen(xs, ys)
159+
scores = [y - (slope * x + intercept) for x, y in zip(xs, ys, strict=True)]
160+
else:
161+
scores = ys # no usable covariate -> original global behaviour
162+
med = statistics.median(scores)
163+
mad = statistics.median(abs(s - med) for s in scores) or 1e-9
164+
return [
165+
(rows[i][0], round(math.exp(ys[i]), 2))
166+
for i in range(len(rows)) if abs(scores[i] - med) > 4 * mad
167+
]
102168

103169
def section(t): print(f"\n### {t}")
104170

105171
def collect(recs, fa, fb):
106172
return [(d["name"], d[fa], d[fb]) for p, fn, d in recs if d.get(fa) and d.get(fb)]
107173

174+
def collect_cpu(recs, fa, fb):
175+
"""Like ``collect`` but tags each pair with its thread-count covariate."""
176+
return [(d["name"], d[fa], d[fb], d.get("threads") or d.get("cores") or 1)
177+
for p, fn, d in recs if d.get(fa) and d.get(fb)]
178+
108179
def main() -> None:
109180
records = {category: load(category) for category in CATEGORIES}
110181
cpus = records["cpu"]; gpus = records["gpu"]
@@ -157,16 +228,25 @@ def main() -> None:
157228
print(msg)
158229

159230
# --- 5. cross-source correlation outliers (KEY contamination detector) ---
231+
# Thread-count-aware (see mad_outliers): the ratio between two CPU benchmarks
232+
# is confounded by core count, so the core-count trend is regressed out and
233+
# each part is judged against the ratio expected for its own core count
234+
# rather than a desktop-dominated global median.
160235
section("CPU cross-source ratio outliers (possible wrong-variant)")
161236
for fa, fb in [("passmark_cpu_mark","cinebench_r23_multi"),
162237
("passmark_cpu_mark","geekbench_multi"),
163238
("cinebench_r23_multi","geekbench_multi"),
164239
("cinebench_2024_multi","cinebench_r23_multi")]:
165-
out = mad_outliers(collect(cpus, fa, fb))
240+
out = mad_outliers(collect_cpu(cpus, fa, fb))
166241
for label, ratio in out:
167242
print(f" [{fa}/{fb}] {label!r}: ratio={ratio}")
168243

169244
# --- 6. GPU cross-source + sanity ---
245+
# Left as a single population on purpose: unlike the CPU thread-count
246+
# confounder, the GPU ratios mix a theoretical spec (fp32_tflops) with
247+
# empirical benchmarks across gaming vs. compute cards and many hardware
248+
# eras, so there is no single clean stratifying variable. These stay
249+
# advisory-only and are surfaced for human review rather than gated.
170250
section("GPU cross-source ratio outliers + sanity")
171251
for fa, fb in [("passmark_g3d_mark","timespy_score"),
172252
("timespy_score","blender_score"),
Lines changed: 123 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,123 @@
1+
"""Unit tests for the core-count-aware cross-source ratio outlier detector in
2+
``integrity_check``.
3+
4+
Cross-source ratios (e.g. cinebench_r23_multi / geekbench_multi) are the key
5+
wrong-variant contamination signal, but the raw ratio is confounded by core
6+
count: the two benchmarks scale differently with parallelism, so the ratio drifts
7+
monotonically with thread count. The original single global median±MAD therefore
8+
flagged entire legitimate core-count strata (every many-core EPYC/Threadripper/
9+
Xeon) as outliers — the same shape of bug as the flat era ceiling that PR #111
10+
fixed. The fix regresses the core-count trend out and runs the outlier test on
11+
the residuals. These tests pin that behaviour: a whole core-count band that only
12+
follows the natural trend must not flag, while a part that is anomalous *for its
13+
own core count* still must, and callers without a covariate keep the original
14+
global behaviour.
15+
"""
16+
17+
from __future__ import annotations
18+
19+
import math
20+
21+
import integrity_check as ic
22+
23+
24+
def _clean_pair_at(
25+
threads: int, trend_slope: float = 0.11, base: float = 0.05
26+
) -> tuple[float, float]:
27+
"""Build an (a, b) pair whose log-ratio sits exactly on the core-count trend.
28+
29+
log(a/b) = base + trend_slope * log(threads); b fixed at 1000.
30+
"""
31+
log_ratio = base + trend_slope * math.log(threads)
32+
b = 1000.0
33+
a = b * math.exp(log_ratio)
34+
return a, b
35+
36+
37+
# A realistic thread-count ladder spanning desktop -> HEDT -> server, each part's
38+
# ratio lying on the same gentle upward core-count trend (the confounder). The
39+
# population is desktop-dominated (like the real catalog: median ~16 threads), so
40+
# a global median is pulled toward the low-core parts and the legitimate high-core
41+
# trend-followers look like outliers to it.
42+
TREND_THREADS = (
43+
[4] * 6 + [8] * 8 + [12] * 10 + [16] * 12 + [24] * 6 + [32] * 4
44+
+ [48, 64, 96, 128, 192, 256]
45+
)
46+
47+
48+
def test_pairs_on_the_core_count_trend_do_not_flag() -> None:
49+
# Every pair follows the same log-ratio-vs-log-threads line, so after the
50+
# trend is regressed out the residuals are ~0 and nothing is an outlier --
51+
# even though the raw ratios span a wide range (the old global test flagged
52+
# the extremes).
53+
pairs = [
54+
(f"part-{t}-{i}", *_clean_pair_at(t), t)
55+
for i, t in enumerate(TREND_THREADS)
56+
]
57+
assert ic.mad_outliers(pairs) == []
58+
59+
60+
def test_raw_global_test_would_have_flagged_the_trend_extremes() -> None:
61+
# Guard/contrast: the SAME data, fed WITHOUT the covariate, reproduces the
62+
# old behaviour and flags the high-core extremes. This documents that the fix
63+
# -- not a change in the data -- is what removes the false positives.
64+
pairs_no_cov = [(f"part-{t}-{i}", *_clean_pair_at(t)) for i, t in enumerate(TREND_THREADS)]
65+
flagged = [label for label, _ in ic.mad_outliers(pairs_no_cov)]
66+
# The many-core trend-followers (which are perfectly legitimate) are what the
67+
# covariate-free test wrongly flags.
68+
assert flagged, "global (covariate-free) test should still flag trend extremes"
69+
assert any("part-256" in f or "part-192" in f for f in flagged)
70+
71+
72+
def test_part_anomalous_for_its_own_core_count_still_flags() -> None:
73+
# A part whose ratio is far off the trend line *for its thread count* is a
74+
# genuine wrong-variant candidate and must survive the detrending.
75+
pairs = [
76+
(f"part-{t}-{i}", *_clean_pair_at(t), t)
77+
for i, t in enumerate(TREND_THREADS)
78+
]
79+
# Inject a 32-thread part with a wildly wrong ratio (e.g. a swapped variant):
80+
bogus_a, bogus_b = _clean_pair_at(32)
81+
pairs.append(("Swapped-variant 32T", bogus_a * 4.0, bogus_b, 32))
82+
flagged = [label for label, _ in ic.mad_outliers(pairs)]
83+
assert "Swapped-variant 32T" in flagged
84+
85+
86+
def test_missing_covariate_recovers_global_behaviour() -> None:
87+
# Three-tuples (no covariate) must behave exactly like the original detector:
88+
# a uniform cluster with one gross outlier flags only the outlier.
89+
pairs = [(f"p{i}", 1000.0 + i, 1000.0) for i in range(20)]
90+
pairs.append(("gross", 100000.0, 1000.0))
91+
flagged = [label for label, _ in ic.mad_outliers(pairs)]
92+
assert flagged == ["gross"]
93+
94+
95+
def test_fewer_than_eight_points_never_flags() -> None:
96+
pairs = [(f"p{i}", 1000.0 * (i + 1), 1000.0, 8) for i in range(7)]
97+
assert ic.mad_outliers(pairs) == []
98+
99+
100+
def test_zero_or_missing_values_are_skipped() -> None:
101+
# a or b of 0/None must not raise and must be dropped from the population.
102+
pairs = [(f"p{i}", 1000.0, 1000.0, 8) for i in range(10)]
103+
pairs += [("zero-a", 0, 1000.0, 8), ("none-b", 1000.0, None, 8)]
104+
# No real outlier among the valid points -> empty, and no exception.
105+
assert ic.mad_outliers(pairs) == []
106+
107+
108+
def test_returned_ratio_is_the_raw_ratio_not_the_residual() -> None:
109+
# The reported number must remain the human-readable a/b ratio so reviewers
110+
# can eyeball it, even though the test runs on residuals.
111+
pairs = [(f"p{i}", *_clean_pair_at(t), t) for i, t in enumerate(TREND_THREADS)]
112+
a, b = _clean_pair_at(32)
113+
pairs.append(("Swapped-variant 32T", a * 4.0, b, 32))
114+
result = dict(ic.mad_outliers(pairs))
115+
assert math.isclose(result["Swapped-variant 32T"], round((a * 4.0) / b, 2), rel_tol=1e-6)
116+
117+
118+
def test_theil_sen_recovers_a_known_slope() -> None:
119+
xs = [math.log(t) for t in (1, 2, 4, 8, 16, 32, 64)]
120+
ys = [0.5 + 0.3 * x for x in xs] # perfect line, slope 0.3
121+
slope, intercept = ic._theil_sen(xs, ys)
122+
assert math.isclose(slope, 0.3, rel_tol=1e-9)
123+
assert math.isclose(intercept, 0.5, rel_tol=1e-9)

0 commit comments

Comments
 (0)