data(cpu): recover 210 CPUs the earlier batches skipped on naming mismatches - #137
Merged
Merged
Conversation
…kipped Recovers records that the merged batches (#132-#135) left untouched because of three separate mismatches in how the pages name things, not because the data was unavailable: - model codes are not always digit-first (Core 2 uses E6600, Atom uses N270), and the table parser required a leading digit; - the Atom page is built entirely from {{cpulist}} templates and had never been run through that parser; - the Opteron page lists a bare model-number column with no branding column, so the family name now comes from the page itself. Same confirm-before-write rule as before (base clock within 0.05 GHz, TDP within 1 W, L3 must be a real capacity and must not equal the row's TDP): 1,610 rows parsed, 1,053 matched a record, 885 passed, 210 still had a gap to fill. Verified against known parts: Clarkdale i5-6xx/i3-5xx 32 nm, Haswell i7-4770R 22 nm, Broadwell i7-5775R 14 nm, Atom 330/D410 45 nm, Atom D2500 32 nm, Opteron Shanghai 6 MB L3. Older families whose rows failed the cross-check (Core 2 E-series, Atom N270, early Opteron) were left empty rather than guessed. Refs #1
Regenerates site/public/v1/cpus for the 210 records changed in the previous commit, with the engine at the pinned submodule commit. Only pages whose content actually differs are committed; 7,745 timestamp-only pages are left alone. Refs #1
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
210 records / 240 fields (
process_node205,l3_cache_mb35). CPUprocess_nodecoverage 56% → 61%,l3_cache_mb54% → 55%.Why these were missed, not missing
After #132–#135 merged, 1,737 CPUs still had no
process_node— and several whole families sat at 0% filled (A-Series 213/213, Ryzen 125/125, Core 2 124/124, Opteron 103/103, Atom 78/78). That looked like absent data. It wasn't; it was three naming mismatches:E6600/Q6600and Atom'sN270/Z530were discarded before any matching happened.{{cpulist}}templates — and had only ever been run through the table parser.Correctness
1,610 rows parsed → 1,053 matched a record → 885 passed the cross-check → 210 still had a gap to fill. Same rules as the merged batches: base clock within 0.05 GHz, TDP within 1 W, L3 must be a real cache capacity and must not equal the row's own TDP.
Spot-checked against known parts, all correct:
Older families were left empty on purpose. Core 2 E-series, Atom N270 and early Opteron rows did not pass the cross-check, so nothing was written for them rather than guessed — which is why no 65 nm or 130 nm values appear here at all.
All 4 source pages HTTP-verified; every record cites the page it came from.
python -m app.validate→ Data validation passed. Dump refreshed for the 210 changed pages only.Closes #1