feat(ingest): parse CPU family and model columns - #107
Merged
Merged
Conversation
Handle rowspan family cells, stacked CPU headers, and th-only SKU rows; preserve unknown threads and deduplicate CPU names against the target dataset. Refs #99
5 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
CPU tables with a rowspan family in column 0 and a SKU in column 1 now produce complete branded model names through the existing Wikipedia CPU ingest collector. The reusable parser handles stacked headers, header units, SKU rows containing only
th, and CPU/GPU separation; Opteron is registered for normal ingest. CPU deduplication uses the requested data root and compares names/slugs symmetrically while ignoring punctuation and manufacturer prefixes, preserving model variants and normalizing PRO placement.Unknown threads remain null, including all unresolved Opteron models. Missing architecture is not filled from the product family. Every proposed record cites the exact Wikipedia page and remains
verified: false.The committed dry-run artifact lists 52 complete additions, 485 already represented models, and 195 incomplete models against TechAPI
developatbf4a0597381cdf208c1fdeceedc1559f2b770bb7. No TechAPI data was changed: the normal weekly workflow cannot select just these two pages, so this branch provides the collector, regressions, and reproducible dry-run output for the standard data-PR path againstdevelop.The live coverage scraper reproduces issue #19's 759 misses and its visible first 30 rows. The total includes 199 EPYC entries outside the two requested pages:
Thus 21 of the listed gaps are ready to fill; 738 are accounted for by the other statuses, not 738 undocumented CPUs. An additional 31 ready models were omitted by the coverage scraper's first-cell traversal, bringing the proposed additions to 52. The coverage scraper's exact unqualified-name comparison needs a separate correction; it is outside this PR's ownership. The audit report explains overlapping missing-field reasons and replay instructions.
Validation:
pytest --cov=app --cov-report=term --cov-fail-under=60: 517 passed, 77.16% coverage;ruff check app tests,mypy app, andpython -m app.validatepass.Refs #99