Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10,675 changes: 10,675 additions & 0 deletions app/ingest/artifacts/amd-cpu-dry-run.json

Large diffs are not rendered by default.

89 changes: 89 additions & 0 deletions app/ingest/artifacts/amd-cpu-report.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,89 @@
# AMD CPU ingest audit

The reusable CPU table parser now handles stacked CPU headers, separate family
and model columns, rowspan carry, and SKU rows made entirely of `th` cells.
It reads units from headers, excludes GPU/NPU columns, retains variant suffixes,
and does not infer threads from a core count. Architecture comes from the source
column or section, never the product family. Opteron is registered in the existing
CPU collector and therefore uses the normal ingest pipeline and weekly workflow.

CPU additions are deduplicated symmetrically against both names and slugs in the
requested TechAPI data root, ignoring punctuation and manufacturer prefixes.
PRO placement is normalized while PRO/non-PRO and H/HS/U/X/HE variants stay
distinct. Records cite their exact Wikipedia page and stay `verified: false`.

## Dry run against TechAPI develop

Dataset revision: `bf4a0597381cdf208c1fdeceedc1559f2b770bb7` (1,479 AMD CPU
records). The artifact records the dataset content hash and source HTML hashes.

| Live tables | Unique models | Ready additions | Already curated | Incomplete |
| --- | ---: | ---: | ---: | ---: |
| Ryzen | 445 | 52 | 383 | 10 |
| Opteron | 287 | 0 | 102 | 185 |
| Total | 732 | 52 | 485 | 195 |

All 185 unresolved Opteron models lack a stated thread count; 76 also lack a
per-row core count. Across both pages, 22 unresolved models lack a usable release
date. These reasons overlap. Missing table specifications are not evidence that
no public specification exists elsewhere; other authoritative sources can fill
them later without guessing.

The dry-run JSON lists all 52 complete records with their intended output paths,
all incomplete models with reasons, and every already represented model.
No TechAPI files were written and no TechAPI data PR was opened. The existing
weekly workflow creates data PRs against `develop`, but it cannot select just
these two pages and this branch's collector is not deployed on `main`; dispatching
it would also ingest unrelated CPU pages. This is the requested dry-run fallback.

## Reconciliation of issue #19's 759 gaps

The live coverage scraper reproduces **759** exact-slug misses and exactly the
30 rows shown in issue #19: Ryzen contributes 277 entries, Opteron 283, and EPYC
199. Thus the original 759 includes a third page beyond the two requested pages.

| Reproduced gap status | Entries |
| --- | ---: |
| Ready to fill through normal ingest | 21 |
| Already curated under a complete branded name | 324 |
| Missing required table specifications | 167 |
| Family tiers, stepping captions, dates, or part-number cells | 48 |
| EPYC entries outside this two-page audit | 199 |
| Total | 759 |

Of the 759, **21 can be filled by the proposed additions; 738 are not new
additions from this audit**, partitioned above. The 167 incomplete gap entries
include 161 with missing threads, 52 with missing cores, and 17 with missing dates
(overlapping reasons). The other **31 of the 52 ready additions** were omitted
by the coverage scraper's first-cell traversal. Every reproduced entry and its
canonical candidate association appears in the JSON's `coverage_reconciliation`.

These are proposed additions, not applied data changes. The coverage issue's
exact comparison of unqualified table cells with branded curated slugs will
continue to produce false positives until the coverage collector is separately
updated; that collector is outside this task's ownership.

## Reproduce

```powershell
python -m app.ingest.cpu_audit `
--page List_of_AMD_Ryzen_processors `
--page List_of_AMD_Opteron_processors `
--coverage-page List_of_AMD_Ryzen_processors `
--coverage-page List_of_AMD_Opteron_processors `
--coverage-page List_of_AMD_Epyc_processors `
--data-root ../TechAPI/data `
--output app/ingest/artifacts/amd-cpu-dry-run.json
```

For an offline replay, add `--html-dir PATH` containing the three downloaded
`List_of_AMD_*_processors.html` files. Without that flag, the engine's regular
Wikipedia fetcher downloads the pages using its declared user-agent.

## Validation

Repository-wide `ruff check app tests`, `mypy app`, and `python -m app.validate`
pass. The full suite passes: **517 tests**, with **77.16% coverage** against the
60% threshold. The focused parser/pipeline selection contains 36 passing tests.

Refs #99
176 changes: 176 additions & 0 deletions app/ingest/cpu_audit.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,176 @@
"""Read-only, reproducible CPU ingest audit; never writes to the dataset.

Example::

python -m app.ingest.cpu_audit --page List_of_AMD_Ryzen_processors \
--page List_of_AMD_Opteron_processors --data-root ../TechAPI/data \
--output amd-cpu-dry-run.json

``--html-dir`` replays previously downloaded HTML instead of fetching pages.
The JSON includes every proposed record and every unresolved unique model.
"""

from __future__ import annotations

import argparse
import hashlib
import json
from collections import Counter
from datetime import UTC, datetime
from pathlib import Path

from app.coverage.sources.wikipedia import fetch_wikipedia_html
from app.coverage.sources.wikipedia_cpu import WikipediaCpu

from .pipeline import run
from .sources.base import IngestCandidate
from .sources.wikipedia_cpu import PAGES, WikipediaCpuIngest


def _entry(candidate: IngestCandidate) -> dict[str, object]:
return {
"output_path": candidate.output_path.as_posix(),
"record": candidate.record,
"missing_fields": list(candidate.missing_fields),
}


def _coverage_audit(
html_by_page: dict[str, str],
candidates: list[IngestCandidate],
ready: set[str],
existing: set[str],
) -> dict[str, object]:
points = {
point.slug: point
for page, html in html_by_page.items()
for point in WikipediaCpu._extract(html, "amd", page)
}
entries = []
for slug, point in sorted(points.items()):
matches = [
c
for c in candidates
if c.source_url == point.url and (c.slug == slug or c.slug.endswith("-" + slug))
]
if not any(c.source_url == point.url for c in candidates):
status = "outside_requested_pages"
elif any(c.slug in ready for c in matches):
status = "ready_to_add"
elif any(c.slug in existing for c in matches):
status = "already_curated"
elif matches:
status = "missing_required_specs"
else:
status = "non_model_or_unparsed_cell"
entries.append(
{
"coverage_slug": slug,
"source_url": point.url,
"status": status,
"candidate_slugs": sorted({c.slug for c in matches}),
"missing_fields": sorted({field for c in matches for field in c.missing_fields}),
}
)
return {
"total": len(points),
"counts": dict(Counter(entry["status"] for entry in entries)),
"entries": entries,
}


def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--page", action="append", required=True, choices=[p[1] for p in PAGES])
parser.add_argument("--data-root", required=True, type=Path)
parser.add_argument("--output", required=True, type=Path)
parser.add_argument("--html-dir", type=Path)
parser.add_argument(
"--coverage-page",
action="append",
default=[],
help="Optional AMD pages whose raw coverage entries should be reconciled.",
)
args = parser.parse_args(argv)
if not args.data_root.is_dir():
parser.error("--data-root must be an existing TechAPI data directory")
candidates: list[IngestCandidate] = []
sources = []
html_by_page: dict[str, str] = {}
for manufacturer, page, family in PAGES:
if page not in args.page:
continue
html = (
(args.html_dir / f"{page}.html").read_text(encoding="utf-8")
if args.html_dir
else fetch_wikipedia_html(page)
)
sources.append(
{
"url": f"https://en.wikipedia.org/wiki/{page}",
"html_sha256": hashlib.sha256(html.encode()).hexdigest(),
}
)
html_by_page[page] = html
candidates.extend(WikipediaCpuIngest._extract(html, manufacturer, page, family))
result = run(candidates, data_root=args.data_root, dry_run=True)
existing = {c.slug: c for c in result.skipped_existing}
incomplete = {c.slug: c for c in result.skipped_incomplete if c.slug not in existing}
missing = Counter(field for c in incomplete.values() for field in c.missing_fields)
snapshot = hashlib.sha256()
manufacturers = {c.manufacturer for c in candidates}
curated_paths = sorted(
path for maker in manufacturers for path in (args.data_root / "cpu" / maker).rglob("*.json")
)
for path in curated_paths:
snapshot.update(path.relative_to(args.data_root).as_posix().encode())
snapshot.update(b"\0")
snapshot.update(path.read_bytes())
payload = {
"generated_at": datetime.now(UTC).isoformat(),
"dry_run": True,
"include_drafts": False,
"sources": sources,
"curated_cpu_snapshot": {
"records": len(curated_paths),
"sha256": snapshot.hexdigest(),
},
"counts": {
"candidate_rows": len(candidates),
"unique_models": len({c.slug for c in candidates}),
"would_add": len(result.written),
"already_existing": len(existing),
"incomplete": len(incomplete),
"missing_fields": dict(missing),
},
"would_add": [_entry(c) for c in result.written],
"already_existing": sorted(existing),
"incomplete": [
{"slug": c.slug, "missing_fields": list(c.missing_fields), "source_url": c.source_url}
for c in incomplete.values()
],
}
if args.coverage_page:
for page in args.coverage_page:
if page not in html_by_page:
html_by_page[page] = (
(args.html_dir / f"{page}.html").read_text(encoding="utf-8")
if args.html_dir
else fetch_wikipedia_html(page)
)
payload["coverage_reconciliation"] = _coverage_audit(
{page: html_by_page[page] for page in args.coverage_page},
candidates,
{c.slug for c in result.written},
set(existing),
)
args.output.parent.mkdir(parents=True, exist_ok=True)
args.output.write_text(
json.dumps(payload, indent=2, ensure_ascii=False) + "\n", encoding="utf-8"
)
print(json.dumps(payload["counts"]))
return 0


if __name__ == "__main__":
raise SystemExit(main())
51 changes: 51 additions & 0 deletions app/ingest/cpu_identity.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
"""Conservative, symmetric CPU identity checks for additions-only ingest."""

from __future__ import annotations

import json
import re
from pathlib import Path

from app.coverage.normalize import slugify


def cpu_key(value: str, manufacturer: str) -> str:
tokens = slugify(value, manufacturer=manufacturer).split("-")
if tokens and tokens[0] == manufacturer:
tokens.pop(0)
# Wikipedia uses both "1700X PRO" and "PRO 1700X" for the same SKU.
if "pro" in tokens:
tokens = [token for token in tokens if token != "pro"] + ["pro"]
return "".join(tokens)


def same_cpu(left: str, right: str) -> bool:
if not left or not right:
return False
if left == right:
return True
short, long = sorted((left, right), key=len)
if len(short) < 4 or not long.endswith(short):
return False
# Bare 1200 must not match 41200. Suffixes (X, U, HE, PRO) remain identity.
prefix = long[: -len(short)]
return (
not (short[0].isdigit() and prefix[-1].isdigit())
or re.search(r"[a-z]\d{1,2}$", prefix) is not None
)


def curated_cpu_keys(data_root: Path, manufacturer: str) -> set[str]:
keys: set[str] = set()
for path in (data_root / "cpu" / manufacturer).rglob("*.json"):
try:
record = json.loads(path.read_text(encoding="utf-8"))
except (OSError, json.JSONDecodeError):
continue
if not isinstance(record, dict):
continue
for field in ("slug", "name"):
value = record.get(field)
if isinstance(value, str):
keys.add(cpu_key(value, manufacturer))
return keys
19 changes: 17 additions & 2 deletions app/ingest/pipeline.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,7 @@

from app.coverage.curated import curated_slugs

from .cpu_identity import cpu_key, curated_cpu_keys, same_cpu
from .sources.base import IngestCandidate


Expand Down Expand Up @@ -59,12 +60,24 @@ def run(
result = IngestResult()
curated_by_category: dict[tuple[str, str], set[str]] = {}
written_slugs: set[tuple[str, str, str]] = set()
cpu_keys: dict[str, set[str]] = {}

for candidate in candidates:
key = (candidate.category, candidate.manufacturer)
if candidate.category == "cpu":
if candidate.manufacturer not in cpu_keys:
cpu_keys[candidate.manufacturer] = curated_cpu_keys(
data_root, candidate.manufacturer
)
identity = cpu_key(candidate.slug, candidate.manufacturer)
if any(same_cpu(identity, other) for other in cpu_keys[candidate.manufacturer]):
result.skipped_existing.append(candidate)
continue
if key not in curated_by_category:
curated_by_category[key] = curated_slugs(
candidate.category, candidate.manufacturer
curated_by_category[key] = (
set()
if candidate.category == "cpu"
else curated_slugs(candidate.category, candidate.manufacturer)
)
if candidate.slug in curated_by_category[key]:
result.skipped_existing.append(candidate)
Expand All @@ -78,6 +91,8 @@ def run(
continue

written_slugs.add(run_key)
if candidate.category == "cpu":
cpu_keys[candidate.manufacturer].add(identity)
target = data_root / candidate.output_path
if not dry_run:
target.parent.mkdir(parents=True, exist_ok=True)
Expand Down
Loading
Loading