Skip to content

data(cpu): fill process node and L3 cache for 310 Pentium/Celeron CPUs - #134

Merged
Seungpyo1007 merged 2 commits into
mainfrom
data/cpu-pentium-celeron-specs
Jul 29, 2026
Merged

Seungpyo1007 merged 2 commits into
mainfrom
data/cpu-pentium-celeron-specs

Conversation

@Seungpyo1007

@Seungpyo1007 Seungpyo1007 commented Jul 29, 2026 •

Copy link
Copy Markdown
Member

What

310 records gain 500 previously-null fields — process_node 287, l3_cache_mb 213 — from the Wikipedia Intel Celeron and Pentium list pages.

Running total with #132 and #133: CPU process_node coverage 24% → 53%, l3_cache_mb 19% → 53%.

Parser change this batch needed

The Xeon pages (#133) each covered a single microarchitecture, so the node could be taken from the page. Celeron and Pentium pages mix generations — one page spans Covington (250 nm) through modern parts — so that shortcut would have been wrong.

The parser now resolves the node from the enclosing section heading:

=== "Coppermine-128" (180 nm) ===
=== "Prescott-256" (90 nm) ===

and leaves the field empty when a heading carries no node, rather than inheriting the previous section's value. Verified across the historical range: Celeron 266 → 250 nm (Covington), Celeron 420 → 65 nm (Conroe-L). The Xeon results were re-checked after the change and are unchanged (Skylake page still 151 rows, all 14 nm).

Correctness

  • Same confirm-before-write rule: core count must match exactly and TDP within 1 W when both sides have it. 1 of 319 matches was dropped.
  • Applied values sit in the expected ranges — nodes 65/45/32/22/14/10 nm, L3 1–6 MB (Celeron/Pentium parts have small caches).
  • Source URLs were HTTP-verified before writing. One generated link returned 404 (..._Pentium_Dual_Core_..., the real title is hyphenated); it was resolved through the API to the canonical page rather than left broken in the data.

Only empty fields are written, nothing existing is overwritten, verified untouched.

python -m app.validate → Data validation passed. Dump refreshed for the 310 changed pages only.

Closes #1

310 records gain 500 previously-null fields (process_node 287,
l3_cache_mb 213) from the Wikipedia Intel Celeron and Pentium list pages.

These pages mix many generations, so the page-level node used for the Xeon
batch does not apply. The parser now resolves the node from the enclosing
section heading -- '"Coppermine-128" (180 nm)', '"Prescott-256" (90 nm)' --
and leaves it empty when the heading carries none, rather than inheriting a
neighbouring section's value.

As before a row is used only when the data already held agrees (core count
exactly, TDP within 1 W); 1 of 319 matches was dropped on that check. Source
URLs were HTTP-verified before writing -- one generated link 404'd and was
corrected to the canonical page.

Refs #1
Regenerates site/public/v1/cpus for the 310 records changed in the previous
commit, with the engine at the pinned submodule commit. Only pages whose
content actually differs are committed; 7,645 pages that differed solely by
created_at/updated_at are left alone.

Refs #1
@github-actions github-actions Bot added data Dataset changes enhancement New feature or request labels Jul 29, 2026
@Seungpyo1007
Seungpyo1007 merged commit 02c8556 into main Jul 29, 2026
5 checks passed
@github-project-automation github-project-automation Bot moved this from Todo to Done in TechAPI-Project Jul 29, 2026
@Seungpyo1007
Seungpyo1007 deleted the data/cpu-pentium-celeron-specs branch July 29, 2026 07:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

data Dataset changes enhancement New feature or request

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

TechAPI dataset roadmap and status: all categories (1989-2026)

1 participant