Skip to content

ci: scope every CI job to the changed data instead of the full dataset #350

Description

@Seungpyo1007

Why

The 2026-09-29/30 mass-agent push (18 concurrent workers on one develop) exposed how expensive our CI is: every PR re-reads the whole dataset (~190k records / ~292k dump files) no matter how little changed, dump refresh takes ~33 min and re-runs on every rebase, and the automation volume coincided with the TechEngineBot account suspension.

Goal

Make every CI job scale with the size of the change, not the size of the dataset.

Plan

  1. Unblock – drop the suspended TechEngineBot from pr-metadata assignees (404s every PR).
  2. Scope calculator (TechEngine) – changed paths -> {category, files, full_run_required}; falls back to a full run when code/schema changed.
  3. Scoped loading – load_category accepts a file list; brand/soc slug sets for FK checks stay full (small).
  4. verify-comment – fetch-depth: 1, read data/_verify/status.json instead of rescoring everything, score once.
  5. validate – PRs validate changed files (+FK targets); full validation only on develop/main push and nightly.
  6. dump – per-category and per-path generation, category matrix; move dump refresh off the per-PR critical path (needs a decision on the "dump is the last commit of every data PR" rule).
  7. deploy-pages – precompute the homepage history timeline so the build no longer needs fetch-depth: 0.
  8. Cleanup – stale worktrees/branches (29 worktrees, 111 local branches), stale dump-refresh/* PRs.

Related: #297 (automation), TechEngine #97-#100.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

ciCI and workflow changes

Type

No type

Projects

Relationships

None yet

Development

No branches or pull requests

Issue actions