Skip to content

docs(catalog): a home for a project's gitignored data (HF trial) + template prompt - #5

Merged
R0SEWT merged 2 commits into
mainfrom
docs/data-storage
Sep 27, 2026
Merged

R0SEWT merged 2 commits into
mainfrom
docs/data-storage

Conversation

@R0SEWT

@R0SEWT R0SEWT commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

Summary

Every data project scaffolded by the kit gitignores data/{bronze,silver,gold}, but nothing says where that data lives or how to rebuild it. This records a research-setup evaluation in the catalog and adds a fill prompt to the CLAUDE.md template, so each new project answers that question. HF is not made a default.

Changes

  • catalog/setup-options.md: new entry, "a home for a project's gitignored data". Candidates:

    • HF dataset repos: strong for publishing our own data; weak as working storage, because deletes don't free quota without a squash.
    • HF Storage Buckets: good for private working data. A bucket is created public unless --private is passed.
    • DVC and git-annex + rclone: only worth it if dataset versions must be pinned to commits.
    • Current practice, which stays valid: $SILVER on an own server (tesis_redes) and regenerable silver (infelix).

    Verdict: trial, since no project in the stack uses HF yet.

  • templates/CLAUDE.md.tmpl: a "Where it lives" fill prompt in Data Conventions. It lists regenerate-from-source, $SILVER, a private HF bucket, and an HF dataset repo (card + license), and warns never to publish third-party or personal data.

Kept as a prompt rather than a default, to respect the kit's rule that templates stay faithful to real repos.

Testing

  • bash -n scripts/scaffold.sh and scaffold.sh … --dry-run: OK. No script changes.
  • Ran a real scaffold into a temp dir (--no-github --no-mcp). The rendered CLAUDE.md shows the new prompt under Data Conventions.
  • Maintenance data comes from PyPI and GitHub (2026-09-24); storage facts come from HF's storage-limits and storage-buckets docs.

Related

Comes from the backup and data-storage split discussed for chrome-helper (R0SEWT/chrome-helper#10). No bead: I didn't touch this repo's beads DB because the main checkout has an uncommitted .beads/issues.jsonl change.

🤖 Generated with Claude Code

Summary by Sourcery

Document where gitignored project data should live and prompt new projects to record its storage and regeneration approach.

Enhancements:

  • Document options for storing and rebuilding gitignored project data, including current remote-storage practices and a trial of Hugging Face dataset repositories and private storage buckets.

Documentation:

  • Add a catalog entry evaluating data-storage options and their privacy and versioning trade-offs.
  • Add a Data Conventions prompt to generated CLAUDE.md files requiring projects to document where gitignored data lives and how it can be rebuilt.

Tests:

  • Verify the scaffold script syntax and dry-run output, and confirm the new prompt renders in a real scaffold.

Chores:

  • Update plugin metadata for the documentation and template changes.

…t in the CLAUDE.md template

Every data project gitignores data/{bronze,silver,gold}, but the kit never said where that
data lives or how to rebuild it. Evaluated HF dataset repos, HF Storage Buckets, DVC and
git-annex + rclone against the current practice ($SILVER on an own server, regenerable
silver). Verdict: trial — no project in the stack uses HF yet.

Not wired as a default: the template gains a "Where it lives" fill prompt listing the
options, so each new data project states where its data lives and never publishes
third-party or personal data.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry @R0SEWT, you've used your own review budget of 250,000 diff characters for the last 7 days.

You can request another review in 4 days and 2 hours by commenting @sourcery-ai review. Upgrade to get a review now.

@sourcery-ai

sourcery-ai Bot commented Sep 25, 2026

Copy link
Copy Markdown

Reviewer's Guide

This documentation-only change records a research-setup trial for storing gitignored data and adds a rendered CLAUDE.md prompt so each scaffolded data project documents its storage location, rebuild path, and privacy considerations. HF remains an optional, explicitly private-by-default workflow rather than a scaffold default.

Flow diagram for documenting gitignored data storage

flowchart LR
    Layers[Gitignored data layers<br/>bronze, silver, gold] --> Prompt[CLAUDE.md Data Conventions<br/>Where it lives]
    Prompt --> Choice{Choose storage and rebuild path}
    Choice --> Regenerate[Regenerate from source]
    Choice --> Silver[Remote silver via $SILVER]
    Choice --> Bucket[Private HF Storage Bucket]
    Choice --> Dataset[HF dataset repo<br/>card and license]
    Prompt --> Privacy[Never publish third-party<br/>or personal data]
Loading

File-Level Changes

Change Details Files
Document and evaluate storage options for gitignored project data, with HF adoption explicitly treated as a trial rather than a kit default.
  • Add a catalog entry comparing HF dataset repos, HF Storage Buckets, DVC, git-annex/rclone, and existing $SILVER or regenerable-data practices.
  • Record the trial recommendation, privacy constraints, HF bucket visibility warning, and conditions for revisiting the decision.
  • Link the catalog decision to the project-template prompt without changing scaffold defaults.
catalog/setup-options.md
Require newly scaffolded data projects to state where ignored data lives and how it can be rebuilt.
  • Add a Data Conventions fill prompt covering regeneration, $SILVER, private HF buckets, and publishable HF dataset repos.
  • Include guidance to provide dataset cards and licenses for published data and never publish third-party or personal data.
templates/CLAUDE.md.tmpl

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

Comment thread templates/CLAUDE.md.tmpl
<!-- Adjust/remove if this is not a data project. Defaults reflect the medallion + polars/duckdb stack. -->

- **Layers**: raw → `data/bronze/`, cleaned → `data/silver/`, analytic → `data/gold/`. All gitignored.
- **Where it lives**: <!-- fill: gitignored data still needs a home — say where each layer lives and how

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This changes a template shipped by the plugin but doesn't bump .claude-plugin/plugin.json version (still 0.1.0). /project-kit:scaffold runs $CLAUDE_PLUGIN_ROOT/scripts/scaffold.sh, which reads templates from the cached project-kit/0.1.0/ copy.

After merge, projects scaffolded through the plugin still get the old CLAUDE.md with no "Where it lives" prompt, because the cache is keyed on the unchanged version. The PR's claimed effect only shows up when running the script straight from the checkout. PR #8's own contract requires a version bump in the same PR.

Automated review (Claude Code, AI-generated). Validate before acting.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 8dc3f37: plugin bumped to 0.2.0. (Claude Opus 5.5, on behalf of Rody)

Comment thread catalog/setup-options.md
repo (card + license) for data we publish and a **private** HF Storage Bucket for working data; DVC
or git-annex only if a project needs dataset versions pinned to commits. Never publish third-party or
personal data (course material, class recordings, PII). Revisit after one project ships.
- **Wiring**: not a default. `templates/CLAUDE.md.tmpl` gains a "Where it lives" fill prompt in Data

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Convention break: a trial verdict gets wired into the default template. The catalog format says "Wiring (if adopt)", research-setup says "If adopt, offer wiring", and the SQLite precedent says trial means "per-project only, never in the default template".

The rule in CLAUDE.md is: "templates/ ... are kept faithful to real repos (tesis_redes, exp_tesis, infelix) — reuse, don't invent." The new template bullet recommends HF buckets and dataset repos, which this entry admits "no project in the stack uses ... yet". So every new data project gets an unproven option in its default CLAUDE.md. Keep the prompt to the proven options (regenerate / $SILVER) until HF is adopted.

Automated review (Claude Code, AI-generated). Validate before acting.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Kept on purpose (Rody's call). 8dc3f37 records it in the entry as a deliberate exception: the prompt asks the question and installs no HF tooling. (Claude Opus 5.5, on behalf of Rody)

Comment thread templates/CLAUDE.md.tmpl
<!-- Adjust/remove if this is not a data project. Defaults reflect the medallion + polars/duckdb stack. -->

- **Layers**: raw → `data/bronze/`, cleaned → `data/silver/`, analytic → `data/gold/`. All gitignored.
- **Where it lives**: <!-- fill: gitignored data still needs a home — say where each layer lives and how

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The new prompt repeats the $SILVER option that the existing Remote silver bullet (L39) already covers. It also points to catalog/setup-options.md in project-kit, a path that doesn't exist in the scaffolded repo.

The rendered CLAUDE.md now describes remote silver twice in one section, so the two can drift, and the bare relative path can't be resolved from the consumer repo. Fold "Where it lives" into the Remote silver bullet, or drop the $SILVER item, and link the catalog by URL.

Automated review (Claude Code, AI-generated). Validate before acting.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 8dc3f37: it now says "remote silver (see below)" and points to the catalog by URL. (Claude Opus 5.5, on behalf of Rody)

The template change must reach the plugin cache, which is keyed on the
plugin version. Drop the duplicate $SILVER mention and point to the
catalog by URL, since the scaffolded repo has no catalog/.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@R0SEWT
R0SEWT merged commit a2954eb into main Sep 27, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant