Skip to content

docs(onchain-analytics): add current system guide - #74

Merged
thalescb merged 1 commit into
masterfrom
docs/onchain-analytics-start-here
Oct 6, 2026
Merged

thalescb merged 1 commit into
masterfrom
docs/onchain-analytics-start-here

Conversation

@thalescb

@thalescb thalescb commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

Adds a current onboarding guide that distinguishes the live XDC models from the prepared, uncommissioned raw pipeline.

Also corrects status and stale-reference notices across the onchain analytics docs, and removes local scratch receipt paths from the public gate comments.

Where to look

  • Start with projects/onchain-analytics/docs/START_HERE.md (421 lines).
  • Then projects/onchain-analytics/README.md (188 lines).
  • Supporting detail: projects/onchain-analytics/pipeline-v5/README.md (85 lines) and projects/onchain-analytics/docs/03_OPERATIONS.md (75 lines).
  • Small status/history notices: docs/00_VISION.md, docs/01_ARCHITECTURE.md, docs/02_DATA_MODEL.md, docs/release-scope.md, and .shipgate-allow.

Validation

  • Public leak gate: 9/9 files passed, zero warnings or blocks.
  • git diff --check: clean. Documentation-only; no runtime tests run.
  • Scope gate: 820 lines across 9 files; shipped as one guided review as authorized.

@thalescb thalescb mentioned this pull request Oct 6, 2026
6 of 9 tasks
@thalescb
thalescb requested a balanced review from Copilot October 6, 2026 02:45

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The guide contains several schema and operational inaccuracies, including a cron example guaranteed to select an unsupported chain.

Review effort: Balanced
Findings: 5 Low severity

Open (5)
What changed in this PR

Adds a current onboarding guide clarifying live XDC analytics versus the uncommissioned raw pipeline.

Changes:

  • Adds START_HERE.md with architecture, status, safety, and commissioning guidance.
  • Updates operational and pipeline documentation.
  • Marks legacy docs and removes scratch receipt references.
File Description
.shipgate-allow Simplifies public gate rationale comments.
projects/​onchain-analytics/​README.md Updates system overview and production status.
projects/​onchain-analytics/​pipeline-v5/​README.md Clarifies pipeline modes, limits, and readiness.
projects/​onchain-analytics/​docs/​START_HERE.md Adds the current onboarding guide.
projects/​onchain-analytics/​docs/​release-scope.md Updates chain ingestion status.
projects/​onchain-analytics/​docs/​03_OPERATIONS.md Corrects operational commands and safeguards.
projects/​onchain-analytics/​docs/​02_DATA_MODEL.md Clarifies legacy model scope.
projects/​onchain-analytics/​docs/​01_ARCHITECTURE.md Marks outdated architecture guidance.
projects/​onchain-analytics/​docs/​00_VISION.md Labels the document as historical.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

last measured on 2026-09-28. Reconcile the workflow with the keyless identity design and verify its
authentication before using it. Do not add a long-lived service-account key for convenience, and
do not schedule ingestion until production ingestion is authorized. When it is time to automate,
the daily flow is two ordered steps: ingest first, then dbt. The examples below are illustrative.
Comment on lines +108 to +114
### Provenance on every row

Every raw row records where it came from: `source_kind` and `source_id` (which reader answered),
`assurance` (how independently confirmed the read was; see section 5), `capture_id` and
`ingestion_run_id` (which capture and run wrote it), `ingested_at`, and the contract era fields
(`implementation_address`, `era_index`, `era_resolution`). If the era cannot be resolved, the row
says `unresolved` and does not guess.
| Table | Answers |
| - | - |
| `IngestionCoverage` | Which reader read which block range of which contract, what it found, what failed, and how far the result can be trusted. The pipeline writes one row for every attempted range, including failures and refusals. **The next run resumes from this table**, not from the highest block number in the data. A block range with no coverage row was never read. |
| `PipelineRuns` | One row per command, with its exit code and outcome counts |
| Explicit coverage | Resume points come from `IngestionCoverage`. Gaps, refusals, and unreadable ranges are recorded, never inferred from the data. |
| Retries and deadlines | Each HyperSync chunk runs in a child process that is killed after 120 seconds. One contract's whole range is abandoned after 90 minutes, and the blocks not attempted are recorded. Every RPC call has a 30-second deadline. BigQuery and HyperSync calls retry up to 5 times with backoff, and requests are paced. |
| Finality | Ingestion stays behind the chain tip: 15 blocks on XDC, 1,930 blocks (about 32 minutes) on Celo, 94 blocks on Ethereum. A range the reader still holds as reversible is recorded `rollback_eligible` and read again later. |
| Cost and size budgets | Every BigQuery job carries a 10 GiB `maximumBytesBilled` cap. A run may attempt at most 12 contracts. One capture may span 30 days of blocks unless the range is named with `--from` and `--to`, and named ranges are capped at 100,000,000 blocks. A bare `backfill` and a one-sided range are both refused. |
Comment on lines +11 to 12
| Celo | Yes | No, not yet in the new raw pipeline | Primary chain for claims; production ingestion has not run |
| XDC | Yes | Yes | Holds the existing invites and claims dataset |
@thalescb
thalescb merged commit d24f6f0 into master Oct 6, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants