Repository navigation
docs(onchain-analytics): add current system guide - #74
Merged
Merged
Conversation
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The guide contains several schema and operational inaccuracies, including a cron example guaranteed to select an unsupported chain.
Review effort: Balanced
Findings: 5
Open (5)
What changed in this PR
Adds a current onboarding guide clarifying live XDC analytics versus the uncommissioned raw pipeline.
Changes:
- Adds
START_HERE.mdwith architecture, status, safety, and commissioning guidance. - Updates operational and pipeline documentation.
- Marks legacy docs and removes scratch receipt references.
| File | Description |
|---|---|
.shipgate-allow |
Simplifies public gate rationale comments. |
projects/onchain-analytics/README.md |
Updates system overview and production status. |
projects/onchain-analytics/pipeline-v5/README.md |
Clarifies pipeline modes, limits, and readiness. |
projects/onchain-analytics/docs/START_HERE.md |
Adds the current onboarding guide. |
projects/onchain-analytics/docs/release-scope.md |
Updates chain ingestion status. |
projects/onchain-analytics/docs/03_OPERATIONS.md |
Corrects operational commands and safeguards. |
projects/onchain-analytics/docs/02_DATA_MODEL.md |
Clarifies legacy model scope. |
projects/onchain-analytics/docs/01_ARCHITECTURE.md |
Marks outdated architecture guidance. |
projects/onchain-analytics/docs/00_VISION.md |
Labels the document as historical. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| last measured on 2026-09-28. Reconcile the workflow with the keyless identity design and verify its | ||
| authentication before using it. Do not add a long-lived service-account key for convenience, and | ||
| do not schedule ingestion until production ingestion is authorized. When it is time to automate, | ||
| the daily flow is two ordered steps: ingest first, then dbt. The examples below are illustrative. |
Comment on lines
+108
to
+114
| ### Provenance on every row | ||
|
|
||
| Every raw row records where it came from: `source_kind` and `source_id` (which reader answered), | ||
| `assurance` (how independently confirmed the read was; see section 5), `capture_id` and | ||
| `ingestion_run_id` (which capture and run wrote it), `ingested_at`, and the contract era fields | ||
| (`implementation_address`, `era_index`, `era_resolution`). If the era cannot be resolved, the row | ||
| says `unresolved` and does not guess. |
| | Table | Answers | | ||
| | - | - | | ||
| | `IngestionCoverage` | Which reader read which block range of which contract, what it found, what failed, and how far the result can be trusted. The pipeline writes one row for every attempted range, including failures and refusals. **The next run resumes from this table**, not from the highest block number in the data. A block range with no coverage row was never read. | | ||
| | `PipelineRuns` | One row per command, with its exit code and outcome counts | |
| | Explicit coverage | Resume points come from `IngestionCoverage`. Gaps, refusals, and unreadable ranges are recorded, never inferred from the data. | | ||
| | Retries and deadlines | Each HyperSync chunk runs in a child process that is killed after 120 seconds. One contract's whole range is abandoned after 90 minutes, and the blocks not attempted are recorded. Every RPC call has a 30-second deadline. BigQuery and HyperSync calls retry up to 5 times with backoff, and requests are paced. | | ||
| | Finality | Ingestion stays behind the chain tip: 15 blocks on XDC, 1,930 blocks (about 32 minutes) on Celo, 94 blocks on Ethereum. A range the reader still holds as reversible is recorded `rollback_eligible` and read again later. | | ||
| | Cost and size budgets | Every BigQuery job carries a 10 GiB `maximumBytesBilled` cap. A run may attempt at most 12 contracts. One capture may span 30 days of blocks unless the range is named with `--from` and `--to`, and named ranges are capped at 100,000,000 blocks. A bare `backfill` and a one-sided range are both refused. | |
Comment on lines
+11
to
12
| | Celo | Yes | No, not yet in the new raw pipeline | Primary chain for claims; production ingestion has not run | | ||
| | XDC | Yes | Yes | Holds the existing invites and claims dataset | |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Adds a current onboarding guide that distinguishes the live XDC models from the prepared, uncommissioned raw pipeline.
Also corrects status and stale-reference notices across the onchain analytics docs, and removes local scratch receipt paths from the public gate comments.
Where to look
Validation