From 6b98e3d4d247abc8f7cdf20e6915fb6e73d53172 Mon Sep 17 00:00:00 2001 From: Thales Barbosa Date: Mon, 5 Oct 2026 23:40:27 -0300 Subject: [PATCH] docs(onchain-analytics): add current system guide --- .shipgate-allow | 24 +- projects/onchain-analytics/README.md | 188 +++++--- projects/onchain-analytics/docs/00_VISION.md | 5 + .../onchain-analytics/docs/01_ARCHITECTURE.md | 8 + .../onchain-analytics/docs/02_DATA_MODEL.md | 5 + .../onchain-analytics/docs/03_OPERATIONS.md | 75 +++- projects/onchain-analytics/docs/START_HERE.md | 421 ++++++++++++++++++ .../onchain-analytics/docs/release-scope.md | 9 +- .../onchain-analytics/pipeline-v5/README.md | 85 ++-- 9 files changed, 674 insertions(+), 146 deletions(-) create mode 100644 projects/onchain-analytics/docs/START_HERE.md diff --git a/.shipgate-allow b/.shipgate-allow index 8ed9fa6..d81d51d 100644 --- a/.shipgate-allow +++ b/.shipgate-allow @@ -11,8 +11,7 @@ # "was this actually checked?". # --- contract registries: public infrastructure, not per-account data -------- -# These list GoodDollar/Uniswap/Ubeswap contract addresses. Confirmed 2026-09-23 -# that none of them contain user wallets. +# These files list deployed contract addresses, not user wallet balances. docs/subgraph/*.txt :: content:address-list :: subgraph manifests, contract addresses and start blocks only projects/onchain-analytics/contracts/ABIs/*.csv :: content:address-list :: contract ABI export, deployed contract addresses only projects/onchain-analytics/contracts/README.md :: content:address-list :: the contract registry doc, by definition a list of contracts @@ -22,25 +21,10 @@ queries/dune/reserve-analysis/lp-v3-positions.sql :: content:address-list :: Uni projects/onchain-analytics/gd_dbt/seeds/contract_deployments.csv :: content:holder-table :: contract address paired with its deployment BLOCK, not a holder balance projects/onchain-analytics/gd_dbt/seeds/tokens.csv :: content:holder-table :: token contract address paired with its DECIMALS, not a holder balance # event_surface.csv is the event ABI catalogue: one row per (contract, era, event). -# Checked 2026-09-28 before waiving, not assumed. All 347 distinct addresses in the -# file are contract addresses already present in contract_deployments.csv above, -# including the 11 that appear outside the two address columns (all in abi_source, -# all of the form bytecode_equality_with_celo_0x..., naming a registry contract). -# Its only numeric columns are chain_id, era_index, indexed_positions and the two -# era block bounds. No balance column and no wallet address exists in it. -# Receipt: specs/_scratch/coord-unit-01/out/01-event-surface-address-provenance.json +# It contains contract addresses and event/era metadata, not holder balances. projects/onchain-analytics/gd_dbt/seeds/event_surface.csv :: content:holder-table :: contract address paired with its era BLOCK BOUNDS, not a holder balance -# era_intervals.csv and era_boundary_evidence.csv are the era validity-interval table -# and the evidence behind it: which implementation was live between which blocks, and -# how that was established. Checked 2026-09-28 before waiving, not assumed. -# era_intervals.csv 262 rows, 278 distinct addresses, ALL already present in -# contract_deployments.csv above. Numeric columns are chain_id, -# valid_from_block, valid_to_block, era_index. -# era_boundary_evidence.csv 41 rows, 44 distinct addresses, ALL already present above. -# Numeric columns are block numbers plus COUNTS of verification -# checkpoints, events and chunks. -# Neither file has an amount column, a balance, or a wallet address. -# Receipt: specs/_scratch/coord-unit-02-gate/verify-new-seeds.mjs +# These files describe contract implementation eras and the block evidence for their boundaries. +# Their address-and-number pairs are deployment metadata, not per-account holdings. projects/onchain-analytics/gd_dbt/seeds/era_intervals.csv :: content:holder-table :: contract address paired with its era BLOCK BOUNDS, not a holder balance projects/onchain-analytics/gd_dbt/seeds/era_boundary_evidence.csv :: content:holder-table :: contract address paired with block numbers and verification COUNTS, not a holder balance diff --git a/projects/onchain-analytics/README.md b/projects/onchain-analytics/README.md index 11db1d1..38c1995 100644 --- a/projects/onchain-analytics/README.md +++ b/projects/onchain-analytics/README.md @@ -1,6 +1,16 @@ # GoodDollar Onchain Analytics System -A production analytics platform on BigQuery that answers any question about GoodDollar onchain activity — from raw blockchain events through business-ready dashboards. Built for the GoodDollar data team and leadership. +An analytics platform on BigQuery for questions about GoodDollar onchain activity, from raw +blockchain events through business-ready dashboards. Built for the GoodDollar data team and +leadership. + +> **New to this system? Read [`docs/START_HERE.md`](docs/START_HERE.md) first.** It explains what +> is live, what is prepared, how data flows, and which documents are current. +> +> **Status, 2026-10-05:** existing dbt models and dashboards are live on the older XDC +> per-contract tables. The new ingestion pipeline and its raw-table migrations are merged and +> sandbox-tested, but **have not been applied or run in production**. There is no scheduled +> ingestion. --- @@ -9,43 +19,50 @@ A production analytics platform on BigQuery that answers any question about Good The system is composed of five layers, each with a clear responsibility: ``` -┌─────────────────────────────────────────────────────────────────┐ -│ GoodDollar Onchain Analytics System │ -├─────────────────────────────────────────────────────────────────┤ -│ Layer 5 │ Self-Service AI [planned] │ -│ Layer 4 │ Dashboards Looker Studio / Metabase │ -│ Layer 3 │ Marts (L3) Pre-aggregated, KPI-ready │ -│ Layer 2 │ Semantic Models (L2) Business logic, defined once │ -│ Layer 1 │ Staging (L1) Normalized, filtered │ -│ Layer 0 │ Ingestion Pipeline HyperSync → BigQuery raw │ -├─────────────────────────────────────────────────────────────────┤ -│ Governance │ Documentation, glossary, contracts, tests │ -└─────────────────────────────────────────────────────────────────┘ ++-----------------------------------------------------------------------+ +| GoodDollar Onchain Analytics System | ++-----------------------------------------------------------------------+ +| Layer 5 | Self-Service AI [planned] | +| Layer 4 | Dashboards Looker Studio | +| Layer 3 | Marts Pre-aggregated, KPI-ready | +| Layer 2 | Semantic Models Business logic, defined once | +| Layer 1 | Staging Normalized, filtered | +| Layer 0 | Ingestion Pipeline HyperSync / RPC -> BigQuery raw | ++-----------------------------------------------------------------------+ +| Governance | Documentation, glossary, contracts, tests | ++-----------------------------------------------------------------------+ ``` -### Layer 0 — Ingestion Pipeline +### Layer 0 -- Ingestion Pipeline -TypeScript application using [Envio HyperSync](https://docs.envio.dev/docs/HyperSync/overview) for high-throughput historical + live indexing of blockchain events into BigQuery raw tables (`gooddollar.BlockchainEvents.*`). +A TypeScript application that reads contract logs with +[Envio HyperSync](https://docs.envio.dev/docs/HyperSync/overview), or over JSON-RPC on chains with +no HyperSync index. It writes them **undecoded** into the BigQuery raw tables +`gooddollar.BlockchainEvents.RawLogs` and `Transactions`, together with a coverage record of every +range it read. Not yet run in production; see [`docs/START_HERE.md`](docs/START_HERE.md). -→ [`pipeline/`](pipeline/) +-> [`pipeline-v5/`](pipeline-v5/) -### Layers 1–3 — dbt Warehouse (Medallion Architecture) +### Layers 1-3 -- dbt Warehouse (Medallion Architecture) Managed by [dbt Core](https://docs.getdbt.com/) on BigQuery: | Layer | Dataset | What it does | Materialization | |---|---|---|---| -| **Staging (L1)** | `Staging` | Normalizes raw events — lowercase addresses, type casting, filters out partial-day data | Views | -| **Semantic (L2)** | `Semantic` | Defines business entities exactly once — signups, payouts, claims, lifecycles | Views | -| **Marts (L3)** | `Marts` | Pre-aggregated tables shaped for dashboards — daily metrics, funnels, KPIs | Tables | +| **Staging** | `Staging` | Normalizes raw events -- lowercase addresses, type casting, filters out partial-day data | Views | +| **Semantic** | `Semantic` | Defines business entities exactly once -- signups, payouts, claims, lifecycles | Views | +| **Marts** | `Marts` | Pre-aggregated tables shaped for dashboards -- daily metrics, funnels, KPIs | Tables | -→ [`gd_dbt/`](gd_dbt/) +The current models read the older per-contract raw tables (`ClaimContractEvents`, +`InviteContractEvents`). Models over the new `RawLogs` and `Transactions` tables are not built yet. -### Layer 4 — Dashboards +-> [`gd_dbt/`](gd_dbt/) -Google Looker Studio connected to L3 Marts. Future: Metabase for broader self-service. +### Layer 4 -- Dashboards -### Layer 5 — Self-Service AI (planned) +Google Looker Studio connected to the Marts. Future: Metabase for broader self-service. + +### Layer 5 -- Self-Service AI (planned) AI analyst with governed access across all layers, grounded by the semantic layer, business glossary, and disambiguation protocol. @@ -53,73 +70,86 @@ AI analyst with governed access across all layers, grounded by the semantic laye Documentation contracts, business glossary, and AI-readiness gates that every model must pass before production. -→ [`docs/`](docs/) +-> [`docs/`](docs/) --- -## How It Works — The 2-Minute Version +## How It Works -- The 2-Minute Version ### Data Flow (end to end) ``` -Blockchain → Pipeline (L0) → BigQuery raw tables → dbt Staging (L1) → dbt Semantic (L2) → dbt Marts (L3) → Dashboards +Blockchain -> Pipeline -> BigQuery raw tables -> dbt Staging -> dbt Semantic -> dbt Marts -> Dashboards ``` -1. **Pipeline ingests raw events** from the blockchain into BigQuery (`BlockchainEvents.*` tables). This is the only component that talks to the chain. -2. **dbt transforms everything above raw.** One command (`dbt run`) rebuilds Staging → Semantic → Marts in the correct order, automatically. -3. **Dashboards read from Marts only.** Pre-aggregated, tested, documented — what you see is what was defined in code. +1. **The pipeline captures raw logs** from the blockchain into BigQuery (`BlockchainEvents`). It + is the only component that talks to the chain, and it does not decode anything. +2. **dbt transforms everything above raw.** One command (`dbt run`) rebuilds Staging, Semantic and + Marts in dependency order. +3. **Dashboards read from Marts only.** Pre-aggregated, tested, documented -- what you see is what + was defined in code. ### Who Owns What | Concern | Where it lives | Key files | |---|---|---| -| **Chain connection & event decoding** | Pipeline (TypeScript) | `pipeline-v5/src/`, runbook in `pipeline-v5/README.md` | -| **Raw table schemas (DDL)** | `warehouse/L1/` | `01_ClaimContractEvents.sql`, etc. | -| **Data cleaning & normalization** | dbt Staging models (L1) | `gd_dbt/models/staging/` | -| **Business logic** (what is a signup, what is a payout) | dbt Semantic models (L2) | `gd_dbt/models/semantic/` | -| **Dashboard-ready metrics** (daily counts, funnels, KPIs) | dbt Marts (L3) | `gd_dbt/models/marts/` | +| **Chain connection & raw capture** | Pipeline (TypeScript) | `pipeline-v5/src/`, reference in `pipeline-v5/README.md` | +| **Which contracts are captured** | dbt reference seed | `gd_dbt/seeds/contract_deployments.csv` | +| **Which chains are in scope** | Pipeline code | `pipeline-v5/src/control-plane/releaseScope.ts`, explained in `docs/release-scope.md` | +| **Raw table schemas (DDL)** | `warehouse/L1/` | Allowlisted migrations `08_` to `12_`, applied only through `scripts/deploy-warehouse.ps1` | +| **Data cleaning & normalization** | dbt Staging models | `gd_dbt/models/staging/` | +| **Business logic** (what is a signup, what is a payout) | dbt Semantic models | `gd_dbt/models/semantic/` | +| **Dashboard-ready metrics** (daily counts, funnels, KPIs) | dbt Marts | `gd_dbt/models/marts/` | | **Tests & data quality** | dbt schema YAML + custom tests | `gd_dbt/models/*/_*.yml` | | **Documentation** | dbt YAML (column/model descriptions) + `/docs` | Auto-published to [GitHub Pages](https://gooddollar.github.io/data-team/) | | **Business glossary & term definitions** | `docs/06_BUSINESS_GLOSSARY_AND_AI_DISAMBIGUATION.md` | Single source of truth for "what does X mean" | ### Adding a New Contract / Event -Short version (full guide in [`01_ARCHITECTURE.md`](docs/01_ARCHITECTURE.md)): - -1. Register the contract (ABI + config) in the pipeline -2. Create the L1 raw table DDL and declare it as a dbt source -3. Run the pipeline to backfill historical events -4. Add a dbt staging model (normalization) → optional semantic model (business logic) → optional mart (dashboard metrics) -5. Add glossary entries for any new terms +1. Add the contract to the `contract_deployments` seed (chain, proxy address, creation block, one + row per implementation era) and, so models can decode it, its events to the `event_surface` + seed. No pipeline code or table schema changes. Details: + [`pipeline-v5/README.md`](pipeline-v5/README.md#adding-a-contract). +2. Capture it with an explicit, authorized `backfill --from --to` run. Production ingestion is not + authorized yet; see [`docs/START_HERE.md`](docs/START_HERE.md). +3. Add a dbt staging model -> optional semantic model (business logic) -> optional mart (dashboard metrics). +4. Add glossary entries for any new terms. Each layer only reads from the layer directly below it. Logic flows up, never sideways. ### Operating the System -| Task | Command | When | +| Task | Command | Production status (2026-10-05) | |---|---|---| -| Ingest new events from chain | `cd pipeline-v5 && npx tsx src/index.ts daily` | Daily / on-demand | -| Check ingestion against the contracts | `cd pipeline-v5 && npx tsx src/index.ts verify` | After every ingest | -| Rebuild all warehouse layers | `cd gd_dbt && dbt run` | After ingestion | +| Preview exactly what an ingestion would do | `cd pipeline-v5 && npx tsx src/index.ts plan --chains=XDC --addresses=0x.. --from=N --to=N` | Safe: reads no chain, writes nothing | +| Run the pipeline's automated tests | `cd pipeline-v5 && npm test` | Safe: credential-free | +| Ingest, check, repair (`daily`, `backfill`, `coverage`, `verify`, `repair`) | `cd pipeline-v5 && npx tsx src/index.ts ...` | **Not authorized against production.** Needs the migrated schema and a writer identity | +| Rebuild warehouse layers | `cd gd_dbt && dbt run` | Default target writes `dev_sandbox` | | Run data quality tests | `cd gd_dbt && dbt test` | After any `dbt run` | | Browse data catalog + lineage | Visit [gooddollar.github.io/data-team](https://gooddollar.github.io/data-team/) | Anytime | | Add/modify a model | Edit SQL in `gd_dbt/models/`, run `dbt run --select model_name` | Development | -The pipeline and dbt are independent — the pipeline writes raw tables, dbt reads them. They run in sequence (ingest first, then transform), not as a single coupled process. +The pipeline and dbt are independent -- the pipeline writes raw tables, dbt reads them. They run in sequence (ingest first, then transform), not as a single coupled process. --- ## What's Live +Status as of 2026-10-05. Details and measured production state: [`docs/START_HERE.md`](docs/START_HERE.md#2-what-is-live-what-is-prepared-what-is-not-applied). + | Component | Status | Scope | |---|---|---| -| Ingestion pipeline | ✅ Live | XDC chain — UBIScheme + Invite contracts | -| dbt warehouse (10 models) | ✅ Live | 2 staging, 5 semantic, 3 marts | -| Looker Studio dashboards | ✅ Live | Invite funnel, daily metrics, claim activity | -| Pipeline hardening | ✅ Done | Idempotent MERGE write path, historical de-duplication, contract-oracle reconciliation, coverage ledger | -| Scheduled ingestion | 🔄 Next | The workflow exists and has never authenticated. See `pipeline-v5/README.md` | -| Multi-chain expansion | 📋 Planned | Celo, Ethereum | -| Self-service AI | 📋 Planned | Post-pipeline hardening | +| dbt warehouse | Live | Staging, Semantic and Marts over the older XDC per-contract tables (`ClaimContractEvents`, `InviteContractEvents`) | +| Looker Studio dashboards | Live (not re-verified on 2026-10-05) | Invite funnel, daily metrics, claim activity, read from the Marts | +| Older XDC raw tables | Present, not refreshed | Loaded by an earlier pipeline version. The current pipeline does not write them and no job refreshes them | +| Ingestion pipeline (`pipeline-v5`) | Merged, **not run in production** | Writes `RawLogs` and `Transactions`. Tested with the automated suite and small sandbox runs only | +| Raw-table migrations `08` to `12` | Merged and sandbox-rehearsed, **not applied** | Scope and retention approved; waiting on a temporary administrator-provisioned commissioning identity | +| Staging dataset and pipeline writer identity | **Not created** | Required before any production ingestion | +| Production canary / 12-month Celo and XDC ingestion | **Not run, not authorized** | Each needs its own approval after the schema is commissioned and verified | +| Scheduled ingestion | Off | `.github/workflows/pipeline-daily.yml` is manual-only and refuses production datasets, but is not ready to dispatch: it still expects `GCP_SA_KEY`, whose absence was last measured 2026-09-28. Reconcile it with the keyless identity design before use | +| dbt models over `RawLogs`, dashboard cutover | Not built | | +| Release scope | Frozen in code | Celo, XDC, Ethereum. Fuse dropped 2026-09-28. Base and Gnosis not assessed | +| Self-service AI | Planned | | --- @@ -127,44 +157,55 @@ The pipeline and dbt are independent — the pipeline writes raw tables, dbt rea | Path | What | |---|---| -| [`gd_dbt/`](gd_dbt/) | dbt project — all warehouse models, tests, docs, macros | -| [`pipeline-v5/`](pipeline-v5/) | The ingestion pipeline (TypeScript). The only one. Runbook in its own README | +| [`gd_dbt/`](gd_dbt/) | dbt project -- all warehouse models, tests, docs, macros, reference seeds | +| [`pipeline-v5/`](pipeline-v5/) | The ingestion pipeline (TypeScript). The only one. Reference in its own README | | [`warehouse/L1/`](warehouse/L1/) | Raw table DDL (pipeline-written tables, dbt *sources*) | | [`scripts/`](scripts/) | Explicitly allowlisted L1 migration helper and sandbox validator | | [`contracts/`](contracts/) | ABI files, deployment block numbers, contract reference | -| [`docs/`](docs/) | System documentation, data model, operations guide, governance | +| [`docs/`](docs/) | Start-here guide, system documentation, data model, operations guide, governance | --- ## Quick Start +Everything in this section is safe to run today: none of it writes to a production dataset. +Production schema changes and production ingestion need separately approved identities; see +[`docs/START_HERE.md`](docs/START_HERE.md#7-three-separate-operations-three-separate-approvals). + ### Prerequisites - Node.js LTS (v20+) - Google Cloud SDK (`gcloud`, `bq`) -- `gcloud auth application-default login` for metadata reads and labelled sandbox validation -- Production L1 DDL and ingestion use separately approved impersonated identities; see [`03_OPERATIONS.md`](docs/03_OPERATIONS.md) +- `gcloud auth application-default login` for metadata reads, dbt development, and labelled + sandbox validation. A personal login is **not** a production writer: as measured on + 2026-10-05 it can read `BlockchainEvents` but cannot create tables, change schemas, or write rows - Python 3.9+ with dbt-bigquery (`pip install dbt-bigquery`) ### Run the warehouse ```bash cd gd_dbt -dbt run # Build all: Staging → Semantic → Marts +dbt run # Build all: Staging -> Semantic -> Marts (default "dev" target writes dev_sandbox) dbt test # Run schema + data-quality tests dbt docs serve # Browse lineage + docs at localhost:8080 ``` -### Run the pipeline +First-time dbt setup is in [`docs/03_OPERATIONS.md`](docs/03_OPERATIONS.md). + +### Try the pipeline locally ```bash cd pipeline-v5 -cp .env.example .env # Add your ENVIO_API_TOKEN npm install -npx tsx src/index.ts daily # Ingest from chain into the BigQuery raw tables -npx tsx src/index.ts verify # Reconcile what was ingested against the contracts +npm test # credential-free automated test suite +cp .env.example .env # add your ENVIO_API_TOKEN; every mode needs it set +npx tsx src/index.ts plan --chains=XDC --addresses=0x.. --from=N --to=N # reads no chain, writes nothing ``` +Do not run `daily`, `backfill`, `verify`, `coverage`, `repair`, `dedup` or `calibrate` against +production until the schema is commissioned and a writer identity exists. Mode reference: +[`pipeline-v5/README.md`](pipeline-v5/README.md). + ### Inspect the prepared L1 migration (plan-only) ```powershell @@ -181,10 +222,13 @@ impersonation; see [`03_OPERATIONS.md`](docs/03_OPERATIONS.md). | Document | Purpose | |---|---| -| [`00_VISION.md`](docs/00_VISION.md) | Why this system exists and where it's going | -| [`01_ARCHITECTURE.md`](docs/01_ARCHITECTURE.md) | Layer responsibilities, naming, how to add contracts | -| [`02_DATA_MODEL.md`](docs/02_DATA_MODEL.md) | Column-level reference for every table | -| [`03_OPERATIONS.md`](docs/03_OPERATIONS.md) | How to run everything (setup, ingest, build, test) | +| [`START_HERE.md`](docs/START_HERE.md) | **Read first.** Current status, data flow, raw design, readers, safety controls, next steps, vocabulary | +| [`pipeline-v5/README.md`](pipeline-v5/README.md) | Pipeline modes, bookkeeping tables, failure handling | +| [`03_OPERATIONS.md`](docs/03_OPERATIONS.md) | How to run everything (setup, migrations, dbt) | +| [`release-scope.md`](docs/release-scope.md) | Which chains the release covers, and why | +| [`02_DATA_MODEL.md`](docs/02_DATA_MODEL.md) | Column-level reference for the older raw tables and the Semantic/Marts models | +| [`01_ARCHITECTURE.md`](docs/01_ARCHITECTURE.md) | Layer responsibilities and naming. Raw-layer section and contract-adding steps are outdated | +| [`00_VISION.md`](docs/00_VISION.md) | Historical MVP-era motivation; not current architecture or status | | [`04_CONTRACT_MECHANICS.md`](docs/04_CONTRACT_MECHANICS.md) | How GoodDollar smart contracts work and what events they emit | | [`05_ANALYTICS_DOCUMENTATION_CONTRACT.md`](docs/05_ANALYTICS_DOCUMENTATION_CONTRACT.md) | Required docs/tests/AI-readiness gates for new models | | [`06_BUSINESS_GLOSSARY_AND_AI_DISAMBIGUATION.md`](docs/06_BUSINESS_GLOSSARY_AND_AI_DISAMBIGUATION.md) | Business term definitions and AI clarification protocol | @@ -193,8 +237,12 @@ impersonation; see [`03_OPERATIONS.md`](docs/03_OPERATIONS.md). ## Tech Stack -- **Ingestion:** TypeScript, Envio HyperSync +- **Ingestion:** TypeScript, Envio HyperSync, JSON-RPC - **Warehouse:** dbt Core, Google BigQuery - **Dashboards:** Google Looker Studio - **Infrastructure:** GCP project `gooddollar` -- **Target chains:** XDC (live), Celo + Ethereum (planned) +- **Release scope:** Celo, XDC, Ethereum (no production ingestion yet; dashboards use older XDC data) + +--- + +*Docs current as of 2026-10-05 -- onchain-analytics@c34c4c1.* diff --git a/projects/onchain-analytics/docs/00_VISION.md b/projects/onchain-analytics/docs/00_VISION.md index c7ac98c..9da4ead 100644 --- a/projects/onchain-analytics/docs/00_VISION.md +++ b/projects/onchain-analytics/docs/00_VISION.md @@ -1,5 +1,10 @@ # Vision: GoodDollar Onchain Analytics Platform +> **Historical MVP-era vision.** This page records the original XDC invites proof of concept, not +> the current system design or production status. Its references to a decoded, per-contract raw +> layer, the old `pipeline/` folder, and Fuse support are outdated. For the current architecture, +> release scope, and verified status, start with [`START_HERE.md`](START_HERE.md). + ## The problem Today's analytics setup is a Google Apps Script glued to a Google Sheet. Every new question requires a custom scraper or a manual export. We can't cross-reference invite signups with claim activity. We can't cohort users by retention. We can't run a simple "how many invitees who signed up in March made it to 3 claims" query without writing code. diff --git a/projects/onchain-analytics/docs/01_ARCHITECTURE.md b/projects/onchain-analytics/docs/01_ARCHITECTURE.md index 0815a73..1fb516a 100644 --- a/projects/onchain-analytics/docs/01_ARCHITECTURE.md +++ b/projects/onchain-analytics/docs/01_ARCHITECTURE.md @@ -1,5 +1,13 @@ # Architecture +> **Partly outdated (2026-10-05).** The L1 (raw) section below describes the older per-contract +> tables, which current dbt models still read. The new raw design is one universal `RawLogs` table +> plus `Transactions`, written undecoded; see [`START_HERE.md`](START_HERE.md#4-the-raw-layer-design). +> Do not follow steps 2 to 4 of "How to add a new contract": `CONTRACTS` in `config.ts` and the +> `--contracts` option no longer exist. Use "Adding a contract" in +> [`pipeline-v5/README.md`](../pipeline-v5/README.md#adding-a-contract) instead. The L2/L3 sections +> remain accurate. + ## The three layers | Layer | BigQuery dataset | Owns | Cadence | Storage type | diff --git a/projects/onchain-analytics/docs/02_DATA_MODEL.md b/projects/onchain-analytics/docs/02_DATA_MODEL.md index 12ee4d1..e0de7ca 100644 --- a/projects/onchain-analytics/docs/02_DATA_MODEL.md +++ b/projects/onchain-analytics/docs/02_DATA_MODEL.md @@ -4,6 +4,11 @@ The canonical reference for every entity in the warehouse. If a column or busine For naming conventions, partitioning rules, and layer responsibilities see [`01_ARCHITECTURE.md`](01_ARCHITECTURE.md). +> **Scope note (2026-10-05).** The L1 tables below are the older per-contract tables that current +> dbt models and dashboards read. The new universal raw tables, `RawLogs` and `Transactions`, are +> not documented here yet; see [`START_HERE.md`](START_HERE.md#4-the-raw-layer-design) and the +> migration files in [`warehouse/L1/`](../warehouse/L1/). + --- ## L1 — `gooddollar.BlockchainEvents.*` diff --git a/projects/onchain-analytics/docs/03_OPERATIONS.md b/projects/onchain-analytics/docs/03_OPERATIONS.md index 14d4242..9db463d 100644 --- a/projects/onchain-analytics/docs/03_OPERATIONS.md +++ b/projects/onchain-analytics/docs/03_OPERATIONS.md @@ -2,6 +2,11 @@ How to run everything in this repo. Written for someone who has never used BigQuery before. +> **Status, 2026-10-05.** The raw-table migrations in this guide are approved but **not applied +> to production**, and the pipeline has **not been run against production**. Commands that write +> to `BlockchainEvents` are documented for when that is authorized; until then, use the plan-only +> and sandbox paths. Read [`START_HERE.md`](START_HERE.md) for current status and the next step. + --- ## One-time setup (do these before anything else) @@ -37,6 +42,11 @@ gcloud config set project gooddollar The pipeline and `bq` CLI both read these credentials automatically — no passwords stored anywhere in this repo. +A personal login is enough for metadata reads, dbt development in `dev_sandbox`, and the labelled +sandbox validator. It is **not** a production writer: on 2026-10-05 the account that prepared this +release could read `BlockchainEvents` but could not create tables, change schemas, or write rows. +Production schema changes and ingestion use separately approved service-account identities. + ### 4. Install pipeline dependencies ``` @@ -60,36 +70,51 @@ The pipeline is [`pipeline-v5/`](../pipeline-v5/), and it is the only one. Its f [`pipeline-v5/README.md`](../pipeline-v5/README.md); this section is the short version. All commands run from inside `pipeline-v5/`. -### Backfill, load full history +### Preview a run (safe now) + +``` +cd pipeline-v5 +npx tsx src/index.ts plan --chains=XDC --addresses=0x.. --from=N --to=N +``` + +`plan` reads no chain and writes nothing to BigQuery. It lists the exact work units and the budget +verdict, and exits nonzero if anything would be refused. + +### Backfill a named range (needs production authorization) ``` cd pipeline-v5 -npx tsx src/index.ts backfill --contracts=ClaimContractEvents -npx tsx src/index.ts backfill --contracts=InviteContractEvents +npx tsx src/index.ts backfill --chains=XDC --addresses=0x.. --from=N --to=N ``` -Add `--from=N --to=N` to target a range. A run reports its chunk plan and, for every range it -attempted, writes a row to `BlockchainEvents.IngestionCoverage` recording whether every chunk -succeeded. +`backfill` requires both `--from` and `--to`; a bare `backfill` is refused. There is no +`--contracts` option. Always pass `--chains`, because the default selection includes a chain +outside the release scope and the run then cannot exit 0. A run reports its chunk plan and, for +every range it attempted, writes a row to `BlockchainEvents.IngestionCoverage` recording whether +every chunk succeeded. **Re-running the same range is safe and is expected.** The write path is a staging table plus a -`MERGE` on `(network, tx_hash, log_index)`, so a repeated backfill leaves the table -byte-identical. This was not true of the predecessor, which appended through streaming inserts -whose `insertId` de-duplication window is minutes rather than months; re-running a range four -months later wrote 43,000 phantom rows. If you read that older instruction anywhere, it is -wrong. +`MERGE` on `(chain_id, tx_hash, log_index)`, so a repeated backfill leaves one row per log. +This was not true of the predecessor, which appended through streaming inserts whose `insertId` +de-duplication window is minutes rather than months. -### Daily incremental +### Daily incremental (needs production authorization) ``` cd pipeline-v5 -npx tsx src/index.ts daily -npx tsx src/index.ts verify +npx tsx src/index.ts daily --chains=CELO,XDC +npx tsx src/index.ts coverage --chains=CELO,XDC +npx tsx src/index.ts verify --chains=CELO,XDC ``` -`verify` reconciles the warehouse against the contracts' own per-day ledgers and is the only -check here that consults something outside the warehouse. A run that does not reconcile exits -nonzero. +`daily` resumes each contract from its coverage record. While that record is empty, every contract +would resume from its creation block, so the run-size and span limits refuse it; capture history +with explicit `backfill` ranges first. `verify` reconciles the warehouse against the contracts' own +per-day ledgers and is the only check here that consults something outside the warehouse. A run +that does not reconcile exits nonzero. + +Every mode except `plan` starts by checking the bookkeeping tables and records a `PipelineRuns` +row, so even `verify` and `coverage` need the migrated schema and write access. --- @@ -111,9 +136,12 @@ Before any production change, validate the same migration files against a fresh From `pipeline-v5/`: ```powershell -node --import tsx ..\scripts\ops\validate-l0-migrations.mjs ..\..\_scratch\unit-07a-commissioning\sandbox-validation.json +node --import tsx ..\scripts\ops\validate-l0-migrations.mjs ``` +The report is written to `_scratch/schema-migration-validation.json` at the repository root; pass a +different path as the first argument to change it. + This sandbox check exercises the old `PipelineRuns` and `OracleReconciliation` shapes, repeats the migrations, checks their statement types and byte caps, verifies historical fixture rows remain, and proves the sandbox is absent after cleanup. It does not write production tables or ingest chain @@ -190,7 +218,7 @@ model/column docs with `dbt docs serve` (opens ). | `SCHEMA_MISMATCH: has no column(s) …` | The live schema differs from the runtime contract | Stop ingestion. Re-measure the schema and approve a new additive migration; do not recreate the table | | Run exits 1 with skipped chunks | HyperSync rate limiting or a timeout | Read `IngestionCoverage` for the exact ranges, then `backfill --from --to` over them | | `UNCONFIRMED EMPTY RANGE` | A range came back empty and no independent endpoint could confirm it | Not an error to clear by retrying. The watermark deliberately did not advance. Re-run when the endpoints recover | -| `REORG SUSPECTED` | An existing key now sits under a different block hash | Delete and re-ingest that block range | +| `REORG SUSPECTED` / `REORG_APPLIED` | An existing key now sits under a different block hash | No manual action. The MERGE has already rewritten the row whole, including its block facts, and the coverage row records it | | `Unrecognized name` during `dbt run` | A Semantic model references an L1 column that does not exist | Check the L1 schema matches `02_DATA_MODEL.md` | | Mart numbers look wrong | Marts rebuilt before L1 was fully ingested | Run `verify` first. If it reports short days, `repair --days=…`, then `cd gd_dbt && dbt run --select marts` | @@ -198,8 +226,13 @@ model/column docs with `dbt docs serve` (opens ). ## Cron / daily automation (post-MVP) -There is **no daily job set up yet** — the pipeline and dbt are run manually. When it's time to -automate, the daily flow is two ordered steps: ingest first, then dbt. +There is **no daily job set up yet** — the pipeline and dbt are run manually. The GitHub workflow +`.github/workflows/pipeline-daily.yml` runs only on manual dispatch and refuses production +datasets. It is not ready to dispatch: it still expects a `GCP_SA_KEY` secret, whose absence was +last measured on 2026-09-28. Reconcile the workflow with the keyless identity design and verify its +authentication before using it. Do not add a long-lived service-account key for convenience, and +do not schedule ingestion until production ingestion is authorized. When it is time to automate, +the daily flow is two ordered steps: ingest first, then dbt. The examples below are illustrative. **Linux/macOS:** diff --git a/projects/onchain-analytics/docs/START_HERE.md b/projects/onchain-analytics/docs/START_HERE.md new file mode 100644 index 0000000..4f45be8 --- /dev/null +++ b/projects/onchain-analytics/docs/START_HERE.md @@ -0,0 +1,421 @@ +# Start here: GoodDollar onchain analytics + +This guide is for engineers who are new to this system. It covers what the system does, what is +running today, how data moves through it, and what has to happen next. Read it before the +detailed references. It also tells you which of those references are current. + +Status in this guide is as of **2026-10-05**, code at `master` commit `c34c4c1`. + +Statements below fall into three kinds, labelled where it matters: + +- **Measured**: observed directly in production metadata or in a run, with the date. +- **Design**: what the merged code does. Tested by automated tests and in disposable sandbox + datasets, not yet in production. +- **Planned / unverified**: intended next steps, or things nobody has checked yet. + +--- + +## 1. What the system is for + +GoodDollar runs smart contracts on several blockchains. Those contracts emit events, for example +"a user claimed UBI", "an invitee joined", or "a bounty was paid". This system copies those +events into Google BigQuery, turns them into business definitions (what counts as a claim, what +counts as a referral signup), and serves them to dashboards. The goal is to answer questions about +onchain GoodDollar activity with SQL against tested, documented tables, without one-off scripts. + +--- + +## 2. What is live, what is prepared, what is not applied + +| Item | State | Notes | +| - | - | - | +| dbt models and Looker Studio dashboards for XDC invites and claims | **Live (existing)** | They read the older per-contract tables `ClaimContractEvents` and `InviteContractEvents`. An earlier pipeline version loaded those tables. The current pipeline does not write them and nothing refreshes them on a schedule, so check their latest `block_timestamp` before relying on recency. Their live status comes from earlier project documentation and was not re-checked for this guide. | +| Ingestion pipeline, [`pipeline-v5/`](../pipeline-v5/) | **Merged, not run in production** | Writes the new raw tables. It has run only over small block ranges against sandbox datasets. | +| Five additive schema migrations, [`warehouse/L1/`](../warehouse/L1/) files `08` to `12` | **Merged and rehearsed in a sandbox. Not applied to production.** | Scope and retention are approved. Applying them is blocked on access; see sections 7 and 8. | +| Staging dataset `BlockchainEvents_Staging` | **Not created** | Must exist before any production ingestion. | +| Pipeline writer identity and its permissions | **Not provisioned** | | +| Production ingestion (canary or history) | **Not run, not authorized** | | +| dbt models over the new raw tables, and dashboard cutover | **Not built, not done** | Current models still read the older tables. | +| Scheduled ingestion | **Off** | The GitHub workflow is manual-only and refuses production dataset names. Its authentication does not match the keyless identity design; reconcile it before use. | + +The new raw schema has not been commissioned or populated in production. The existing XDC models +and dashboards still use the older per-contract tables; check their freshness before relying on +current data. + +--- + +## 3. How data flows, and who owns each step + +``` + Blockchains (Celo, XDC, Ethereum) + | + | pipeline-v5 (TypeScript; reads via HyperSync or JSON-RPC) + v + BigQuery raw layer: gooddollar.BlockchainEvents + RawLogs, Transactions undecoded chain data + IngestionCoverage, PipelineRuns, bookkeeping: what was read, by which run, + OracleReconciliation and how it compared with the contracts + | + | dbt (gd_dbt/, SQL models) + v + Staging -> Semantic -> Marts cleaning, business definitions, dashboard tables + | + v + Looker Studio dashboards +``` + +| Step | Owned by | Lives in | What it writes | +| - | - | - | - | +| Read the chain | Pipeline | `pipeline-v5/src/` | Nothing by itself | +| Raw capture | Pipeline | `pipeline-v5/src/`; the contract list is the seed `gd_dbt/seeds/contract_deployments.csv` | Raw and bookkeeping tables in `BlockchainEvents` | +| Raw table definitions | Migration files, applied by an administrator | `warehouse/L1/`, `scripts/deploy-warehouse.ps1` | Table and view definitions only, never rows | +| Decode, clean, define business terms | dbt | `gd_dbt/models/` | `Staging`, `Semantic`, `Marts` (or `dev_sandbox` by default) | +| Dashboards | Looker Studio | Outside this repository | Nothing | + +Ground rules: + +- Only the pipeline talks to the chain. +- The pipeline stores nothing decoded. Turning logs into business events is dbt's job. +- dbt reads raw tables and never writes them. +- Each layer reads only from the layer below it. + +A naming note: the raw layer is called **L0** in pipeline code and **L1** in older documents and in +the `warehouse/L1/` folder name. Both mean the `BlockchainEvents` dataset. + +--- + +## 4. The raw layer design + +### `RawLogs`: one row per log, every contract, undecoded + +Every log a captured contract emits becomes one row. A row stores the chain, block, and +transaction it came from, the emitting contract address, all four topic slots, and the data +section **verbatim**. The pipeline does not try to match the log to an expected event. + +- Merge key: `(chain_id, tx_hash, log_index)`. One log, one row. +- Partitioned by month of `block_timestamp`, with **partition filter required**. Every query + against the table must filter on `block_timestamp`. For whole-history reads, use the + `RawLogsAllHistory` view, which carries a wide filter in its definition. +- Column definitions: [`warehouse/L1/09_CreateRawLogs_v1.sql`](../warehouse/L1/09_CreateRawLogs_v1.sql). + +### `Transactions`: one row per transaction that produced a captured log + +- Merge key: `(chain_id, tx_hash)`. +- A reverted transaction emits no logs, so it never appears here. Absence from this table means + "produced no captured log", not "did not happen". +- `TransactionsAllHistory` is its whole-history view. + +### Provenance on every row + +Every raw row records where it came from: `source_kind` and `source_id` (which reader answered), +`assurance` (how independently confirmed the read was; see section 5), `capture_id` and +`ingestion_run_id` (which capture and run wrote it), `ingested_at`, and the contract era fields +(`implementation_address`, `era_index`, `era_resolution`). If the era cannot be resolved, the row +says `unresolved` and does not guess. + +### The bookkeeping tables + +| Table | Answers | +| - | - | +| `IngestionCoverage` | Which reader read which block range of which contract, what it found, what failed, and how far the result can be trusted. The pipeline writes one row for every attempted range, including failures and refusals. **The next run resumes from this table**, not from the highest block number in the data. A block range with no coverage row was never read. | +| `PipelineRuns` | One row per command, with its exit code and outcome counts | +| `OracleReconciliation` | Per protocol day: what the contract's own ledger says against what the warehouse holds | + +### Why raw data stays undecoded + +- A decoding mistake is fixed by changing a dbt model and rebuilding. Nobody has to re-read the + chain. +- A new event on an existing contract needs no schema change. +- No log is lost because nobody had defined a column for it. + +The cost: a raw row means nothing on its own until a model matches it against the event +reference seed, [`gd_dbt/seeds/event_surface.csv`](../gd_dbt/seeds/event_surface.csv). + +### Why raw data is preserved rather than replaced + +- The pipeline never drops or truncates raw tables. Its writer identity is designed to have no + permission to delete tables in the raw dataset. +- Reading the same log twice updates its single row instead of adding a second one. A matched + row is rewritten whole, so its block facts stay correct after a chain reorganisation. +- `dedup` is the only mode that removes rows, and it removes only repeated keys. +- Schema migrations are additive: create-if-absent, or add nullable columns. Existing rows remain, + and older rows read `NULL` in the new columns. +- Retention (approved): no automatic expiration on raw or bookkeeping tables. Only temporary + tables in `BlockchainEvents_Staging` expire, after six hours, once that dataset exists. + +--- + +## 5. Readers: how the pipeline reads each chain + +- **HyperSync** is Envio's indexed blockchain data service. It is fast and needs an + `ENVIO_API_TOKEN`. It is the primary reader wherever an index exists for the chain. +- **JSON-RPC** is the standard API that blockchain nodes expose. The pipeline uses it to read + chains that have no HyperSync index, to check empty ranges, and to read contract state for + `verify`. + +### Release scope (Design, enforced in code) + +The chain list is frozen in `RELEASE_SCOPE_FREEZE`, in +`pipeline-v5/src/control-plane/releaseScope.ts`. The decision is explained in +[`release-scope.md`](release-scope.md). + +| Chain | Chain id | Status | Reader in code | +| - | - | - | - | +| Celo | 42220 | In release | HyperSync; RPC for empty-range checks | +| XDC | 50 | In release | HyperSync; RPC for empty-range checks | +| Ethereum | 1 | Declared in release scope; no ingestion planned for this release | JSON-RPC only, because no HyperSync index resolves for it | +| Fuse | 122 | Dropped 2026-09-28 | The pipeline refuses it | +| Base, Gnosis | n/a | Not assessed (no queries were made against either chain) | None | + +Fuse remains in the network configuration, and the default chain selection includes it. A run +without `--chains` therefore reports Fuse as unsupported and cannot exit 0. Always pass `--chains`. + +### Empty results are not proof of absence + +A reader that returns no logs for a range has not proved that nothing happened there. Public +endpoints can return an empty answer without raising an error. **Measured:** an identical repeated +log query returned zero 7 times in 10 on one day and 3 times in 10 on the day before. One Celo +endpoint returned false zeros 20 to 90 percent of the time, depending on how old the range was. + +What the pipeline does (Design): + +1. If HyperSync reports that it stopped early in a chunk, the chunk counts as a failure and is + retried. It is never treated as empty. +2. If a chunk has no logs for a contract, the pipeline asks **every configured RPC endpoint** for + that contract and range. The range is split to each chain's per-request limit: 1,000 blocks on + XDC and 5,000 on Celo. +3. Then one of three things happens: + - **An endpoint finds a log.** The empty result is refuted. The capture is not recorded as + complete, and a later run reads the range again. The logs RPC found are not written by this + check. + - **At least two endpoints cover the whole range without error and find nothing.** The + emptiness is *corroborated* and coverage may advance past it. This is still not proof. + - **Fewer than two endpoints do so.** The range is recorded `unconfirmed_empty` and coverage + does not advance. + +Limitations of the RPC check: + +- It can refute an empty result, but it cannot prove one. +- It is expensive on sparse ranges. Calls per empty chunk = sub-ranges x endpoints. With + defaults, an empty 20,000-block XDC chunk costs up to 60 calls, and Celo costs up to 12. +- On Ethereum, the check queries the same endpoints that did the read. It repeats the read; it + does not add an independent source. +- Setting `CONFIRM_EMPTY_CHUNKS=false` turns the check off. Leave it on. +- Reading over RPC (Ethereum) costs one extra call per block and two per transaction to fetch + block and transaction details. Public endpoint rate limits make wide ranges slow. + +### Assurance grades + +Each capture is graded by how many independent sources agree on it: + +- **A**: two independent sources each enumerated the range and returned identical results. +- **B**: one source enumerated the range, and a second confirmed its emptiness. +- **C**: one source only. + +A complete HyperSync capture is normally **C**, because one index is one source. An Ethereum RPC +capture where two endpoints agree can be **A**. The grade measures independent confirmation, not +reader quality. + +--- + +## 6. Safety and correctness controls + +All of these are **Design**: they exist in merged code and are exercised by the automated test +suite and by small sandbox runs. + +| Control | What it does | +| - | - | +| Bounded chunks | Ranges are read in fixed block chunks: 20,000 blocks by default on Celo and XDC, 50,000 on Ethereum. A chunk either completes or is recorded as skipped. | +| Block-boundary writes | Buffers are flushed only between blocks, so a write never contains part of a block. | +| Stage plus MERGE | Rows load into a temporary table in `BlockchainEvents_Staging`, then merge into the target on its key. Re-running a range leaves one row per key. Each MERGE is limited to a literal time window derived from the rows, padded one month on each side. | +| One writer at a time | A file lease blocks a second process on the same host from writing to the same table. A waiting writer gives up after 15 minutes, exits nonzero, and records the range as incomplete. A dead same-host process is detected and its lease reclaimed immediately; otherwise any lease older than the one-hour stale timeout can be reclaimed by age. **Limit:** this is not a cross-machine lock, and the full production workload has not tested the one-hour timeout boundary. | +| Explicit coverage | Resume points come from `IngestionCoverage`. Gaps, refusals, and unreadable ranges are recorded, never inferred from the data. | +| Retries and deadlines | Each HyperSync chunk runs in a child process that is killed after 120 seconds. One contract's whole range is abandoned after 90 minutes, and the blocks not attempted are recorded. Every RPC call has a 30-second deadline. BigQuery and HyperSync calls retry up to 5 times with backoff, and requests are paced. | +| Finality | Ingestion stays behind the chain tip: 15 blocks on XDC, 1,930 blocks (about 32 minutes) on Celo, 94 blocks on Ethereum. A range the reader still holds as reversible is recorded `rollback_eligible` and read again later. | +| Cost and size budgets | Every BigQuery job carries a 10 GiB `maximumBytesBilled` cap. A run may attempt at most 12 contracts. One capture may span 30 days of blocks unless the range is named with `--from` and `--to`, and named ranges are capped at 100,000,000 blocks. A bare `backfill` and a one-sided range are both refused. | +| Release scope | Dropped or undeclared chains are refused on every path that reads or writes. | +| Startup schema check | Every mode except `plan` checks the bookkeeping tables' columns first and stops with `SCHEMA_MISMATCH` if the live tables are behind the code. | +| Exit statuses | `0`: at least one unit completed and none was refused, unsupported, incomplete, or failed. `1`: partial, or a read-only check found something. `2`: nothing completed, the scope was empty, the run was refused, or the arguments were wrong. `PipelineRuns` records the same outcome. Any nonzero exit alerts Slack if a webhook is configured. | +| Data verification | `verify` reconciles warehouse counts and amounts, per protocol day, against the contracts' own ledgers. Ledgers exist only for the UBIScheme on Celo and XDC and the Invites contract on XDC. `coverage` lists unread ranges and logs that no event definition matches. `calibrate` measures each source's miss rate. dbt has uniqueness tests on both raw merge keys; they are disabled until the tables exist and are enabled with `--vars '{l0_v4_tables_exist: true}'`. | + +**What this testing does and does not show.** These controls passed the credential-free test +suite, small-range runs against sandbox datasets, and a labelled-sandbox rehearsal of the five +migrations. That rehearsal ran each migration twice, with no change on the second run, and kept +all historical fixture rows. None of this proves a full-history or 12-month production run, the +production permission setup, or production cost and throughput at scale. Those are only shown by +running in production, starting with a canary. + +--- + +## 7. Three separate operations, three separate approvals + +| | Schema commissioning | Data ingestion | dbt transformation | +| - | - | - | - | +| What it does | Creates or alters raw table and view definitions | Reads chains and writes raw rows and bookkeeping | Builds `Staging`, `Semantic`, and `Marts` from raw tables | +| Tool | `scripts/deploy-warehouse.ps1` with one named migration from `warehouse/L1/` | `pipeline-v5` | `gd_dbt` | +| Identity | A temporary schema-commissioner service account, impersonated by an administrator | A separate pipeline writer identity | The credentials in your dbt profile | +| Writes rows? | No | Yes (raw dataset) | Yes (derived datasets only) | +| State on 2026-10-05 | Approved, not applied | Not run, not authorized | Existing models run over the older tables | + +Approving one of these does not authorize the next. + +### Production access roles + +- **Personal login (Application Default Credentials).** This is read access, not a production + writer. Use it for local development and read-only checks, not as the production writer. Verify + effective permissions in the target environment; production commissioning and ingestion use + separately provisioned identities. +- **Schema commissioner (proposed, not provisioned).** A dedicated service account with no + downloadable key. On the project it may only run jobs. On `BlockchainEvents` it may create, + list, get, and read tables and update table definitions (required by `ALTER TABLE`). It cannot + write rows, delete anything, or change access. It is granted for a single window of at most 30 + minutes, then revoked, and the revocation is verified. +- **Pipeline writer (proposed, separate).** On the raw dataset it may create tables, read, and + write rows, but not delete tables or the dataset. On `BlockchainEvents_Staging` it may also + delete tables. That is why staging has its own dataset. +- **Administrator.** Someone who can create service accounts, add and remove role bindings, grant + temporary impersonation (Service Account Token Creator), and create the staging dataset. The + administrator who grants access also removes it. Owning the `BlockchainEvents` dataset lets an + owner change that dataset's access list. It does not by itself include the project-level + permissions to create service accounts or grant impersonation, so confirm those separately. + +No service-account keys are used. The migration helper passes the impersonated identity only to +the `bq` process it starts, and never changes persistent `gcloud` settings. + +--- + +## 8. The safe next step (as of 2026-10-05) + +1. **Commission the approved schema through a temporary, narrow identity.** An administrator + provisions the schema commissioner and grants themselves temporary impersonation on it. They + then apply the five migrations one at a time, in this order: `09_CreateRawLogs_v1.sql`, + `08_PipelineRunsOutcome_v1.sql`, `10_AddOracleReconciliationCompatibility_v1.sql`, + `11_CreateRawLogsAllHistory_v1.sql`, `12_CreateTransactionsAllHistory_v1.sql`. Finally, they + remove the grants. The exact command and stop conditions are in + [`03_OPERATIONS.md`](03_OPERATIONS.md#l1-raw-tables----allowlisted-additive-migrations). +2. **Verify, live.** Check each of the following: + - **Schemas.** `RawLogs` has 25 columns, monthly partitioning, and a required partition filter. + `PipelineRuns` has 41 columns. `OracleReconciliation` has nullable `chain_id` and + `contract_address` and keeps `table_id`. Both all-history views exist. + - **Records.** Row counts are unchanged: `PipelineRuns` 22, `OracleReconciliation` 265, + `IngestionCoverage` 10, `Transactions` 0. No expiration is set on permanent tables. + - **Permissions.** A fresh effective-permission check shows the commissioner's mutation + permissions denied after revocation. +3. **Only then, authorize a bounded canary.** A canary is a deliberately small first production + ingestion. Before it runs: + - Create `BlockchainEvents_Staging` with a six-hour default table expiration. + - Provision the pipeline writer identity. + - Run `plan` for an explicit range. + + Then run `backfill` with `--chains`, `--addresses`, `--from`, and `--to`, followed by + `coverage` and `verify`. Finally, check the resulting `PipelineRuns` and `IngestionCoverage` + rows. The canary's contracts and block range have not been chosen yet. + +**What schema approval does not authorize:** the full 12-month Celo and XDC ingestion, scheduling, +or moving dbt models and dashboards onto the new tables. Each one is a separate decision, made +after the canary result is reviewed. + +--- + +## 9. What you can safely run today + +None of these commands writes to a production dataset. + +```bash +cd pipeline-v5 +npm install +npm test # credential-free automated test suite +cp .env.example .env # every mode needs ENVIO_API_TOKEN set, even plan +npx tsx src/index.ts plan --chains=XDC --addresses= --from= --to= +``` + +`plan` reads no chain and writes nothing to BigQuery. It lists the exact work and checks it +against the budgets. + +```powershell +# from projects/onchain-analytics/ +.\scripts\deploy-warehouse.ps1 -Migration 09_CreateRawLogs_v1.sql # prints the target only +``` + +```bash +cd gd_dbt +dbt run # default "dev" target writes to the dev_sandbox dataset, not production +dbt test +``` + +dbt setup is in [`03_OPERATIONS.md`](03_OPERATIONS.md#staging-semantic-marts--dbt). + +**Do not run against production** until the steps in section 8 are complete: `daily`, +`backfill`, `repair`, `dedup`, `verify`, `coverage`, or `calibrate`. Each of these starts with the +bookkeeping-table check and records a `PipelineRuns` row, so each needs the migrated schema and a +writer identity. + +--- + +## 10. Finding your way around + +| Path | What is there | +| - | - | +| [`pipeline-v5/`](../pipeline-v5/) | The ingestion pipeline. Entry point `src/index.ts`. Readers: `src/reader.ts`, `src/hypersync.ts`, `src/rpc.ts`. Orchestration: `src/pipeline.ts`. Resume and gaps: `src/coverage.ts`, `src/repair.ts`. Exit codes: `src/outcome.ts`. Chains and settings: `src/config.ts`. | +| [`warehouse/L1/`](../warehouse/L1/) | Raw table DDL. Files `08` to `12` are the allowlisted migrations. `06_L0Contract_v4.sql` is a design reference; **do not run it**. | +| [`scripts/`](../scripts/) | `deploy-warehouse.ps1`, the migration helper. `ops/validate-l0-migrations.mjs`, the sandbox rehearsal. | +| [`gd_dbt/`](../gd_dbt/) | dbt project: models, tests, and the reference seeds (contracts, events, chains, tokens). | +| [`contracts/`](../contracts/) | ABIs and contract reference material | +| [`docs/`](.) | This guide and the references below | + +### Which documents are current + +| Document | Status | +| - | - | +| This guide | Current as of 2026-10-05 | +| [`pipeline-v5/README.md`](../pipeline-v5/README.md) | Current. Detailed reference for modes, bookkeeping tables, and failure handling | +| [`03_OPERATIONS.md`](03_OPERATIONS.md) | Current for setup, migrations, and dbt. Pipeline commands corrected to match the code | +| [`release-scope.md`](release-scope.md) | Current | +| [`02_DATA_MODEL.md`](02_DATA_MODEL.md) | Describes the older per-contract tables and the `Semantic` and `Marts` models that dashboards use today. Does not describe `RawLogs` or `Transactions`. | +| [`01_ARCHITECTURE.md`](01_ARCHITECTURE.md) | **Partly outdated.** Its raw-layer section describes the older per-contract design. Its "how to add a contract" steps use options that no longer exist; do not follow them. | +| [`00_VISION.md`](00_VISION.md) | Historical MVP-era motivation; not current architecture or status. | +| `04_CONTRACT_MECHANICS.md`, `05_ANALYTICS_DOCUMENTATION_CONTRACT.md`, `06_BUSINESS_GLOSSARY_AND_AI_DISAMBIGUATION.md` | Contract behaviour, model documentation rules, and business terms. Not tied to the raw-layer design. | +| [dbt docs site](https://gooddollar.github.io/data-team/) | Model and column lineage for the dbt project | + +If an older document contradicts this guide or the code, the code is the authority. Do not run a +pipeline command from an older document unless it matches the usage text at the top of +`pipeline-v5/src/index.ts` or the pipeline README. + +--- + +## 11. Vocabulary + +| Term | Meaning | +| - | - | +| Block | A batch of transactions the chain adds at one time. Has a number, a hash, and a timestamp. | +| Block range | An inclusive span of block numbers, for example `--from=100 --to=199`. The unit of reading and of coverage. | +| Chain tip | The newest block a chain has produced. | +| Finality | The point after which a block will not be replaced. The pipeline stays a fixed number of blocks behind the tip. | +| Reorganisation (reorg) | The chain replacing recent blocks with different ones. A log can move to another block. | +| Log / event | A record a contract emits during a transaction, such as `UBIClaimed`. "Event" is the contract's definition; "log" is one occurrence of it. | +| Topic, `topic0` | Indexed fields of a log. `topic0` is normally the event's signature hash and identifies which event it is. | +| `log_data` | The non-indexed part of a log, stored as hex. | +| Era | A period during which a proxy contract pointed to one implementation. Upgrades start a new era. | +| HyperSync | Envio's indexed service for reading chain data quickly. Primary reader where an index exists. | +| RPC (JSON-RPC) | The standard request API of a blockchain node. `eth_getLogs` is its log query. | +| False zero | An empty answer from an endpoint for a range that actually contains logs, returned without an error. | +| Raw layer | The `BlockchainEvents` dataset: undecoded chain data plus bookkeeping. Called L0 in code, L1 in older docs. | +| Merge key | The columns that identify one row, used by `MERGE` to update rather than duplicate. | +| MERGE | A BigQuery statement that inserts new keys and updates existing ones in one step. | +| Staging dataset | `BlockchainEvents_Staging`: temporary tables used for each write, separate from production so the writer never needs delete rights there. Unrelated to dbt's `Staging` layer. | +| Partition filter | A required `WHERE` condition on `block_timestamp` that stops a query from scanning the whole table. | +| Capture | One reader reading one or more contracts over one block range in one run. | +| Coverage, coverage record | `IngestionCoverage`: which ranges were read, how, and with what result. A range with no row was never read. | +| Coverage frontier | The edge of the last clean capture. `daily` resumes from here. | +| Assurance | The A/B/C grade for how independently a capture was confirmed. | +| Oracle | A contract's own public per-day ledger, used by `verify` as an external check. | +| Protocol day | The contract's own day boundary. For the UBIScheme it starts at 12:00 UTC, not midnight. | +| dbt | The SQL transformation tool that builds `Staging`, `Semantic`, and `Marts` from raw tables. | +| Staging / Semantic / Marts | dbt layers: light cleaning; business definitions; dashboard-shaped tables. | +| ADC | Application Default Credentials: the personal Google login that local tools use. | +| Service account | A non-human Google identity with its own permissions. | +| Impersonation | Acting as a service account through short-lived tokens, without a key file. | +| Migration | One additive DDL file that changes a raw table or view definition. | +| Commissioning | Applying approved migrations to production. Changes definitions, writes no rows. | +| Canary | A deliberately small first production ingestion, used to prove the write path and permissions before anything larger. | +| Release scope | The frozen list of chains this release may read: Celo, XDC, Ethereum. | diff --git a/projects/onchain-analytics/docs/release-scope.md b/projects/onchain-analytics/docs/release-scope.md index 4cf2d96..729e679 100644 --- a/projects/onchain-analytics/docs/release-scope.md +++ b/projects/onchain-analytics/docs/release-scope.md @@ -8,7 +8,7 @@ without something failing. | Chain | In the release | Ingested in this release | Note | | - | - | - | - | -| Celo | Yes | Yes | Primary chain for claims | +| Celo | Yes | No, not yet in the new raw pipeline | Primary chain for claims; production ingestion has not run | | XDC | Yes | Yes | Holds the existing invites and claims dataset | | Ethereum | Yes | No | Declared, deliberately not ingested here. Reserve paused since 2023-12-17, and no model reads an Ethereum contract today | | **Fuse** | **No, dropped 2026-09-28** | No | See below | @@ -23,7 +23,8 @@ can be declared in scope and still not be captured yet; Ethereum is exactly that Fuse data is disproportionately expensive to obtain, and the cost was holding up everything else. Specifically: -- There is no HyperSync index for Fuse, so it is the only chain that needs a raw RPC adapter. +- There is no HyperSync index for Fuse, so every Fuse read would go through the slower JSON-RPC + reader. - The one archive-capable Fuse endpoint this project could find is rate limited to roughly 54 reads per hour, and both enumerating readers cap at 20,000 logs per response. - Fuse prunes its transaction index, so a receipt lookup cannot establish a contract's creation @@ -36,8 +37,8 @@ serve it to the same standard as the rest, and the decision was to ship the rest **Fuse can be added back as its own piece of work.** Nothing has been thrown away: Fuse history is permanently readable from the chain itself, and the inventory this project already measured is -still in the repository (see the next section). Bringing it back means building the RPC reader and -re-running ingestion, not rediscovering what is there. +still in the repository (see the next section). Bringing it back means returning it to the release +scope and running ingestion over the existing RPC reader, not rediscovering what is there. ## What "dropped" means for data that is already recorded diff --git a/projects/onchain-analytics/pipeline-v5/README.md b/projects/onchain-analytics/pipeline-v5/README.md index 9dfbb85..035d24c 100644 --- a/projects/onchain-analytics/pipeline-v5/README.md +++ b/projects/onchain-analytics/pipeline-v5/README.md @@ -4,6 +4,10 @@ This is the **only** pipeline in this repository. It reads contract logs from En or over JSON-RPC on the chains no index serves, and writes them into BigQuery **undecoded**, then checks its own work against the contracts. +> **Status, 2026-10-05:** not yet run against production. The raw-table migrations it depends on +> are approved but not applied, and there is no writer identity or staging dataset yet. Start with +> [`../docs/START_HERE.md`](../docs/START_HERE.md) for current status and the next step. + Two older versions existed, `pipeline/` and `future/pipeline-v4/`, and both are gone. What they could do that this one could not has been ported, and what they did that was wrong is the reason several of the rules below exist. @@ -13,7 +17,8 @@ several of the rules below exist. ## Nothing is decoded here, and that is the point The warehouse holds one row per log entry, with all four topic slots and the data blob stored -verbatim, for every contract on Celo, XDC, Fuse and Ethereum. It does not hold a column per event +verbatim, for every contract on the chains in the release scope (Celo, XDC and Ethereum; see +[`../docs/release-scope.md`](../docs/release-scope.md)). It does not hold a column per event field, and it does not try to match a log against an expected event at ingestion time. That is a deliberate trade with a stated cost. A row carries no meaning until a model joins it to @@ -86,6 +91,9 @@ cp .env.example .env # then fill in ENVIO_API_TOKEN gcloud auth application-default login ``` +A personal login is enough for `plan`, the tests, and sandbox work. It is not a production writer; +production ingestion runs under a separately approved writer identity. + The BigQuery tables are created by allowlisted migrations in `../warehouse/L1/`, not by the pipeline. The multi-statement `06_L0Contract_v4.sql` is a reference, not a deployment command. Before production commissioning, run the labelled-sandbox validator described in @@ -128,16 +136,20 @@ never pass through this process and are bounded instead by the size of the table ### One writer at a time -A cross-process lease excludes a second writer on the same host from merging into the same table. +A cross-process lease excludes a second writer on the same host from writing to the same table. Two concurrent captures of one range previously produced two rows per key while both processes exited zero and both recorded the range as completely captured. A writer that cannot take the lease within `WRITE_LOCK_WAIT_MS` is refused, exits nonzero, and records the range as incomplete -so the next run reads it again. A lease whose holder is no longer running is reclaimed at once; -one from another host expires after `WRITE_LOCK_STALE_MS`. +so the next run reads it again. A dead same-host process is detected and its lease reclaimed +immediately. A lease older than `WRITE_LOCK_STALE_MS` (60 minutes by default) can also be +reclaimed by age, even if a same-host process still exists; another host's process cannot be +checked directly. The production workload has not tested this timeout boundary, so a run that +reaches it should be treated as a safety incident and checked before retrying. This file lock +does not coordinate writers on separate hosts. --- -## The seven modes +## The eight modes ``` npx tsx src/index.ts [options] @@ -145,15 +157,25 @@ npx tsx src/index.ts [options] | Mode | What it does | Writes | | - | - | - | +| `plan` | Says exactly what a run would do and checks it against the budgets. Reads no chain | Nothing | | `daily` | Ingests from each contract's coverage frontier to the chain tip, minus the finality margin | `RawLogs`, `Transactions` | -| `backfill` | Ingests a named range, or each contract's whole history from its creation block | `RawLogs`, `Transactions` | +| `backfill` | Ingests a named range. `--from` and `--to` are both required | `RawLogs`, `Transactions` | | `verify` | Reconciles the warehouse against a contract's own ledger, per protocol day | Nothing except a reconciliation record | | `coverage` | Reports every block range not covered by a clean capture, and every contract with no capture at all | Nothing | | `dedup` | Collapses repeated natural keys, in place | `RawLogs`, `Transactions` | | `repair` | Re-reads every range the coverage ledger records as not covered, then re-checks | `RawLogs`, `Transactions` | | `calibrate` | Repeats one identical query against every source and reports each one's miss rate | Nothing | -Options: `--chains=A,B`, `--addresses=0x..`, `--from=N --to=N`, `--days=N,N`, `--dry-run`. +Every mode except `plan` also checks the bookkeeping tables at startup and records a +`PipelineRuns` row. + +Options: `--chains=A,B`, `--addresses=0x..`, `--from=N --to=N`, `--days=N,N`, `--dry-run`, +`--max-capture-blocks=N`, `--max-captures=N`. + +Budgets: a run may attempt at most 12 contracts and one capture at most 30 days of blocks, unless +raised with the last two options or, for the span, named with `--from` and `--to` (capped at +100,000,000 blocks). A one-sided range is refused. Always pass `--chains`: the default selection +includes Fuse, which is outside the release scope and is reported as unsupported. Exit codes: `0` everything attempted completed and reconciled, `1` partial, `2` nothing succeeded or the arguments were wrong. Every mode fails closed. A run that skipped a chunk, or @@ -170,10 +192,10 @@ npx tsx src/index.ts coverage npx tsx src/index.ts verify ``` -`daily` is what the scheduled workflow runs. `coverage` says what the warehouse does and does not +`daily` is the incremental run. `coverage` says what the warehouse does and does not cover, which is the question an empty query result cannot answer on its own. `verify` is cheap relative to being wrong, and it is the only check here that consults something outside the -warehouse. +warehouse. Pass `--chains` to each. ## Adding a contract @@ -200,7 +222,7 @@ what is missing and claims no coverage, so an empty result over it can be read c | Symptom | What to run | | - | - | | `coverage` reports open gaps | `repair`, which re-reads exactly those ranges and re-checks the ledger afterwards | -| `coverage` reports a contract with no capture at all | `backfill --addresses=0x..`, because no range has ever been read for it | +| `coverage` reports a contract with no capture at all | `backfill --chains=.. --addresses=0x.. --from=N --to=N`, because no range has ever been read for it | | `verify` reports days as `missing` | `repair`, then `verify --days=...` | | `verify` reports days as `duplicated` | `dedup --dry-run`, then `dedup` | | `verify` reports days as `amount_unreadable` | A value in `log_data` exceeded what the decoder can represent. It is reported rather than summed as a zero | @@ -230,31 +252,30 @@ rows are left where they are; this program simply no longer adds to them. ## Scheduling -`.github/workflows/pipeline-daily.yml`, 01:00 UTC daily, plus manual dispatch. - -**It has never successfully run.** `PipelineRuns` holds eight rows, every one of them from a -laptop, none from a runner, across the seven weeks since the workflow was merged. The pipeline -log shows the GCP credential setup failing four different ways on 2026-08-18. Before relying on -the schedule, dispatch it manually and confirm a row appears in `PipelineRuns` with a runner -hostname. +There is no scheduled ingestion. `.github/workflows/pipeline-daily.yml` runs only on manual +dispatch, requires a target dataset and a typed confirmation, and refuses `BlockchainEvents`, +`Staging`, `Semantic` and `Marts`. It is for sandbox runs only, but is not ready to dispatch: it +still expects a `GCP_SA_KEY` secret, whose absence was last measured on 2026-09-28. Reconcile its +authentication with the keyless identity design before use. Do not add a long-lived service-account +key for convenience. Scheduling production ingestion is a separate decision that follows a +successful production canary. --- ## Known limits, stated rather than discovered later -- **No historical backfill has been run.** The pipeline has been exercised over small ranges on - XDC and Ethereum against a sandbox dataset. The full history is a separate, larger piece of - work with its own cost and access decisions. -- **Two of the four chains have no index.** `fuse.hypersync.xyz` and `eth.hypersync.xyz` do not - resolve, so those chains are enumerated over JSON-RPC, which costs two calls per transaction - and one per block on top of the log query. One explorer in that set publishes a quota of ten - reads per eleven minutes; hydration rotates across the available endpoints and backs off hard - on a rate limit, but a wide range there is slow by construction. +- **No historical backfill has been run, and nothing has been run against production.** The + pipeline has been exercised over small ranges on XDC and Ethereum against a sandbox dataset. + The full history is a separate, larger piece of work with its own cost and access decisions. +- **Ethereum has no index.** `eth.hypersync.xyz` does not resolve (nor does `fuse.hypersync.xyz`, + for the dropped Fuse chain), so Ethereum is enumerated over JSON-RPC, which costs two calls per + transaction and one per block on top of the log query. Hydration rotates across the available + endpoints and backs off hard on a rate limit, but a wide range there is slow by construction. - **Assurance grade A needs two independent readers that both answer.** A complete capture from a single index earns C, because one source is one source however good it is. That is the - definition rather than a judgement on the reader, and it means the two chains with no index can - reach a higher grade than the two with one. Raising the index chains to B means confirming - non-empty ranges against a second source, which is real cost and is not done today. + definition rather than a judgement on the reader, and it means Ethereum, with no index, can + reach a higher grade than Celo and XDC, which have one. Raising the index chains to B means + confirming non-empty ranges against a second source, which is real cost and is not done today. - **`era_resolution` is `era_map_lookup` or `unresolved`, never `slot_read_at_block`.** The era is resolved from the reference seed's block ranges rather than by reading the proxy slot at each row's own block, which would be one archive call per log. A block the seed does not cover @@ -265,7 +286,9 @@ hostname. - **Empty-chunk confirmation is expensive on sparse ranges.** Each empty chunk costs one JSON-RPC query per endpoint per sub-range. `CONFIRM_EMPTY_CHUNKS=false` turns it off, and turning it off is how a false zero becomes a permanent gap. -- **A `daily` run over a contract with no coverage is a full backfill of that contract.** The +- **A `daily` run over a contract with no coverage would be a full backfill of that contract.** The coverage ledger is the only thing that can say a range was read, so a contract with no coverage - row resumes from its creation block. That is deliberate: Inferring coverage from the rows that - happen to be present is the defect this pipeline was rewritten to remove. + row resumes from its creation block, and the run-size and span budgets then refuse the run + rather than start a history read nobody asked for. Capture history with explicit `backfill` + ranges. Inferring coverage from the rows that happen to be present is the defect this pipeline + was rewritten to remove.