Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 4 additions & 20 deletions .shipgate-allow
Original file line number Diff line number Diff line change
Expand Up @@ -11,8 +11,7 @@
# "was this actually checked?".

# --- contract registries: public infrastructure, not per-account data --------
# These list GoodDollar/Uniswap/Ubeswap contract addresses. Confirmed 2026-09-23
# that none of them contain user wallets.
# These files list deployed contract addresses, not user wallet balances.
docs/subgraph/*.txt :: content:address-list :: subgraph manifests, contract addresses and start blocks only
projects/onchain-analytics/contracts/ABIs/*.csv :: content:address-list :: contract ABI export, deployed contract addresses only
projects/onchain-analytics/contracts/README.md :: content:address-list :: the contract registry doc, by definition a list of contracts
Expand All @@ -22,25 +21,10 @@ queries/dune/reserve-analysis/lp-v3-positions.sql :: content:address-list :: Uni
projects/onchain-analytics/gd_dbt/seeds/contract_deployments.csv :: content:holder-table :: contract address paired with its deployment BLOCK, not a holder balance
projects/onchain-analytics/gd_dbt/seeds/tokens.csv :: content:holder-table :: token contract address paired with its DECIMALS, not a holder balance
# event_surface.csv is the event ABI catalogue: one row per (contract, era, event).
# Checked 2026-09-28 before waiving, not assumed. All 347 distinct addresses in the
# file are contract addresses already present in contract_deployments.csv above,
# including the 11 that appear outside the two address columns (all in abi_source,
# all of the form bytecode_equality_with_celo_0x..., naming a registry contract).
# Its only numeric columns are chain_id, era_index, indexed_positions and the two
# era block bounds. No balance column and no wallet address exists in it.
# Receipt: specs/_scratch/coord-unit-01/out/01-event-surface-address-provenance.json
# It contains contract addresses and event/era metadata, not holder balances.
projects/onchain-analytics/gd_dbt/seeds/event_surface.csv :: content:holder-table :: contract address paired with its era BLOCK BOUNDS, not a holder balance
# era_intervals.csv and era_boundary_evidence.csv are the era validity-interval table
# and the evidence behind it: which implementation was live between which blocks, and
# how that was established. Checked 2026-09-28 before waiving, not assumed.
# era_intervals.csv 262 rows, 278 distinct addresses, ALL already present in
# contract_deployments.csv above. Numeric columns are chain_id,
# valid_from_block, valid_to_block, era_index.
# era_boundary_evidence.csv 41 rows, 44 distinct addresses, ALL already present above.
# Numeric columns are block numbers plus COUNTS of verification
# checkpoints, events and chunks.
# Neither file has an amount column, a balance, or a wallet address.
# Receipt: specs/_scratch/coord-unit-02-gate/verify-new-seeds.mjs
# These files describe contract implementation eras and the block evidence for their boundaries.
# Their address-and-number pairs are deployment metadata, not per-account holdings.
projects/onchain-analytics/gd_dbt/seeds/era_intervals.csv :: content:holder-table :: contract address paired with its era BLOCK BOUNDS, not a holder balance
projects/onchain-analytics/gd_dbt/seeds/era_boundary_evidence.csv :: content:holder-table :: contract address paired with block numbers and verification COUNTS, not a holder balance

Expand Down
188 changes: 118 additions & 70 deletions projects/onchain-analytics/README.md

Large diffs are not rendered by default.

5 changes: 5 additions & 0 deletions projects/onchain-analytics/docs/00_VISION.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,10 @@
# Vision: GoodDollar Onchain Analytics Platform

> **Historical MVP-era vision.** This page records the original XDC invites proof of concept, not
> the current system design or production status. Its references to a decoded, per-contract raw
> layer, the old `pipeline/` folder, and Fuse support are outdated. For the current architecture,
> release scope, and verified status, start with [`START_HERE.md`](START_HERE.md).

## The problem

Today's analytics setup is a Google Apps Script glued to a Google Sheet. Every new question requires a custom scraper or a manual export. We can't cross-reference invite signups with claim activity. We can't cohort users by retention. We can't run a simple "how many invitees who signed up in March made it to 3 claims" query without writing code.
Expand Down
8 changes: 8 additions & 0 deletions projects/onchain-analytics/docs/01_ARCHITECTURE.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,13 @@
# Architecture

> **Partly outdated (2026-10-05).** The L1 (raw) section below describes the older per-contract
> tables, which current dbt models still read. The new raw design is one universal `RawLogs` table
> plus `Transactions`, written undecoded; see [`START_HERE.md`](START_HERE.md#4-the-raw-layer-design).
> Do not follow steps 2 to 4 of "How to add a new contract": `CONTRACTS` in `config.ts` and the
> `--contracts` option no longer exist. Use "Adding a contract" in
> [`pipeline-v5/README.md`](../pipeline-v5/README.md#adding-a-contract) instead. The L2/L3 sections
> remain accurate.

## The three layers

| Layer | BigQuery dataset | Owns | Cadence | Storage type |
Expand Down
5 changes: 5 additions & 0 deletions projects/onchain-analytics/docs/02_DATA_MODEL.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,11 @@ The canonical reference for every entity in the warehouse. If a column or busine

For naming conventions, partitioning rules, and layer responsibilities see [`01_ARCHITECTURE.md`](01_ARCHITECTURE.md).

> **Scope note (2026-10-05).** The L1 tables below are the older per-contract tables that current
> dbt models and dashboards read. The new universal raw tables, `RawLogs` and `Transactions`, are
> not documented here yet; see [`START_HERE.md`](START_HERE.md#4-the-raw-layer-design) and the
> migration files in [`warehouse/L1/`](../warehouse/L1/).

---

## L1 — `gooddollar.BlockchainEvents.*`
Expand Down
75 changes: 54 additions & 21 deletions projects/onchain-analytics/docs/03_OPERATIONS.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,11 @@

How to run everything in this repo. Written for someone who has never used BigQuery before.

> **Status, 2026-10-05.** The raw-table migrations in this guide are approved but **not applied
> to production**, and the pipeline has **not been run against production**. Commands that write
> to `BlockchainEvents` are documented for when that is authorized; until then, use the plan-only
> and sandbox paths. Read [`START_HERE.md`](START_HERE.md) for current status and the next step.

---

## One-time setup (do these before anything else)
Expand Down Expand Up @@ -37,6 +42,11 @@ gcloud config set project gooddollar

The pipeline and `bq` CLI both read these credentials automatically — no passwords stored anywhere in this repo.

A personal login is enough for metadata reads, dbt development in `dev_sandbox`, and the labelled
sandbox validator. It is **not** a production writer: on 2026-10-05 the account that prepared this
release could read `BlockchainEvents` but could not create tables, change schemas, or write rows.
Production schema changes and ingestion use separately approved service-account identities.

### 4. Install pipeline dependencies

```
Expand All @@ -60,36 +70,51 @@ The pipeline is [`pipeline-v5/`](../pipeline-v5/), and it is the only one. Its f
[`pipeline-v5/README.md`](../pipeline-v5/README.md); this section is the short version. All
commands run from inside `pipeline-v5/`.

### Backfill, load full history
### Preview a run (safe now)

```
cd pipeline-v5
npx tsx src/index.ts plan --chains=XDC --addresses=0x.. --from=N --to=N
```

`plan` reads no chain and writes nothing to BigQuery. It lists the exact work units and the budget
verdict, and exits nonzero if anything would be refused.

### Backfill a named range (needs production authorization)

```
cd pipeline-v5
npx tsx src/index.ts backfill --contracts=ClaimContractEvents
npx tsx src/index.ts backfill --contracts=InviteContractEvents
npx tsx src/index.ts backfill --chains=XDC --addresses=0x.. --from=N --to=N
```

Add `--from=N --to=N` to target a range. A run reports its chunk plan and, for every range it
attempted, writes a row to `BlockchainEvents.IngestionCoverage` recording whether every chunk
succeeded.
`backfill` requires both `--from` and `--to`; a bare `backfill` is refused. There is no
`--contracts` option. Always pass `--chains`, because the default selection includes a chain
outside the release scope and the run then cannot exit 0. A run reports its chunk plan and, for
every range it attempted, writes a row to `BlockchainEvents.IngestionCoverage` recording whether
every chunk succeeded.

**Re-running the same range is safe and is expected.** The write path is a staging table plus a
`MERGE` on `(network, tx_hash, log_index)`, so a repeated backfill leaves the table
byte-identical. This was not true of the predecessor, which appended through streaming inserts
whose `insertId` de-duplication window is minutes rather than months; re-running a range four
months later wrote 43,000 phantom rows. If you read that older instruction anywhere, it is
wrong.
`MERGE` on `(chain_id, tx_hash, log_index)`, so a repeated backfill leaves one row per log.
This was not true of the predecessor, which appended through streaming inserts whose `insertId`
de-duplication window is minutes rather than months.

### Daily incremental
### Daily incremental (needs production authorization)

```
cd pipeline-v5
npx tsx src/index.ts daily
npx tsx src/index.ts verify
npx tsx src/index.ts daily --chains=CELO,XDC
npx tsx src/index.ts coverage --chains=CELO,XDC
npx tsx src/index.ts verify --chains=CELO,XDC
```

`verify` reconciles the warehouse against the contracts' own per-day ledgers and is the only
check here that consults something outside the warehouse. A run that does not reconcile exits
nonzero.
`daily` resumes each contract from its coverage record. While that record is empty, every contract
would resume from its creation block, so the run-size and span limits refuse it; capture history
with explicit `backfill` ranges first. `verify` reconciles the warehouse against the contracts' own
per-day ledgers and is the only check here that consults something outside the warehouse. A run
that does not reconcile exits nonzero.

Every mode except `plan` starts by checking the bookkeeping tables and records a `PipelineRuns`
row, so even `verify` and `coverage` need the migrated schema and write access.

---

Expand All @@ -111,9 +136,12 @@ Before any production change, validate the same migration files against a fresh
From `pipeline-v5/`:

```powershell
node --import tsx ..\scripts\ops\validate-l0-migrations.mjs ..\..\_scratch\unit-07a-commissioning\sandbox-validation.json
node --import tsx ..\scripts\ops\validate-l0-migrations.mjs
```

The report is written to `_scratch/schema-migration-validation.json` at the repository root; pass a
different path as the first argument to change it.

This sandbox check exercises the old `PipelineRuns` and `OracleReconciliation` shapes, repeats the
migrations, checks their statement types and byte caps, verifies historical fixture rows remain,
and proves the sandbox is absent after cleanup. It does not write production tables or ingest chain
Expand Down Expand Up @@ -190,16 +218,21 @@ model/column docs with `dbt docs serve` (opens <http://localhost:8080>).
| `SCHEMA_MISMATCH: <table> has no column(s) …` | The live schema differs from the runtime contract | Stop ingestion. Re-measure the schema and approve a new additive migration; do not recreate the table |
| Run exits 1 with skipped chunks | HyperSync rate limiting or a timeout | Read `IngestionCoverage` for the exact ranges, then `backfill --from --to` over them |
| `UNCONFIRMED EMPTY RANGE` | A range came back empty and no independent endpoint could confirm it | Not an error to clear by retrying. The watermark deliberately did not advance. Re-run when the endpoints recover |
| `REORG SUSPECTED` | An existing key now sits under a different block hash | Delete and re-ingest that block range |
| `REORG SUSPECTED` / `REORG_APPLIED` | An existing key now sits under a different block hash | No manual action. The MERGE has already rewritten the row whole, including its block facts, and the coverage row records it |
| `Unrecognized name` during `dbt run` | A Semantic model references an L1 column that does not exist | Check the L1 schema matches `02_DATA_MODEL.md` |
| Mart numbers look wrong | Marts rebuilt before L1 was fully ingested | Run `verify` first. If it reports short days, `repair --days=…`, then `cd gd_dbt && dbt run --select marts` |

---

## Cron / daily automation (post-MVP)

There is **no daily job set up yet** — the pipeline and dbt are run manually. When it's time to
automate, the daily flow is two ordered steps: ingest first, then dbt.
There is **no daily job set up yet** — the pipeline and dbt are run manually. The GitHub workflow
`.github/workflows/pipeline-daily.yml` runs only on manual dispatch and refuses production
datasets. It is not ready to dispatch: it still expects a `GCP_SA_KEY` secret, whose absence was
last measured on 2026-09-28. Reconcile the workflow with the keyless identity design and verify its
authentication before using it. Do not add a long-lived service-account key for convenience, and
do not schedule ingestion until production ingestion is authorized. When it is time to automate,
the daily flow is two ordered steps: ingest first, then dbt. The examples below are illustrative.

**Linux/macOS:**

Expand Down
Loading
Loading