# Data dictionary — `dataset.csv` (N=6 rows, 20 columns)

Companion to `dataset.csv` at
`docs/research/erc-8004-four-chain-census-2026-09/dataset.csv`. Every column, in
schema-emission order, its type, its units, its null semantics, and the
`PROVENANCE.json` key the value traces to. This file is the reader's map for
turning the 6 CSV rows back into the freeze envelopes on disk.

## Row identity

One row per **measurement slice**, not per agent, not per chain. The six slices
are three whole-history censuses (Ethereum, Base, Arbitrum), one whole-history
raw-count backfill (BSC standard registry), two disjoint cluster samples of the
BSC standard registry (early era and late era), and one probe of the BRC8004
BNB-team fork. Row key is the `chain` column value.

## Columns (schema-emission order)

| # | Column | Type | Units | Null semantics | Provenance |
|--:|---|---|---|---|---|
| 1 | `chain` | string | identifier | Always populated. One of `ethereum`, `base`, `arbitrum`, `bnb_standard`, `bnb_standard_late`, `bnb_brc8004_fork`. Row key. | Assigned by builder. |
| 2 | `registry_address` | string | hex-address | Always populated. `0x8004A169FB4a3325136EB29fA0ceB6D2e539a432` for the standard ERC-8004 rows (Ethereum, Base, Arbitrum, both BSC standard rows); `0xfA09B3397fAC75424422C4D28b1729E3D4f659D7` for the BRC8004 BNB-team fork. | Constant per-row from spec / probe envelope. |
| 3 | `method` | enum (string) | measurement method | Always populated. One of `full_rescan` (Ethereum / Base / Arbitrum: every `Transfer(from=0x0)` since deployment plus every `tx.from` resolved), `full_raw_count_backfill` (BSC standard row: raw count is whole-history, sender columns are early-era-sampled lower bounds), `cluster_sample_uniform_tuple` (BSC early sample, K=1500 uniform-in-tuple), `cluster_sample_windowed` (BSC late sample and BRC8004 probe, K=60 / K=40 non-overlapping 2,000-block windows). | Assigned by row. |
| 4 | `window_start_block` | integer | block-number | Always populated. Deployment block for census rows; low bound of the sample window for BSC-late and BRC8004 rows; deployment block for the BSC-standard row (its window covers the full history). | `PROVENANCE.json` entries with `column: window_start_block` (source keys `s3_method.first_block`, `deployment_block`, `late_range[0]`). |
| 5 | `window_end_block` | integer | block-number | Always populated. Checkpoint block for census rows; head-checkpoint block for the BSC standard row; high bound of the sample window for BSC-late and BRC8004 rows. | `PROVENANCE.json` entries with `column: window_end_block` (source keys `s3_method.last_block`, `head_checkpoint`, `late_range[1]`, `head_checkpoint_block`). |
| 6 | `window_start_utc` | string | ISO8601 UTC (may be empty) | Empty on rows where the block→UTC lookup was not captured in the freeze envelope. Empty on `bnb_standard`, `bnb_standard_late`, and `bnb_brc8004_fork` (the BSC freeze envelopes did not record UTC anchors); populated on Ethereum, Base, Arbitrum. | Read verbatim from the 065[456] freeze envelope's `window_start_utc` field where present. |
| 7 | `window_end_utc` | string | ISO8601 UTC (may be empty) | Empty on the same three BSC rows for the same reason; populated on Ethereum, Base, Arbitrum. | Read verbatim from the 065[456] freeze envelope's `window_end_utc` field where present. |
| 8 | `registrations` | integer | count | Always populated. Raw count of `Transfer(from=0x0)` events in the window. **Never confuse with `distinct_senders`.** For `bnb_standard` this is the whole-history raw count (352,415 = 352,359 backfill + 56 delta) even though the sender columns on the same row are early-era-sampled — this asymmetry is the reason the `notes` column carries `raw_count_is_whole_history`. For `bnb_brc8004_fork` this is the count observed inside the K=40 sample and is 0 (see § "Row-specific reading notes"). | `PROVENANCE.json` entries with `column: registrations` (source keys `s2_reconciliation.raw_registration_count`, `raw_registrations`, `total_late_tuples_pooled`, `tuples_pooled`). |
| 9 | `distinct_senders` | integer | count | Always populated. Distinct-sender count for the slice. **Interpret with column 10.** For the three census rows this is the exact distinct-sender count over the full raw-count set (100% `tx.from` coverage). For the two BSC sample rows and the BRC8004 row this is a lower bound from a cluster sample — see column 10 and the `notes` column. | `PROVENANCE.json` entries with `column: distinct_senders` (source keys `s2_reconciliation.distinct_registrant_deduped`, `distinct_senders_in_sample_lower_bound`). |
| 10 | `distinct_senders_is_lower_bound` | boolean | boolean (`true`/`false`) | Always populated. **This column enforces trap 2 by schema.** `false` on the three census rows (Ethereum, Base, Arbitrum — 100% `tx.from` coverage means the value in column 9 is the true count). `true` on `bnb_standard` (the sender count is early-era-sampled, not whole-history), `bnb_standard_late` (windowed sample), and `bnb_brc8004_fork` (windowed probe). A downstream parser that sees `true` here must treat column 9 as `≥`, not `=`, and must not sum it with any other row's column 9. | Assigned by row per method. |
| 11 | `top1_share` | float | 0–1 fraction | Always populated. Share of registrations in the window attributable to the single largest sender. For census rows this is a point measurement; for sample rows read together with columns 13/14 (Wilson CI95). Zero on `bnb_brc8004_fork` (0 registrations in the sample). | `PROVENANCE.json` entries with `column: top1_share` (source keys `s2_reconciliation.top1_sender_share`, `top1_sender_share`). |
| 12 | `top10_share` | float | 0–1 fraction | Always populated. Share of registrations attributable to the ten largest senders. Same read semantics as column 11. Zero on `bnb_brc8004_fork`. | `PROVENANCE.json` entries with `column: top10_share` (source keys `s2_reconciliation.top10_sender_share`, `top10_sender_share`). |
| 13 | `top1_share_ci95_low` | float \| empty | 0–1 fraction | **Empty on census rows (Ethereum, Base, Arbitrum) — the census is a point measurement, so a confidence interval is not defined.** Populated on the two BSC sample rows and the BRC8004 row with the Wilson CI95 lower bound on the top-1 share. On BRC8004 it is 0 (0-of-0 sample). | `PROVENANCE.json` entries with `column: top1_share_ci95_low` (source key `top1_sender_share_ci95[0]`). |
| 14 | `top1_share_ci95_high` | float \| empty | 0–1 fraction | Same rule as column 13 — empty on census rows, populated on sample / probe rows with the Wilson CI95 upper bound. | `PROVENANCE.json` entries with `column: top1_share_ci95_high` (source key `top1_sender_share_ci95[1]`). |
| 15 | `top10_share_ci95_low` | float \| empty | 0–1 fraction | Same rule as column 13 — empty on census rows, populated on sample / probe rows with the Wilson CI95 lower bound on the top-10 share. | `PROVENANCE.json` entries with `column: top10_share_ci95_low` (source key `top10_sender_share_ci95[0]`). |
| 16 | `top10_share_ci95_high` | float \| empty | 0–1 fraction | Same rule as column 13 — empty on census rows, populated on sample / probe rows with the Wilson CI95 upper bound on the top-10 share. | `PROVENANCE.json` entries with `column: top10_share_ci95_high` (source key `top10_sender_share_ci95[1]`). |
| 17 | `hhi` | float \| empty | 0–1 scale | Empty where not defensible. Empty on the three census rows because their freeze envelopes recorded the top-1 and top-10 aggregates but did not retain the full sender-count distribution needed to compute HHI over all senders. Empty on `bnb_brc8004_fork` (0 registrations). Populated on `bnb_standard` (0.010379) and `bnb_standard_late` (0.000963). | `PROVENANCE.json` entries with `column: hhi` (source key `hhi_0_1_scale`). |
| 18 | `sample_seed` | integer \| empty | integer | **Empty on census rows** (no sample was drawn). Populated on the three sampled rows with the exact seed used to draw the sample: `bnb_standard` = 659, `bnb_standard_late` = 660, `bnb_brc8004_fork` = 661. Named seeds make the sampled rows reproducible; a third party running the same window with the same seed will land the same tuples. | `PROVENANCE.json` entries with `column: sample_seed` (source key `seed`). |
| 19 | `n_resolved` | integer | count | Always populated. Distinct transactions whose `tx.from` was resolved into the sender pool for this row. Equals `registrations` on the three census rows (100% coverage). On sample rows this is the sample size, not the total: 1,500 tuples in the BSC early sample, 1,200 in the BSC late sample, 0 in the BRC8004 probe. **This is the number that anchors the Wilson CI95 columns.** | `PROVENANCE.json` entries with `column: n_resolved` (source keys `s2_reconciliation.raw_registration_count`, `n_resolved_tuples`, `n_resolved`). |
| 20 | `notes` | string | text | Short caveat pointers, semicolon-delimited. Empty on the two "clean" census rows (Ethereum, Base). Populated on rows with a reader-facing caveat: `small_n_caveat` (Arbitrum), `early_era_only;cluster_sample_uniform_tuple;raw_count_is_whole_history` (BSC standard), `late_era_only;cluster_sample_windowed` (BSC late), `fork_effectively_empty` (BRC8004). | Assigned by row per method. |

## Why the CI95 columns are null on census rows

Wilson CI95 is a confidence interval on a **binomial share estimate from a
sample**. The three census rows (Ethereum, Base, Arbitrum) each resolved
`tx.from` for 100% of the raw registrations in the window — there is no
sampling error to bound, so a CI95 is undefined. The two BSC sample rows and
the BRC8004 probe row draw shares from cluster samples, so a Wilson CI95
applies and is reported.

Downstream users must not fill in the empty CI95 cells on census rows with
zeros or with `top1_share` itself — those cells are structurally undefined, not
"tight to zero width." A tool that needs a CI on a census row should treat it
as an exact measurement (width 0 conceptually, but the recorded cells stay
empty to prevent that from being mistaken for a computed statistic).

## Row-specific reading notes

- **`bnb_standard`** — `registrations = 352,415` is the whole-history raw count
  (352,359 from the backfill in `.0658-bnb-standard-census.json` plus 56 from
  the reconciliation delta in `.0659-bnb-standard-reconcile.json`). All the
  sender-side columns on this row (`distinct_senders`, `top1_share`,
  `top10_share`, the four CI95 columns, `hhi`, `n_resolved`) are drawn from the
  **early-era** cluster sample in `.0659c-bnb-sender-sample.json` and are
  bounded to the block range 79,027,268 – 104,927,268. The row is deliberately
  asymmetric to preserve the whole-history raw count in the same table as the
  three censuses while making the sender-side lower bound explicit in the
  `notes` column.
- **`bnb_standard_late`** — this is a **separate row** for the late-era cluster
  sample only. It does not restate the raw count from `bnb_standard`; instead
  its `registrations` field carries the raw count observed inside the K=60
  sampled 2,000-block windows within the late range (`total_late_tuples_pooled`
  in the source envelope), which is a lower bound on the late-range total, not
  the total itself. Splitting the two rows enforces trap 2 by schema.
- **`bnb_brc8004_fork`** — the BRC8004 BNB-team fork. Same event signature and
  same on-chain surface as the standard ERC-8004 registry but a **different
  contract address**. The K=40 sample of 2,000-block windows across the full
  deployment-to-head range observed 0 mints, so every count and share column
  reads 0 (including the Wilson CI95 columns) and `hhi` is empty. Nine mints
  were observed by task 0661 outside the K=40 sample (in the first 50,000
  blocks after deployment); the row's `notes` column carries
  `fork_effectively_empty`. Do not read this row as "BRC8004 has zero
  registrations" — read it as "the K=40 sample of BRC8004's full block range
  observed zero registrations," which the `notes` column names.

## Companion files

- `dataset.csv` — the 6 rows described here
- `dataset.json` — the same 6 rows re-emitted as a nested JSON object plus the
  attribution and offer blocks
- `dataset.jsonld` — schema.org `Dataset` metadata block (also inlined into
  `index.html`)
- `PROVENANCE.json` — one entry per numeric CSV cell, tracing to source key +
  source envelope path (this file is what makes every value in the CSV
  auditable)
- `SHA256SUMS` — sha256 of the four data files
- `METHOD.md` — how each row was produced, per-row window / RPC / seed /
  reconciliation history
- `GRADE.md` — the self-grade against the seven claim traps

## License

CC BY 4.0 — see `METHOD.md` § "License".
