# GRADE — 2026-09-17 four-chain census

**Verdict:** the dataset supports the specific windowed claims below. It does
**not** support the broader claims listed under each trap. The seven claim
traps below are enforced in the data (via the `distinct_senders_is_lower_bound`
column, separate BSC rows, empty CI95 cells on census rows) AND named here in
prose so a reader who never opens the CSV still carries them.

## What this dataset does support

- **A four-chain raw-registration total on the ERC-8004 IdentityRegistry as of
  2026-09-17,** anchored to per-chain checkpoint blocks (Ethereum 25,994,481
  at 03:20:47Z UTC; Base 51,415,064 at 04:31:15Z UTC; Arbitrum 505,982,474 at
  04:31:28Z UTC; BSC 122,349,961). Every row is anchored to a `(chain,
  window_start_block, window_end_block)` triple, not to a clock time.
- **An exact distinct-sender count on the three census chains** — Ethereum
  11,352 / 50,820, Base 29,289 / 88,500, Arbitrum 142 / 1,502 — with 100%
  `tx.from` coverage and 0 fails.
- **A defensible statement that early-era top-1 and top-10 sender shares on
  BSC are materially higher than late-era shares**, backed by two disjoint
  cluster samples with named seeds (659, 660) and non-overlapping Wilson
  CI95s. `top1` 9.67% (8.27–11.27%) → 0.25% (0.09–0.73%); `top10` 14.20%
  (12.52–16.06%) → 2.50% (1.76–3.55%).
- **A probe result that the BRC8004 BNB-team fork is effectively empty** at
  the K=40 windowed sampling cadence (0 mints observed inside the sample).
- **Full auditable provenance** — every numeric cell in `dataset.csv` traces
  through `PROVENANCE.json` to a specific `source_key` inside a specific
  frozen envelope on disk. The five artifact files carry sha256s in
  `SHA256SUMS` that `sha256sum -c` passes at HEAD.

## What this dataset does NOT support — the seven claim traps

### Trap 1 — raw-count-not-columned

**352,415 is a raw registration count, not a distinct-sender count.** The BSC
standard row (`bnb_standard`) carries the whole-history raw count in the
`registrations` column, but every sender-side column on that same row is
drawn from an early-era K=1500 cluster sample bounded to blocks 79,027,268 –
104,927,268 and marked `distinct_senders_is_lower_bound = true`. A reader who
slots 352,415 into the same mental column as Ethereum's 11,352 or Base's
29,289 has mistaken registrations for distinct senders. The dataset enforces
this at column level (the `distinct_senders_is_lower_bound` boolean and the
`notes` marker `raw_count_is_whole_history` sit adjacent to the offending
row), but the trap is real enough that it deserves prose here too. Ratios
inside the census chains land at ~9.5% (Arbitrum) to ~22.3% (Ethereum); the
BSC ratio is much lower but by how much cannot be stated from this dataset.

### Trap 2 — sender counts are lower bounds

**On rows where `distinct_senders_is_lower_bound = true` (rows 4, 5, 6), the
`distinct_senders` column is a lower bound on the true distinct-sender count
in the slice, not the count itself.** Do not sum row 4's 1,289 and row 5's
1,117 and report the sum as "BSC distinct senders" — the two rows are
disjoint cluster samples of a common range and their sum is neither a total
nor a rate. Do not extrapolate either row upward. The dataset enforces trap 2
by row split (BSC is not a single row with a blended figure) and by the
boolean column; this prose is the human-readable version of the same claim.

### Trap 3 — sample/cluster bias

**Cluster samples over an uneven density are order-of-magnitude comparisons,
not point estimates.** BSC registration density is not uniform across its
43M-block deployment-to-head range; both the K=1500 early-era sample and the
K=60 late-era windowed sample are drawn from windows that may over-weight
dense stretches. The Wilson CI95 columns bound sampling error under a
binomial assumption, which is a lower bound on total error when clustering is
present. Read the BSC shares as ranking evidence and as early-vs-late
difference evidence, not as calibrated point estimates against the census
chains.

### Trap 4 — no liveness rate

**The dataset publishes no per-chain liveness rate.** The finished publication
copy (`docs/drafts/erc-8004-four-chain-census-2026-09.md` § "What this does —
and does not — say") documents why: the rate-publishable chains all read
100% on 8 host clusters with a batch-mint dominance flag set at this
checkpoint, so publishing the per-chain split would be worse than not
publishing. The single liveness figure retained in this window is the
previously measured BSC 0.15% live figure, which stands with its own method
sentence and does not appear as a column in `dataset.csv`. A downstream user
who wants "% of ERC-8004 agents that answer HTTP" cannot get it from this
dataset.

### Trap 5 — Arbitrum small-n

**Arbitrum's `top1_share = 18.2%` and `top10_share = 39.1%` are small-n
artifacts.** The chain has 1,502 total registrations across 142 senders;
concentration statistics at that volume are dominated by sample-size effects,
the same way a 10-employee company shows high headcount concentration. The
`notes` column on row 3 carries `small_n_caveat` for exactly this reason.
Arbitrum's top-1 and top-10 shares are **not** evidence that Arbitrum is
"captured" — they are evidence that Arbitrum has ~100× less volume than Base
and ~30× less than Ethereum.

### Trap 6 — BRC8004 empty

**The BRC8004 BNB-team fork row reports 0 mints in a K=40 windowed sample —
not "BRC8004 has zero lifetime registrations."** Task 0661 observed nine mints
outside the K=40 sample (in the first 50,000 blocks after deployment); the
sample happened to miss them. A defensible upper bound on lifetime BRC8004
registrations is low thousands, almost certainly <1% of the standard
registry's 352,359 — and the BRC8004 sender set is disjoint from both BSC
standard-registry samples. The row's `notes` column carries
`fork_effectively_empty`, and it does not appear in the four-chain headline
table for the same reason: it is not census-comparable.

### Trap 7 — external-totals disagreement disclosed

**Public ERC-8004 totals published by third-party explorers and press mentions
disagree with each other by orders of magnitude, and this dataset makes that
disagreement visible rather than laundering any of those figures into the
census.** No external figure is re-cited here without a fetchable source, so
the audit chain of this file terminates at the four chains and six rows it
publishes on its own. This dataset publishes per-chain rows, not a summed row,
and treating a cross-chain sum as authoritative would itself be a mistake
given the checkpoint drift between rows. The disagreement is disclosed in the
finished copy (`docs/drafts/erc-8004-four-chain-census-2026-09.md` § "What
this does — and does not — say" final bullet) and re-flagged here so a reader
who cites the dataset carries the disclosure with them. This dataset does not
reconcile the outlier public figures line-by-line; it publishes its own
numbers, its own method, its own block-time interval, and lets a third party
point their own archive RPC at the same window.

## Trap presence — self-check

Each trap is named with the token given above so a downstream check can find
it by grep:

- Trap 1 header token: `raw-count-not-columned`
- Trap 2 header token: `sender counts are lower bounds`
- Trap 3 header token: `sample/cluster bias`
- Trap 4 header token: `no liveness rate`
- Trap 5 header token: `Arbitrum small-n`
- Trap 6 header token: `BRC8004 empty`
- Trap 7 header token: `external-totals disagreement disclosed`

Every one of the seven tokens appears exactly once as a `### Trap N — …`
heading above.

## Grade

**CONFIRMED — as a set of windowed measurements traced to disk, with seven
named traps carried alongside.** The publish decision is not "these figures
are the state of ERC-8004" — the publish decision is that these are the four
chains, four numbers, one dedup, and one BSC cluster split, at these block
windows, with these seven caveats attached. A reader who copies the census
figures without the traps is misusing the dataset; the dataset is emitted so
that reader has no plausible-deniability defense (the traps live in the CSV
schema, in `DATA-DICTIONARY.md`, in `METHOD.md`, and here).

## Not a claim about future dates

**A next-day resample can and will change the counts.** New mints between the
checkpoint blocks in this dataset and any later re-run will move the raw
counts and, on chains with any concentration structure, will move the top-1
and top-10 shares. This dataset pins an observation to a specific block
interval — nothing more. Corrections are done at a new dated `/data/...` URL,
not by overwriting this one.

## Companion files

- `dataset.csv` — the 6 rows
- `dataset.json` — same 6 rows re-emitted as JSON with attribution + offer
- `dataset.jsonld` — schema.org `Dataset` metadata block
- `PROVENANCE.json` — one entry per numeric cell, traced to source envelope
- `SHA256SUMS` — sha256 of the four data files
- `DATA-DICTIONARY.md` — column-by-column reader's map (20 columns)
- `METHOD.md` — per-row window, RPC, seed, reconciliation history
- `index.html` — the staged (research-dir-only) HTML page
- `../../drafts/erc-8004-four-chain-census-2026-09.md` — finished publication copy (T5a)
- `../../drafts/erc-8004-four-chain-census-dataset-spec.md` — build spec (T5a)
