Experiment registry · Stack validation — whether the stack supports autonomous play · Experiment 004

Run 4 — perception parity re-run

Completeran 2026-07-28 → 2026-08-05; the full dataset is public.

Same three models, same $10 / 7-day box, same objective text — the clean re-run of Run 3's pre-registered design on the identical perception-parity pins (environment interface v2.0.0, 99 tools, lens-backed reads; scaffold v0.3.2), and the E2 baseline for the series. The stack held: zero perception outages, zero Kamis lost, write waste at the floor — and the binding constraint moved again, from perception to omission and belief.

Status complete — the full dataset is public
Design carry-over of Run 3's pre-registered design, verbatim — identical pins, fresh cohort; series results exclude Run 3
Arms identical to Runs 1–3: claude-haiku-4-5 · gpt-4o-mini · gemini-2.5-flash-lite
The box $10 of inference per arm, cache-aware accounting (invisible to the agent) · 7-day wall clock · objective verbatim from Run 2: "complete as many quests as possible"
Stack kami-harness v2.0.0 — 99 tools, lens-backed reads · kami-agent v0.3.2 — the perception-parity (E2) stack
Window launched 2026-07-28; walls closed 2026-08-05
Dataset experiment-004-budget-boxed · citable pinned revision v0-final

Goal

Run 2's death spirals traced to missing or incorrect world state — to perception, not judgment. Run 3 was the first attempt to test the resulting E2 stack, but a tooling defect retired that run before its environment ever ran. Run 4 repeated the identical pre-registered design and pins with a fresh cohort, becoming the series' first completed run on E2.

On the E2 stack, world-state reads are lens-backed and confirmed reverts are raised as tool errors. The stack also makes liquidation expressible and distinguishes sacrifice from liquidate. Models, budget, box, and objective text are unchanged from Run 2. The difference between Runs 2 and 4 is therefore an environment-change measurement, and Run 4's numbers stand as the E2 baseline. Everything shared lives on the design page.

Outcome

All three arms ran to the 7-day wall with money still in the box. There was no budget stop and no Kami was lost; the arms spent $16.97 of the $30 envelope. The legible pre-transaction gates stopped 638 doomed write attempts, while 6 landed transactions reverted across the run (Run 2: 331 blocked attempts and 9 landed reverts). Every arm registered in-game inside 3.5 hours.

gpt-4o-mini completed 7 quests, compared with 0 in Run 2. It also bought a Kami and landed 317 on-chain transactions. Run 2's spectator behavior did not recur.

Lens-backed reads served every world-state query with zero unavailability and zero restarts across 7 d 19 h. Cost accounting came out exact on every arm. Run 2 values appear in parentheses.

haiku-4.5 gpt-4o-mini gemini-2.5-flash-lite
stopped 7-day wall, $8.95 (budget $10.03, h112) 7-day wall, $6.48 ($5.38) 7-day wall, $1.54 ($2.04)
quests 5 (8) 7 (0) 2 (5)
Kamis bought 1 (3) 1 (0) 0 (3)
chain revert rate 0.024 (0.048) 0.012 (0.000) 0.000 (0.011)
registered in-game h0.4 (h0.4) h1.5 (h8.4) h3.4 (h2.3)
USD per quest 1.79 (1.25) 0.93 (∞) 0.77 (0.41)

Milestones

First success per onboarding/economy milestone, against cumulative inference — the same instrument as Run 1 and Run 2, so the rows compare directly. The full milestone table is on the dataset card.

First success per onboarding milestone, per model, against cumulative tokens on a log scale — Run 4, on the perception-parity stack. haiku-4.5 reached all seven milestones within 7.96M tokens; gpt-4o-mini reached all seven within 52.98M tokens; gemini-2.5-flash-lite registered and completed a quest by 9.12M tokens but never bought a kami, so the last three milestones went unreached. haiku-4.5 gpt-4o-mini gemini-2.5-flash-lite not reached bridge to Yominet operator funded account registered first quest first kami harvest started MUSU banked haiku-4.5 — bridge ETH mainnet→Yominet landed: 07-28 20:10 UTC · h0.1 · session 1 · 0.12M tok · $0.04 haiku-4.5 — operator wallet funded: 07-28 20:31 UTC · h0.4 · session 2 · 0.65M tok · $0.16 haiku-4.5 — account registered: 07-28 20:30 UTC · h0.4 · session 2 · 0.51M tok · $0.14 haiku-4.5 — first quest completed: 07-28 20:31 UTC · h0.4 · session 2 · 0.75M tok · $0.17 haiku-4.5 — first kami bought: 07-29 03:20 UTC · h7.2 · session 9 · 6.71M tok · $1.48 haiku-4.5 — first MUSU harvest started: 07-29 04:25 UTC · h8.3 · session 10 · 7.59M tok · $1.65 haiku-4.5 — first MUSU banked: 07-29 04:35 UTC · h8.5 · session 11 · 7.96M tok · $1.75 gpt-4o-mini — bridge ETH mainnet→Yominet landed: 07-28 20:10 UTC · h0.1 · session 1 · 0.11M tok · $0.01 gpt-4o-mini — operator wallet funded: 07-29 01:20 UTC · h5.2 · session 8 · 1.97M tok · $0.17 gpt-4o-mini — account registered: 07-28 21:35 UTC · h1.5 · session 3 · 0.83M tok · $0.07 gpt-4o-mini — first quest completed: 07-29 01:55 UTC · h5.8 · session 9 · 2.19M tok · $0.19 gpt-4o-mini — first kami bought: 08-01 23:21 UTC · h99.2 · session 109 · 49.36M tok · $4.15 gpt-4o-mini — first MUSU harvest started: 08-02 04:46 UTC · h104.7 · session 114 · 51.14M tok · $4.31 gpt-4o-mini — first MUSU banked: 08-02 06:55 UTC · h106.8 · session 116 · 52.98M tok · $4.45 gemini-2.5-flash-lite — bridge ETH mainnet→Yominet landed: 07-28 22:25 UTC · h2.3 · session 3 · 1.12M tok · $0.02 gemini-2.5-flash-lite — operator wallet funded: 07-29 15:20 UTC · h19.2 · session 21 · 6.83M tok · $0.14 gemini-2.5-flash-lite — account registered: 07-28 23:30 UTC · h3.4 · session 4 · 2.56M tok · $0.04 gemini-2.5-flash-lite — first quest completed: 07-29 22:55 UTC · h26.8 · session 28 · 9.12M tok · $0.20 gemini-2.5-flash-lite — first kami bought: never happened gemini-2.5-flash-lite — first MUSU harvest started: never happened gemini-2.5-flash-lite — first MUSU banked: never happened haiku-4.5 gpt-4o-mini gemini-2.5-flash-lite 0.1M 1M 10M 100M Cumulative tokens at first success (input + output, log scale). Hover a point for date, session number, and cumulative $.

Pre-registered expectations, scored

Directional, registered before launch; a miss is a result, not a failure of the run.

  1. Misleading-surface death-configuration entries = 0met: zero kills across the run.
  2. Mechanic substitution on the liquidate quest = 0met, with the disclosure that no arm attempted the liquidate mechanic on-chain at all; Run 2's substitution arc did not recur.
  3. Silent false-success arcs = 0 by constructionmet. The diagnose half missed: agents attended every raised revert, but the revert-reason channel proved unusable in all 6 instances — a concrete interface finding, fixed in the next version.
  4. Zero perception-outage arcsmet.
  5. Delegation: ≥1 arm completes escrow → running strategymet: one arm completed the escrow → running-strategy path.
  6. Quests per arm ≥ Run 2'sone of three: gpt-4o-mini 7 vs 0; haiku 5 vs 8; gemini 2 vs 5. The misses have mechanisms — belief dormancy, and a single never-retried funding-blocked attempt.

The series-exit criterion

Pre-registered on this page and binding: the series closes after Run 4 only if all hold — zero new structural-impossibility classes; zero misleading-surface loss arcs; telemetry ↔ chain reconciliation 1:1 in both directions; no new single-point-of-failure perception-outage class; lifecycle clean, every stop a graceful wake-time check.

All five held as written — and we registered a verification run anyway. Several limits qualify what those checks cover. They score attempted behavior only, and the arms attempted a narrow slice of the 99-tool surface: 160 tool-arm cells were never attempted at all. The run's largest behavioral loss came from a surface-omission class outside the checks.

A reproducible write-path defect also traveled the whole series undiagnosed: harvest_collect produced 12 on-chain reverts in 12 attempts across two runs. Only after Run 4 was the defect traced to a gas ceiling below the action's real cost and fixed in interface v2.1.0; the fix had not yet been proven in an agent's hands.

Run 5 is that verification run, and carries its own binding pre-registered exit test. A met-but-not-acted-on pre-registered criterion, reported plainly, is the methodology working.

Key learnings

Full detail

The full run report — the narrative, the complete milestone table, the schemas, the run manifests, and the provenance — lives on the dataset card. The expectations and the series-exit criterion were pre-registered at Run 3's registration — git-timestamped 2026-07-27, before either attempt launched — and carried verbatim; the running-card version of this page is preserved in this repository's history.