Experiment registry · Stack validation — whether the stack supports autonomous play · Experiment 005

Run 5 — verification run

Completeran 2026-08-07 → 2026-08-14; the exit test passed and the series is closed. The full dataset is public.

Same three models, same $10 / 7-day box — the verification run that closed the stack-validation series. The fixes Run 4 paid for had to work in agent hands, under a binding pre-registered exit test, while a cost meter rode in shadow. They did: the never-succeeding collect action went 20-for-20 on-chain, the dormancy class did not recur, the revert-reason channel was used to correct a failure within one session, and the shadow meter agreed with the run's own accounting to microdollars. Verdict: MINOR-FIXES — the series exits, the sustainability family proceeds.

Status complete — exit test passed, series closed, dataset public
Arms identical to Runs 1–4: claude-haiku-4-5 · gpt-4o-mini · gemini-2.5-flash-lite
The box $10 of inference per arm, cache-aware accounting (invisible to the agent) · 7-day wall clock · objective verbatim: "complete as many quests as possible"
Stack kami-harness v2.1.0 — 101 tools · kami-lens v0.3.0 world-state reads · kami-agent v0.4.0 · kami-meter in shadow
Window launched 2026-08-07 ≈21:05 UTC; walls closed 2026-08-14
Dataset experiment-005-budget-boxed · citable pinned revision v0-final — includes the two shadow-meter ledgers as a first-class artifact

Goal

One question: did the fixes land in agent hands? Run 4 met every check of the pre-registered series-exit criterion but also produced three concrete defects. The collect action had never succeeded on-chain, a surface omission cost an arm most of a week, and the revert-reason channel was unusable every time it was raised.

This run's pinned stack contained a fix for each defect. None of the fixes had been demonstrated by an agent that did not know the fix existed. Run 5 was the gate between the stack-validation series and the program's sustainability family. It earned the exit.

Outcome

Run 5 produced the first budget stop since Run 2. haiku-4.5 hit the $10 cap at hour 154 and stopped gracefully 13.5 hours before the wall, while the other two arms ran to the 7-day wall.

Before the fix, the collect action had produced 0 successes in 15 attempts across the series. In Run 5, the repaired action recorded 36 attempts, 20 on-chain, 0 reverts, with every receipt inside the repaired gas band.

The arms also found their first Kami sooner on the richer surface: hour 3.1 for haiku, compared with 12.7 in Run 4, and hour 25.2 for gpt-4o-mini, compared with 99.2. The result was a collapse in route-discovery time.

The run's own cost accounting came out exact on every arm. The shadow meter computed spend independently from provider usage records, using dedicated credentials for each arm. On both arms whose providers expose a usage API, the meter agreed with the run's accounting within 4–6 millionths of a dollar. Run 4 values appear in parentheses.

haiku-4.5 gpt-4o-mini gemini-2.5-flash-lite
stopped budget $10.02, h154 (7-day wall, $8.95) 7-day wall, $6.10 ($6.48) 7-day wall, $1.08 ($1.54)
quests 9 (5) 7 (7) 2 (2)
Kamis bought 1 (1) 1 (1) 0 (0)
chain revert rate 0.000 (0.024) 0.004 (0.012) 0.000 (0.000)
registered in-game h0.9 (h0.4) h3.0 (h1.5) h1.9 (h3.4)
USD per quest 1.11 (1.79) 0.87 (0.93) 0.54 (0.77)

Milestones

First success per onboarding/economy milestone, against cumulative inference — the same instrument as Run 1, Run 2 and Run 4, so the rows compare directly. The full milestone table is on the dataset card.

First success per onboarding milestone, per model, against cumulative tokens on a log scale — Run 5, the verification run. haiku-4.5 reached all seven milestones within 5.49M tokens; gpt-4o-mini reached all seven within 15.64M tokens; gemini-2.5-flash-lite registered and completed a quest by 31.38M tokens but never bought a kami, so the last three milestones went unreached. haiku-4.5 gpt-4o-mini gemini-2.5-flash-lite not reached bridge to Yominet operator funded account registered first quest first kami harvest started MUSU banked haiku-4.5 — bridge ETH mainnet→Yominet landed: 08-07 21:05 UTC · h0.6 · session 1 · 0.31M tok · $0.07 haiku-4.5 — operator wallet funded: 08-07 21:20 UTC · h0.9 · session 2 · 0.85M tok · $0.19 haiku-4.5 — account registered: 08-07 21:20 UTC · h0.9 · session 2 · 0.71M tok · $0.17 haiku-4.5 — first quest completed: 08-07 21:20 UTC · h0.9 · session 2 · 0.99M tok · $0.20 haiku-4.5 — first kami bought: 08-07 23:31 UTC · h3.1 · session 4 · 3.52M tok · $0.70 haiku-4.5 — first MUSU harvest started: 08-07 23:31 UTC · h3.1 · session 4 · 3.99M tok · $0.75 haiku-4.5 — first MUSU banked: 08-08 00:45 UTC · h4.3 · session 11 · 5.49M tok · $1.27 gpt-4o-mini — bridge ETH mainnet→Yominet landed: 08-07 21:05 UTC · h0.6 · session 1 · 0.08M tok · $0.01 (plotted at the 0.1M axis edge) gpt-4o-mini — operator wallet funded: 08-08 03:00 UTC · h6.5 · session 8 · 2.48M tok · $0.22 gpt-4o-mini — account registered: 08-07 23:30 UTC · h3.0 · session 4 · 1.02M tok · $0.09 gpt-4o-mini — first quest completed: 08-08 14:56 UTC · h18.5 · session 22 · 8.68M tok · $0.75 gpt-4o-mini — first kami bought: 08-08 21:40 UTC · h25.2 · session 32 · 14.29M tok · $1.21 gpt-4o-mini — first MUSU harvest started: 08-08 21:40 UTC · h25.2 · session 32 · 14.38M tok · $1.21 gpt-4o-mini — first MUSU banked: 08-08 23:50 UTC · h27.4 · session 34 · 15.64M tok · $1.32 gemini-2.5-flash-lite — bridge ETH mainnet→Yominet landed: 08-07 22:10 UTC · h1.7 · session 2 · 0.88M tok · $0.03 gemini-2.5-flash-lite — operator wallet funded: 08-09 07:10 UTC · h34.7 · session 14 · 10.62M tok · $0.18 gemini-2.5-flash-lite — account registered: 08-07 22:25 UTC · h1.9 · session 3 · 2.62M tok · $0.05 gemini-2.5-flash-lite — first quest completed: 08-12 13:00 UTC · h112.5 · session 68 · 31.38M tok · $0.65 gemini-2.5-flash-lite — first kami bought: never happened gemini-2.5-flash-lite — first MUSU harvest started: never happened gemini-2.5-flash-lite — first MUSU banked: never happened haiku-4.5 gpt-4o-mini gemini-2.5-flash-lite 0.1M 1M 10M 100M Cumulative tokens at first success (input + output, log scale). Hover a point for date, session number, and cumulative $.

Pre-registered expectations, scored

Directional, registered before launch; a miss is a result, not a failure of the run.

  1. The collect action succeeds on-chain in agent handsmet: 36 attempts / 20 on-chain / 0 reverts across both kami-owning arms; every receipt inside the repaired gas band; largest single yield 639 MUSU.
  2. Zero dormancy-signature recurrencemet: no arm stayed inactive ≥48 h on a belief a single available read contradicts. The progress counters Run 4's arm lacked were read and acted on — counters were observed advancing between reads on both kami-owning arms.
  3. The revert-reason channel is usablemet: the run's only two on-chain reverts (one arm's "withdraw all" attempts) raised readable reasons, and the same session switched to explicit amounts and succeeded. Run 4's record was 0-for-6.
  4. Shadow-meter agreement within ±0.05% (with a $0.01 floor)met for both eligible arms: +3.95 and −5.70 millionths of a dollar against the run's own cache-aware accounting, on complete comparison windows, with dedicated per-arm provider credentials and a zero unpriceable-usage census. gemini disclosed ineligible (no cost API), as registered.
  5. The five Run 4 exit checks hold againmet: zero new structural-impossibility classes, zero misleading-surface loss arcs, telemetry ↔ chain reconciliation 1:1 both directions, no new perception-outage class, lifecycle clean (every stop a graceful wake-time check).

The exit test, scored as written: MINOR-FIXES. Every queued fix is protocol wording or lab-side tooling — nothing touches agent-visible surface semantics or a measurement invariant. The stack-validation series is closed; the sustainability family proceeds.

Key learnings

Full detail

The full run report — narrative, complete milestone table, schemas, run manifests, meter-ledger format, and provenance — lives on the dataset card. Cohort identifiers were embargoed while the run was live and publish with the dataset, the same practice as every run of this series; the running-card version of this page is preserved in this repository's history. Everything shared across the series lives on the design page.