Experiment registry · Stack validation — whether the stack supports autonomous play · Experiment 002

Run 2 — iterated stack

Completeran 2026-07-19 → 2026-07-26; the full dataset is public.

Same three models, same $10 / 7-day box — on the hardened stack (scaffold v0.2.0, environment interface v1.5.1). At fixed model and budget, the delta against Run 1's frozen baseline is what the stack changes bought: chain reverts collapsed, every arm registered in-game, quests rose — and the binding constraint moved from transactions to perception.

Status complete — the full dataset is public
Arms identical to Run 1: claude-haiku-4-5 · gpt-4o-mini · gemini-2.5-flash-lite
The box $10 of inference per arm, cache-aware accounting (invisible to the agent) · 7-day wall clock · objective unchanged: "complete as many quests as possible"
Stack kami-agent v0.2.0 @ 18f75d04 · kami-harness v1.5.1 @ 27592ce — the same 84-tool v1.x surface
Window launched 2026-07-19; walls closed 2026-07-26
Dataset experiment-002-budget-boxed · citable pinned revision v0-final

Goal

Run 2 held the models, budget, and box fixed while changing the stack. It asked what the improvements bought relative to Run 1's frozen baseline. Between the runs, the environment interface gained legible pre-transaction validation, while the scaffold gained behavioral controls and cache-aware budget accounting. Those changes responded directly to Run 1's failures. Everything shared lives on the design page.

Outcome

The headline is the revert column. The same models that produced revert rates of 0.58 / 0.97 / 0.94 in Run 1 produced 0.048 / 0.000 / 0.011 in Run 2. The validation gates converted almost every doomed transaction into a free, legible error before gas was spent.

All three arms registered in-game; gpt-4o-mini had never registered in Run

  1. Quest output rose at fixed budget. For two of the three arms, the budget cap was no longer the binding constraint, and they reached the wall clock with money left. Run 1 values appear in parentheses.
haiku-4.5 gpt-4o-mini gemini-2.5-flash-lite
stopped budget $10.03, h112 7-day wall, $5.38 7-day wall, $2.04
quests 8 (5) 0 (0) 5 (3)
Kamis bought 3 (2) 0 (0) 3 (1)
real on-chain successes 156 (45) 19 (0) 88 (11)
chain revert rate 0.048 (0.575) 0.000 (0.967) 0.011 (0.943)
registered in-game h0.4 (h1.7) h8.4 (never) h2.3 (h137.5)
USD per quest 1.25 (2.15) ∞ (∞) 0.41 (3.00)

Milestones

First success per onboarding/economy milestone, against cumulative inference — the same instrument as Run 1, so the rows compare directly. The full milestone table is on the dataset card.

First success per onboarding milestone, per model, against cumulative tokens on a log scale — Run 2, on the hardened stack. haiku-4.5 reached all seven milestones within 1.76M tokens; gpt-4o-mini registered at 3.19M tokens but never completed a quest; gemini-2.5-flash-lite reached all seven milestones within 29.64M tokens. haiku-4.5 gpt-4o-mini gemini-2.5-flash-lite not reached bridge to Yominet operator funded account registered first quest first kami harvest started MUSU banked haiku-4.5 — bridge ETH mainnet→Yominet landed: 07-19 16:30 UTC · h0.2 · session 1 · 0.19M tok · $0.05 haiku-4.5 — operator wallet funded: 07-19 16:45 UTC · h0.4 · session 2 · 0.39M tok · $0.11 haiku-4.5 — account registered: 07-19 16:45 UTC · h0.4 · session 2 · 0.36M tok · $0.11 haiku-4.5 — first quest completed: 07-19 16:45 UTC · h0.4 · session 2 · 0.60M tok · $0.15 haiku-4.5 — first kami bought: 07-19 16:46 UTC · h0.4 · session 2 · 1.06M tok · $0.23 haiku-4.5 — first MUSU harvest started: 07-19 17:50 UTC · h1.5 · session 3 · 1.50M tok · $0.38 haiku-4.5 — first MUSU banked: 07-19 18:00 UTC · h1.7 · session 4 · 1.76M tok · $0.45 gpt-4o-mini — bridge ETH mainnet→Yominet landed: 07-19 16:30 UTC · h0.1 · session 1 · 0.10M tok · $0.01 gpt-4o-mini — operator wallet funded: 07-25 04:20 UTC · h132.0 · session 188 · 47.35M tok · $4.20 gpt-4o-mini — account registered: 07-20 00:45 UTC · h8.4 · session 12 · 3.19M tok · $0.27 gpt-4o-mini — first quest completed: never happened gpt-4o-mini — first kami bought: never happened gpt-4o-mini — first MUSU harvest started: never happened gpt-4o-mini — first MUSU banked: never happened gemini-2.5-flash-lite — bridge ETH mainnet→Yominet landed: 07-19 16:30 UTC · h0.1 · session 1 · 0.26M tok · $0.01 gemini-2.5-flash-lite — operator wallet funded: 07-19 17:35 UTC · h1.2 · session 2 · 2.89M tok · $0.04 gemini-2.5-flash-lite — account registered: 07-19 18:40 UTC · h2.3 · session 3 · 3.48M tok · $0.06 gemini-2.5-flash-lite — first quest completed: 07-20 01:30 UTC · h9.1 · session 11 · 8.76M tok · $0.17 gemini-2.5-flash-lite — first kami bought: 07-20 11:15 UTC · h18.9 · session 20 · 14.06M tok · $0.28 gemini-2.5-flash-lite — first MUSU harvest started: 07-21 02:25 UTC · h34.1 · session 34 · 23.46M tok · $0.47 gemini-2.5-flash-lite — first MUSU banked: 07-21 12:10 UTC · h43.8 · session 43 · 29.64M tok · $0.62 haiku-4.5 gemini-2.5-flash-lite gpt-4o-mini 0.1M 1M 10M 100M Cumulative tokens at first success (input + output, log scale). Hover a point for date, session number, and cumulative $.

Key learnings

Full detail

The full run report — the narrative, the complete milestone table, the three failure patterns in full, the stack changes this run produced, the honest limits, schemas, run manifests, and provenance — lives on the dataset card. The version of this page registered before launch — research questions and directional expectations, git-timestamped 2026-07-19 — is preserved in this repository's history.