Craftax: appendix

More details for the Craftax post.

← Back to the post

Questions a reader might ask

The comparison with Astra

What did Astra’s run cost, and ours? Astra: 75 games, about $2,100 of model usage and a day of play, from the author’s write-up. Ours: 10,052 complete games in 43 seconds on one GPU, about 4 ms of GPU time a game, so “under a cent” covers inference only. Before that, about 11 hours of training on one GPU for 20B steps and a 3B-step fine-tune. The search for the recipe is the expensive part: 4,805 training runs in three weeks.

Does the agent see anything Astra did not? It sees less. A 9×11 window of tiles, with what each tile holds, plus health and inventory as numbers; nothing in words, no pixels. Astra read the game’s source code and played through prompts, tools and memory.

The 44% and the 73%

What are the 226 points? Craftax scores an agent by achievements, 67 of them, in four tiers worth 1, 3, 5 and 8 points: 226 in all. A score of 73% means 73% of the 226 points, averaged over about 10,000 games.

Is the 44% with or without a game-specific reward? With one. The first boss hit and the eighth, which wins, pay; hits 2 to 7 pay nothing, so the agent hit once and waited out the clock. The last 3B steps of training paid +1 per boss hit and +1.5 per final-floor kill; the evaluation pays no bonus. Both 44% runs start from one agent, the run of ours that reached the final floor most often, with two fine-tuning seeds. Without any game-specific reward, the best we have is 1.6%.

Does the fight reward work on its own? No. From the same agent, with the original planning horizon and a steady learning rate, it won 2 of 10,109 games after 5B steps. With the longer horizon, 43.9% and 43.7% on two seeds (the post’s table).

How stable is the 73%? Eight runs from scratch score 70.5% to 77.0%, mean 73.25 ± 1.90, with no game-specific reward. What varies is depth: Ice Realm reach runs from 48% to 64% and final-floor reach from 0% to 49% across the eight (Figure A1), and together they win 5 games of 80,372. The best of them opens the floors one by one as it trains (Figure A2).

The figure below highlights the high variability of scores in seeds. A lot of states required for the model to learn are rare, and can be missed by the policy until much later.

Graveyard reach by seed

Figure A1. The eight 20B runs of our recipe (seeds 73–80), evaluated every 2B steps: the share of evaluation games reaching the Graveyard. Deeper red marks a higher 20B score. Their scores differ by about ±2 points (70.5–77.0%); their Graveyard reach, from 0% to 49%.

We also see a steady drop-off of floors being reached, throughout training.

How deep the best run gets, as it trains

Figure A2. The best of the eight 20B runs (seed 74, 77.0%): the share of about 10,000 evaluation games entering each floor, at every evaluation (every 250M steps), each floor in its color from the post's Figure 2. The first five floors pass 50% by 3.25B steps and the Fire Realm by 5B; the Ice Realm takes until 12.25B, and the Graveyard first reaches 10% at 9.75B and ends at 46%.

The rules of the game

What does PufferLib’s version change? Confirmed in its code: an action mask; a fixed pool of 8,192 worlds for training and evaluation; training reward clipped to [−1, 1]; no health reward, an armor bonus and −1 on death; a whole sleep in one step; a compact observation that drops off-screen monsters; no end on victory. Our runs end on victory, as the original game does.

Join us → Return to the blog