A 15M-parameter RL agent trained from scratch vs GPT-6 Astra: who would win?
Our AI scientists optimized a 15M-parameter RL agent harness in three weeks. A model trained from random weights defeats Craftax's final boss, a first for any RL model, in 44% of its games. GPT-6 Astra won 1 game in 75 and needed a day to do it. Ours needed 10 seconds.
David and Goliath
Last month, RL researcher Roger Creus Castanyer beat Craftax with OpenAI’s GPT-6 Astra. It was the first time anyone had shown an AI defeating the game’s final boss. Astra won 1 game in 75, and the winning game took almost a day.
Our agent is a 15M-parameter network trained from random weights. It was trained tabula rasa, meaning it had seen no human games, had no language model, and had no written rules. It defeats the final boss in 44% of its games. A win takes about 10 seconds and costs under a cent.
No person wrote its training recipe. Our AI scientists did, in three weeks and 4,805 experiments.
Figure 1. Trackinizer during the Craftax campaign. The campaign's records grow in the order they were made; then two moments told later in the post open as their own neighbourhoods, and the campaign log scrolls to the line that caught the scoring tool.
Craftax is a 2D survival game: Minecraft-style crafting on top of a nine-floor dungeon. The player sees a 9×11 window around itself, plus its health and inventory. It must eat, drink and sleep, craft tools and armour, learn spells, and fight monsters that get tougher on every floor. The last floor holds the Necromancer. Land eight hits on him and you win.
What it learned
Roger had spent years trying to beat Craftax with RL. He wrote:
“I struggle to imagine a small neural network, without the general knowledge an LLM starts with, learning all of these behaviours through reward maximisation alone.”
The best RL agent before ours stopped earlier. Daphne Cornelisse trained it with PufferLib and described where it ends: in the Fire Realm, floor 7. “Then it meets a pigman.”
Pigmen are immune to fire and block 90% of physical damage. Only ice hurts them, and the agent cannot see an enemy’s health. None of Daphne’s games reached the next floor. She asked whether “some domain knowledge [is] needed to guide the search?”
No new knowledge was needed, on the same rules. Across eight training runs, our agents reach the Ice Realm in 56% of games. Here is what the winners do, found by watching 1,284 of their games.
50 winning games in one world
Figure 2. 50 winning games played by our agent from the same start in one world, as ghosts on the 48×48 board of each floor; one game is highlighted in green. Below, each game's score over its length. The boards follow the green game down; use the controls to pause or jump.
Twelve tricks the 15M model learned by RL from gameplay
No one wrote these moves into the agent. It found each one by trial and reward: billions of games, a score at the end of each. Two in full; click the others to watch.
[CLICK TO EXPAND]
It digs a little bed before it sleeps.
It goes back upstairs for groceries.
It builds bridges across lava.
It places stone blocks as its own shield.
It knows which spell to use.
It shoots where monsters are about to step.
It almost never gets hit when one hit means death.
It waits for its moment against the final boss.
It lights up dark rooms before going in.
An earlier agent found a loophole in the boss fight.
Where Astra fumbled
Astra’s winning game shows where general knowledge falls short. It knew the rules, and it still fumbled fights our agent handles routinely:
| Situation | Astra, in its one win | Our agent, across its wins |
|---|---|---|
| A monster closes in | Lets two trolls reach it; health falls 13 → 4 (clip) | Steps away or walls it off; unhurt in 97–99% of one-hit waves |
| A moving target | Casts three fireballs down a lane the troll has left (clip) | Funnels monsters around corners and fires where they will step |
| Terrain | Mistakes a shrub for a wall and steps into a projectile (clip) | Builds piers and lava bridges; blocks about 18 shots a game with stone |
The Astra moments come from one game on a different version of the rules. They are examples, not error rates.
We pointed our AI research team at PufferLib’s version of Craftax. We ported their C/CUDA implementation to our priml harness, which is written in Python/PyTorch. Pufferlib’s best recipe produced 61% return, consisting of a stack of MinGRUs (efficient RNNs), which were hyperparameter tuned to get a much higher score than previous state-of-the-art (~20%).
But one small problem: that pesky pigman again!

It’s invulnerable to everything except ice attacks, and because the the RL agent is trained from scratch, there’s nothing that tells it that ice should beat fire.
But, as it turns out, teaching the agent this concept is not strictly necessary. Through some clever engineering, our AI Scientist team was able to discover a few key techniques to make it work.
Here is what the team produced:
- Past the pigman. Our agents reach the Ice Realm in 56% of games.
- An agent that beats the game. After 20B steps of play, it reaches the the final floor, the graveyard, and can beat the final boss in 2% of games. After fine-tuning a small reward for fighting enemies, it defeats the boss in 44% of games.
- 73% Achievement Return. 8 seeds at 20B steps of training, ± 1.9%. The best public result at that budget is 55%, or 61% with 8 GPUs.
- One GPU per agent. ~11 hours of training for 20B steps on a single H100.
- A world model you can play, trained on the agents’ own games.
How deep agents get
Figure 3. Share of evaluation games reaching each floor, in game order, and winning, as each change is added. Blue: PufferLib's public run. Red: ours, trained on a single GPU. Scores are calculated as the % of the 226 available achievement points.
All of these numbers are on PufferLib’s rules: the agent is told which actions are possible at each step, every game is drawn from a fixed set of 8,192 worlds, and a whole sleep counts as one step. The first of these, the action mask, matters most. It rules out actions that cannot work yet, such as crafting a tool without the materials or taking a ladder before the floor’s kills are made. Astra, in the opening, played the original game with none of those changes, and its 75 games include the ones it used to develop its harness.
More compute was not what moved the score. The best public agent reached 61% after 120B steps, and our own 100B-step run finished at 74%. What moved it was research: 4,805 training runs, most of them in the campaign’s first three weeks, and an estimated 8.7 trillion game steps, to find out what to change. That search is the expensive part, and it is the part our AI scientists do.
Plot twist: Astra is one of those AI scientists, working in harmony with the rest of the team. So yes: you can train models that train models better than you.
The recipe the AI scientists found
The AI scientists started from PufferLib’s recipe: PPO with a recurrent (MinGRU) memory. On one GPU it scored 49% of Craftax’s achievement points in our rerun. They changed six things:
- A board encoder. A small convolutional network reads the 9×11 map. A second encoder reads health and inventory. Both feed the memory. Removing it drops the final score from 74% to 51%.
- An action-effect head. During training, the agent also predicts whether each action will change what it sees.
- Rewards in proportion. PufferLib clips every reward to [−1, 1], so entering the Fire Realm paid the same as collecting wood. Dividing rewards by 8 instead keeps big achievements big.
- Frontier practice. Each time a game reaches a new score level, the trainer saves it, along with the agent’s memory. 20% of training games restart from those saves, favouring levels the agent rarely reaches. It is a player reloading a save before the hard part. It builds on Go-Explore (Ecoffet et al.) and SCALAR (Zabounidis et al.).
- A stall cap. Training games end 10,000 steps after the last reward.
- A boss-fight reward. The game pays nothing for boss hits 2–7. So the agent hit once, then waited out the clock. A 3B-step fine-tune paid +1 per boss hit and +1.5 per final-floor kill, with a longer planning horizon. Wins went from 2% to 44%..
On the last point, the most challenging part for the network was to learn how to defeat the final boss. The rewards are extremely sparse: one when hitting the boss, and one when killing the boss. It requires thousands of steps in between to complete: an agent needs to hit the boss, then clear the wave of spawned enemies, otherwise the boss is invulnerable. This is where 98% of runs die, as the models struggle to assign credit with killing all the waves. The AI Scientists discovered an optional tweak to add a small reward for defeating enemies. As a result the completion rate jumped from 2% to 44%, showing that the model is having trouble knowing what to do when it reaches the final boss.
We set the goal and the rules, suggested some training settings. The rest came from the AI scientists. Changes 1–5 score 73% at 20B steps, averaged over 8 runs. PufferLib’s public run scored 55% at the same budget on 8 GPUs. Each of our runs took about 11 hours on one GPU.
Here are some of the problems the AI scientists solved on the way:
The RL agent could not read the board
In the game. The RL agent sees a 9×11 window of tiles around it, plus its health and inventory. Everything it will ever know about the world arrives through that window. Our rerun of PufferLib’s recipe plateaued at 48.6%.
In the lab. The AI scientists rebuilt how the RL agent sees. A small convolutional network reads the board, a second encoder reads health and inventory, and both feed the RL agent’s memory. The learning algorithm stayed the same: PPO with a recurrent memory, as in PufferLib’s recipe. People guided parts of this step, including the value-loss settings. The first strong score was 66.8%. Rerun with nothing changed, the same recipe scored 57.6%, and 49.9% on another seed. One run is not a result, and from then on every claim rests on 3 to 8 runs. The check that settled it came later: remove the board network from the final recipe and the score falls from 74.1% to 51.1%, in every one of three matched pairs.
From PufferLib's recipe to ours
Figure 4. Each dot is one run trained from scratch and scored at 20B steps on about 10,000 games; bars mark means. Columns are milestones in order, with days counted from the campaign's launch. The third column is this step: 66.8%, then the same recipe again at 57.6% and 49.9%. The last three are the next box. Dashed line: PufferLib's own 8-GPU public run, 55.4%; our 1-GPU rerun of its recipe scored 48.6%. Practice is the frontier practice in the next box; our recipe is practice plus the reward scaling, first rebuilt on our PyTorch port, then run on 8 seeds.
After. With the new eyes, RL agents average 55% across seven runs and begin to reach the Ice Realm, in 6% to 24% of games depending on the run.
The pigman was not the problem
In the game. The Fire Realm’s enemies are immune to fire and block most physical damage. PufferLib’s write-up stops there and asks whether better exploration could get past them, “or is some domain knowledge needed to guide the search?” No public RL agent we found reaches the next floor. Ours reached it in a minority of games, then turned back.
In the lab. The first clue came from watching. RL agents that reached the Ice Realm turned around within about 17 steps and wandered for tens of thousands more. So the AI scientists ran an experiment on the moment itself. They saved 100 games at the point where the RL agent had made the eight kills that unlock the way down, and replayed each one under 12 conditions. From the campaign’s notebook, verbatim, as an AI scientist wrote it:
Among the 90 original failures, privileged full-state routing rescued 47; visible-trigger routing rescued one. Health, necessities, and both without routing rescued zero. […] Forty-six of 47 route-only rescues came from controls that never exposed the ladder, making discovery the dominant causal bottleneck.
In plain terms: restoring the RL agent’s health or supplies saved none of the 90 lost games, and steering it to the ladder saved 47. The RL agent could already beat the pigmen. It could not find the way down.
Two changes followed. One came from reading code. An AI scientist going through the trainer saw every reward clipped to the range −1 to 1. Clipping is a common default that keeps training stable. Here it meant that entering the Fire Realm, an 8-point achievement, paid the same as collecting wood. Dropping the clip and dividing rewards by 8 keeps them in proportion.
The other was practice. Each time a game reaches a new score level, the trainer saves it together with the RL agent’s memory, so it does not wake up at the save with no idea how it got there. The trainer then restarts 20% of training games from those saves, favoring the levels the RL agent rarely reaches. It is a player reloading a save before the hard part. We call it frontier practice, and it builds on Go-Explore (Ecoffet et al.) and SCALAR (Zabounidis et al.). Evaluation games start normally, with no saves.
Neither change was enough alone. Practice by itself reached the Ice Realm in 16% of games. Together, in three matched pairs, they raised the score from 53.7% to 72.5%, and we cannot say how much each contributed.
It took both changes
Figure 5. Share of evaluation games reaching the Ice Realm and the final floor (the Graveyard), by recipe at 20B steps. Practice v1 and v2 are two versions of frontier practice; our recipe is practice v2 with the reward scaling. Practice alone does not get there. Adding the reward scaling does on one seed, and the 8-run panel confirms the pair. The two changes were not separated at full scale with many seeds, and final-floor reach still varies from 0% to 49% across the eight runs.
After. In all eight runs the RL agent reaches the Ice Realm, in 48% to 64% of its games, and on average it reaches the final floor in 26%, with a score of 73%. So the answer to PufferLib’s question, on PufferLib’s own rules: no extra knowledge was needed. More practice in the right places, with rewards kept in proportion, was enough.
Nobody paid the RL agent to finish the fight
In the game. The final floor holds the Necromancer. It takes eight hits, and each hit brings a new wave of enemies. In one of our deepest runs, 39% of games reached it. None won: 0 of 10,007.
In the lab. The AI scientists replayed the lost fights and counted. From the notebook, which counts floors from zero, so its floor 8 is the final floor:
Timeouts (1,698): 97% end with no living hostiles and the spawn timer expired […], i.e. the boss is vulnerable for ~66k ticks and the agent never walks over to melee it. […] Hits 2-7 are unrewarded and training’s 10k stall cap ends 86% of floor-8 training episodes after the first hit, so the late fight gets no learning signal. Spell spam (~8.5k casts) cannot damage the boss.
The RL agent stood off from a defenseless boss and cast spells that could not hurt it. The reason is how the game pays. Hits 2 to 7 earn no reward, so nothing told the RL agent that those hits were progress. A probe confirmed it: its own estimate of how well it was doing barely moved from one hit to the next.
The first route used no game-specific reward, because the campaign’s rules, set by people, allowed none. The AI scientists replayed the RL agent’s own wins, learned in smaller steps, and anchored the weights to the parent model so the wins would not fade. That one chain of fine-tunes, on one seed, won 581 of 36,908 evaluation games, 1.6%. On the way the scoring tool turned out to be hiding wins. The notebook again:
[…] the first 2,048 completions of 2,048 agents are each agent’s shortest episode. True eval win rate of this family ~2%, matching training.
Winning games are long, so a tool that kept each RL agent’s shortest game missed them. The AI scientists measured the bias on the first day, steered by training games instead, then fixed the tool.
The direct route pays for progress in the fight: +1 per boss hit and +1.5 per final-floor kill, with a longer planning horizon (a discount factor of 0.9998 instead of 0.9994, which stretches the RL agent’s foresight from about 1,700 steps to about 5,000). The AI scientists had run this reward early as a diagnostic, where it won about a quarter of its games, and set it aside under the same rule. People later lifted that rule. Each arm below starts from one RL agent, a different run from the 1.6% chain, and trains for 3B more steps unless noted. Evaluation pays no bonus.
| Fine-tuning arm | Boss wins / games | Win rate |
|---|---|---|
| Control: no bonus, nothing else changed | 0 / 10,004 | 0% |
| Bonus, steady learning rate (5B steps) | 2 / 10,109 | 0.02% |
| Bonus, learning rate decaying to zero | 3,505 / 10,043 | 34.9% |
| Bonus, longer horizon, learning rate decaying to zero | 4,292 / 10,039 | 42.8% |
| Bonus, longer planning horizon | 4,415 / 10,052 | 43.9% |
| Same, second fine-tuning seed | 4,383 / 10,024 | 43.7% |
The reward does not work alone. It needs the longer horizon, or a learning rate that decays to zero.
After. The RL agent walks up to the boss and finishes the fight in 44% of its games. Winning that often without paying for each hit is the open problem.
Unlocked by Trackinizer
Each AI scientist is an LLM agent with its own copy of the code and access to a GPU cluster. Every question, experiment and verdict goes into Trackinizer, our open-source Epistemic Knowledge Graph. Each new experiment starts from what is already known. Thousands of linked records turned 4,805 runs into one line of research, not 4,805 guesses.
Two moments from the record, the ones Figure 1 opens:
- The pigman was not the problem. Agents that reached the Ice Realm turned back within about 17 steps. The scientists saved 100 games at the moment the way down opened, then replayed each under 12 conditions. Restoring health or supplies saved none of the 90 lost games. Steering the agent to the ladder saved 47. It could beat the pigmen; it could not find the way down. That finding led to frontier practice.
- The scoring tool was hiding wins. It kept each agent’s shortest game, and winning games are long. The scientists measured the bias on the first day and fixed the tool.
Plot twist: Astra is one of those AI scientists. The large model did the research once. The small model plays 10,000 games in 43 seconds. So yes: you can train models that train models better than you.
The loop has also run on NanoChat and ARC-AGI. It needs a score to push and a way to run it again and again. The tools are open source: Trackinizer, Priml and Configgle. Bring us a hard problem on Discord.