Teaching Snake to play itself
There is no movement logic in this code. The agent gets eleven booleans and a score, and works out the rest by dying a lot.
Snake is a good first reinforcement learning problem because the rules fit in a sentence and the failure is obvious. I wanted to write the whole loop myself: the game, the state, the network and the training step, so that none of it was a black box I had imported.
What the agent is allowed to see
The temptation is to hand the network the whole board. I gave it eleven booleans instead:
- three danger flags: is there a collision straight ahead, to the right, to the left
- four direction flags: which way the snake is currently heading
- four food flags: is the food left, right, above, below
The important thing is that the danger flags are relative to where the snake is facing, not absolute. "Danger to the right" means different board squares depending on the heading, and the state code works that out before the network ever sees it.
That choice does most of the work. It means the same eleven values describe the same situation regardless of which way the snake happens to be pointing, so a lesson learned while moving up transfers to moving left for free. Give the network absolute coordinates instead and it has to learn each orientation separately from scratch.
Three actions, not four
The action space is straight, right turn, left turn, encoded one-hot. There is no "move up". Snake cannot reverse into itself, so an absolute four-direction action space contains one illegal move in every state, and you end up writing rules to mask it. Relative turns make every action legal in every state and there is nothing to special-case.
The reward stays boring on purpose
Plus ten for eating, minus ten for dying, zero the rest of the time. That is all of it.
It is tempting to shape it: a small reward for moving toward the food, a penalty for moving away. That usually backfires. Reward the snake for closing distance and it will learn to hover next to the food, collecting approach rewards without committing, because the shaped signal is easier to farm than the real one. Leaving the reward sparse means the only way to score is to actually eat.
There is one guard: a game ends if it runs past a hundred frames per segment of snake without eating. Otherwise an agent that has learned to avoid dying but not to eat will circle forever and the training loop stalls on an episode that never ends. It is a timeout, not a reward signal, so it does not teach the agent anything except that stalling is not a strategy.
Exploration by counting games
Epsilon is 80 - games_played, compared against a random
draw from 0 to 200. So the first game is about a 40% chance of a random
move, and after eighty games the randomness is gone entirely and the
agent is purely following its own predictions.
This is crude compared to a proper decay schedule, but it is one line and the behaviour is completely predictable: early games are noisy exploration, later ones are the policy. Nothing here needed more than that.
Two training calls per move
Every step trains twice. Once immediately on the single transition that just happened, and once on a random batch of a thousand transitions sampled from a replay buffer of the last hundred thousand.
The replay sampling is the part that matters. Consecutive frames of a Snake game are almost identical, so training only on them in order means every gradient step points the same direction and the network overfits to whatever the snake is doing right now. Sampling randomly across the buffer breaks that correlation.
The network itself is two linear layers with a ReLU between them, eleven inputs to two hundred and fifty six hidden to three outputs. Nothing exotic. The target is the standard Bellman update: the reward, plus the discounted best predicted value of the next state, unless the game ended, in which case it is just the reward.
Watching it get better is the point
The first few dozen games are painful. It walks into walls, doubles back into itself, and scores nothing. Then somewhere around game eighty the random moves stop and the score starts climbing, and it is genuinely strange to watch something work out a strategy you never wrote down.
What I took from it is that almost all of the design was in the state representation. The network is textbook and the training step is the Bellman equation copied honestly. Deciding what eleven numbers the agent gets to see is where the actual thinking went.