Run 02 · July 21, 2026
How the Grand Prix learned to drive
· Tahir Lone

Of the three missions at the AES camp, the Grand Prix was the one I rebuilt until it stopped being a game about cars and became a game about teaching. This is the story of how it works, and why it works that way.
You do not program the car
Reinforcement learning is the branch of AI everyone has read headlines about and almost nobody has felt. The textbook version runs on reward functions and patience. A fourteen-year-old has a keyboard and twenty minutes. So the Grand Prix skips every definition and hands over the wheel: you drive the circuit yourself, and your rookie AI starts taking notes the moment you move.
Every finished lap becomes training data. Clean sectors bank rewards. Wall scrapes log penalties. The rookie is not studying the track. It is studying you: your braking points, your lines, your habits, good and bad alike.
Breeding speed
Then comes the part that looks like magic and is actually evolution in a hurry. Every press of RUN TRAINING breeds twenty-four test drivers from your rookie's notes and keeps the fastest, generation after generation, while you watch the whole swarm from a helicopter. Lap times fall. Some generations grind and find nothing. Some find serious pace all at once.
The budget is deliberate: five trainings per teaching lap, then the rookie needs a fresh lesson. That one rule carries the entire point of the mission. Better laps in, better driver out.
You cannot grind your way to a fast car. You can only teach your way there.
Why the physics had to be real
The car has weight, grip, and opinions about hairpins. It had to. If the car is a toy, the demonstration is a toy, and the rookie learns nothing worth watching. A lap runs about a minute and a half: long enough to have a personality, short enough that a scored attempt fits inside a camp afternoon. All of it runs in the browser, and every finished drive compresses into a replay small enough to send racing against the whole field.
The seat that matters
Doodle Dojo teaches by example. Inbox Defender teaches by judgment. The Grand Prix completes the set: teaching by demonstration, where the student becomes the dataset. Watching a car lap a circuit carrying your habits is the fastest explanation of machine learning I have ever found. Nobody in that seat ever asked me what training data means.
The Grand Prix is still open, and the field never sleeps. Train a car and see whose habits it picks up.