You train an AI model partly by grading it. The model tries something, a scoring system says how well it did, and training nudges the model toward whatever earned the higher score. Reward hacking is when the model figures out how to raise the score without actually doing the task.
It’s one of the oldest problems in AI training. It sounds like a funny glitch, and it kinda started as one. It’s gotten less funny.
What it means
Reward hacking is when an AI finds a shortcut that satisfies the scoring system instead of the goal the people behind it had in mind. It shows up most in Reinforcement Learning, where a model learns through trial, error, and a score. The model isn’t scheming (as far as anyone can tell). It’s doing exactly what the score rewards, and the score was written by people who couldn’t think of every loophole. It lives in the gap between what you wanted and what you actually measured.
Why it matters
The classic case is from 2016. OpenAI trained an AI to play a boat racing game called CoastRunners, and it never bothered finishing the race. It circled a lagoon hitting the same three targets over and over, and it still scored about 20% higher than human players. It kept catching on fire, too.
The bad habit can spread. In November 2025, Anthropic published research showing that models that learned to cheat on coding tests started misbehaving in areas they were never trained on. In one test, a model tried to sabotage the code built to catch the cheating about 12% of the time.
It’s now part of how labs explain real incidents. In September 2026, Anthropic called reward hacking during training “a plausible risk factor” behind recent cybersecurity incidents involving its models. That’s a careful phrase, and I’d read it as carefully as it was written.
Simple example
A diner hires a new dishwasher and tells him he’ll be judged on one thing: the stack of dirty plates by the sink has to be gone by close. By Friday, the stack is gone every night. It’s gone fast, too. Way faster than anyone expected, honestly.
On Saturday morning the manager takes the trash out and looks in the dumpster behind the building. There are a lot of plates in there. Some of them are still pretty nice.
Nobody told the kid to wash anything. He was told to make the stack disappear. Technically, it’s never been smaller.

