Why did rewards make monkeys worse at solving puzzles?
One of the most surprising findings in psychology came from Harry Harlow's puzzle-box experiments.
Monkeys happily solved mechanical puzzles with no food reward at all. The puzzles themselves seemed to be enough.
Then Harlow introduced food rewards.
Performance didn't improve.
In many cases, it actually got worse.
Years later, psychologists found the same pattern in humans. Children who loved drawing often drew less after they began receiving expected rewards for it. This became known as the overjustification effect.
Why should reward reduce motivation?
Within the Rozum Framework, the answer is surprisingly simple.
The reward doesn't just change behavior.
It changes what the system is optimizing.
Without an external reward, exploration is driven by one objective: increasing the future Reliability of Output (RO) by reducing uncertainty. Solving the puzzle improves the agent itself.
Within RF, introducing an expected reward changes the optimization landscape. The system shifts from exploration-driven uncertainty reduction toward efficient resource acquisition, making the Existence Axiom increasingly dominant in determining behavior.
Now exploration is no longer "free."
Every unnecessary action costs time and energy. The system naturally shifts from maximizing knowledge to optimizing resources.
Nothing is "wrong" with the learner.
The optimization target has changed.
From this perspective, both Harlow's monkeys and the overjustification effect are not anomalies.
They are exactly what we should expect from any adaptive system operating under Open World conditions.
Curiosity optimizes future reliability.
Rewards optimize immediate survival.
Confusing the two has hidden an important architectural distinction for decades – a distinction that becomes most evident precisely in the monkeys.
This leads to a testable prediction: expected rewards should suppress exploration only when they successfully shift the system's objective from uncertainty reduction to resource acquisition. If we find a way to decouple the reward from survival optimization, the curiosity should remain intact.
If rewards change what an agent optimizes rather than simply increasing motivation, what experiment would reveal the difference?
If you want to participate, please contact rozum.framework {at} gmail.com
