Why Rats Learn Without Reward
Why do rats learn mazes when there's no reward?
Tolman's classic latent learning experiments showed that rats build cognitive maps of mazes even when no food is present. This was a serious problem for classical behaviorism, which assumed that learning requires immediate reinforcement.
Tolman's answer was that behavior is goal-directed. But this raises another question:
What exactly is the goal when there's no food to find?
Within the Rozum Framework (RF), the answer is different.
The goal is not food. The goal is maximizing the Reliability of Output (RO).
Under Open World conditions, the best way to maximize future reliability is often to reduce uncertainty before it becomes relevant. Exploration isn't a search for reward -- it is a search for information that will make future decisions more reliable.
Mechanically, the rat is updating the certainty of statements (CS) while leaving the certainty of algorithms (CA) unchanged. It learns facts about the environment ("left corridor = dead end", "right corridor = exit") while using the same navigation strategy. When food finally appears, the existing algorithm operates on a far more reliable model of the world, immediately producing high RO.
From this perspective, latent learning is not an anomaly or an exception to reinforcement learning.
It is what any adaptive system must do when future uncertainty cannot be eliminated in advance.
Curiosity is not reward-seeking.
It is uncertainty reduction.
Or, in RF terms, сuriosity is the computational obligation of any conscious adaptive system operating under Open World constraints.
- https://www.linkedin.com/posts/bayaro_cognitivescience-aialignment-neuroscience-share-7485384447975079937-MBUY
- https://www.linkedin.com/posts/bayaro_cognitivescience-aialignment-neuroscience-share-7489967527326846976-UnMm
If you want to participate, please contact rozum.framework {at} gmail.com
