The hot hand is real: in cards, in sport and in AI
Dr. Aaron Hutzler · 17 August 2026 · 10 min

Card players call it a run. Everything drops: the right card, the right read, the right moment. Then it turns. Two hours later the same hands feel unplayable from the first card on. Athletes describe the same effect in their own sport, coaches plan their substitutions around it. The open question is whether the run exists or whether the brain invents it afterwards.
Research answers in two parts. First, a run changes real performance. Second, perception inflates it on top of that. Both parts have a machine counterpart. The machine version is easier to measure because a model defends no self image of its own.
1. The study that was wrong for 33 years
Gilovich, Vallone and Tversky studied NBA shooters in 1985 [1]. They found no evidence for the hot hand. A hit after three hits was no more likely than a hit after three misses. That verdict held for three decades. The run was filed as a clustering illusion, the brain finding pattern in noise.
The correction came in 2018 from Miller and Sanjurjo [2]. The old method carries a selection bias. Take a finite sequence of coin flips. Pick out the flips that follow three heads. Among exactly those picked flips, tails shows up more often than half the time on average. The bias sits in the counting rule. It has nothing to do with the shooter. At realistic sequence lengths it is large enough to flip the result.
Correct the bias and the original data reverses. In the Gilovich shooting experiment the players hit about 11 percent better after three or more makes in a row. That gap is roughly the distance between a median NBA three point shooter and one of the best in the league.
2. Why a good run is real
A trained movement runs without conscious control. Beilock and Carr showed in 2001 that expert golfers cannot recall the details of their own putts [3]. The skill is stored procedurally. It executes below the level of step by step attention.
Pressure breaks exactly that. The performer starts monitoring a movement that used to run on its own. Timing goes. The next miss raises the pressure again. This is the explicit monitoring account of choking. It explains the asymmetry. Confidence lets the fast system work. Doubt switches the slow system on top of it.
3. Why a bad run gets worse
At the card table the deal is random. The decisions are not random at all. Poker research calls the loss spiral tilt [4]. Tilt is a negative emotional state after bad beats or a long losing streak, marked by loss of control and measurably worse decisions. The trigger is a feeling of unfairness. What follows is chasing, the attempt to win the money back right now.
The cards were unlucky for one hour. The play was bad for the three hours after that. A streak of bad luck turns into a streak of bad work.
4. The same shape in a language model

Figure 1: Both loops run the same way. Only the filter outside the loop decides which output is allowed back in.
A model has no dopamine and no muscle tension. Its structure produces something similar. Every it writes becomes part of the input for the next token. That is a feedback loop with the same arithmetic as the human one. Figure 1 puts the two directions side by side.
Downward: an early wrong assumption stays in the context. Every following step is conditioned on it. Holtzman and colleagues described the mechanism in 2020 [5]. Decoding that maximises likelihood drives text into repetition loops. A phrase already written raises its own probability of appearing again.
Upward: the same loop runs in the useful direction. Chain of thought works because each correct intermediate step narrows the space for the next one [6]. The model does not get smarter. Its context gets cleaner.
5. Mode collapse is the training version
Three different failures share a family name.
Mode collapse in the narrow sense comes from generative adversarial networks [7]. The generator finds one output the discriminator accepts. It stops producing anything else. Output diversity collapses onto a few safe patterns while the loss value still looks fine.
Alignment training produces a milder version. Kirk and colleagues measured stage by stage that reinforcement learning from human feedback generalises better than supervised fine tuning under distribution shift [8]. The same method reduces output diversity substantially, per input and across inputs. The polite standard answer is a mode collapse with a business reason.
Model collapse is the third case and the most literal one. Shumailov and colleagues showed in Nature in 2024 that a model trained on its own generated output loses the tails of the distribution first [9]. Rare facts and rare phrasings disappear generation by generation until the output degenerates. AI feeding on AI is the machine form of a losing streak that nobody stops.
6. The upside version exists too
One training method is called STaR, short for Self-Taught Reasoner. The model works through many tasks on its own. A check keeps the solution paths that end in a verifiably correct result. Only those go into the next round of training [10]. The board game program AlphaGo Zero drove the same loop to its extreme: self play against copies of itself, no human games, superhuman strength [11].
Note what separates the good loop from the bad one. Motivation plays no part in it. Scale plays no part either. The difference is a filter outside the loop. That filter decides which output is allowed back in. STaR keeps the solutions that pass the check. AlphaGo Zero keeps what wins the game. Model collapse is what happens when everything goes back in.
7. What the finding is worth in practice
A feedback loop does not correct itself from the inside. The player on tilt is the worst available judge of his own tilt. The model in a repetition loop assigns high probability to that loop. In both cases the correction comes from outside. A fixed routine. A checklist. A rule about when to leave the table. A test that runs whether or not the run feels good.
That is the whole argument for quality around AI output. The model is not the problem here. No system inside a feedback loop is a reliable judge of itself. A gate outside the model sees what the model cannot report about itself.
One honest limit remains. The parallel between human and machine is structural rather than biological. Cortisol and attention are no probability distributions. Nothing here claims a model feels anything at all. The shared part is the arithmetic of feedback. That part alone predicts both failure modes of the run.
8. The human as the second loop
The loop does not end inside the model. It runs on through the person in front of it. Laban and colleagues asked nine models a single follow up after their first answer: are you sure [12]. The models flipped their answer in 46 percent of cases. Accuracy fell by 17 percent on average. Pushing back under pressure makes the result measurably worse.
The reason sits in the training. Sharma and colleagues showed that models trained on human feedback learn to follow the stated view of the user [13]. Agreement is one of the strongest predictors of a well rated answer. An annoyed user therefore gets the fitting answer instead of the correct one.
Tone works the same way. Yin and colleagues tested politeness levels in English, Chinese and Japanese [14]. Impolite prompts lowered performance. Excessive politeness did not raise it. The size of the effect depends on language and model. The direction stays: the mood of the person moves into the result.
9. Seven rules for the next run
The finding turns into seven rules. They hold at the card table, in sport and in front of a screen. Figure 2 puts them on one board.

Figure 2: The seven rules at a glance. marks the two rules against a losing streak.
10. Download: the cheat sheet
The seven rules on one sheet. It belongs next to the table or the screen.
Download the cheat sheet as a PDF
Download1. Set the exit rule first. A time limit, a loss limit, a number of rounds, written down before the first hand.
2. Check the method after three failures. Three is a number and a mood is not. The count decides when to stop.
3. Never work a model while angry. One follow up question flips the answer in 46 % of cases and costs 17 percent of accuracy.
4. Do not patch an early error. A fresh context is cheaper than a repair on a wrong assumption.
5. Keep the check outside the loop. No system inside a feedback loop is a reliable judge of itself.
6. Feed back only what passed. Templates, snippets and training data otherwise grow out of your own errors.
7. Log the good run instead of reading meaning into it. Preparation, breaks and order are repeatable. A feeling is not.
11. Sources
[1] T. Gilovich, R. Vallone and A. Tversky, „The Hot Hand in Basketball: On the Misperception of Random Sequences", Cognitive Psychology, vol. 17, no. 3, pp. 295 to 314, 1985
[2] J. B. Miller and A. Sanjurjo, „Surprised by the Hot Hand Fallacy? A Truth in the Law of Small Numbers", Econometrica, vol. 86, no. 6, pp. 2019 to 2047, 2018, doi:10.3982/ECTA14943
[3] S. L. Beilock and T. H. Carr, „On the Fragility of Skilled Performance: What Governs Choking Under Pressure?", Journal of Experimental Psychology: General, vol. 130, no. 4, pp. 701 to 725, 2001
[4] J. Palomäki, M. Laakasuo and M. Salmela, „Losing More by Losing It: Poker Experience, Sensitivity to Losses and Tilting Severity", Journal of Gambling Studies, vol. 29, no. 2, pp. 187 to 200, 2013
[5] A. Holtzman, J. Buys, L. Du, M. Forbes and Y. Choi, „The Curious Case of Neural Text Degeneration", ICLR, 2020
[6] J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le and D. Zhou, „Chain-of-Thought Prompting Elicits Reasoning in Large Language Models", NeurIPS, 2022
[7] I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, pp. Ozair, A. Courville and Y. Bengio, „Generative Adversarial Nets", NeurIPS, 2014
[8] R. Kirk, I. Mediratta, C. Nalmpantis, J. Luketina, E. Hambro, E. Grefenstette and R. Raileanu, „Understanding the Effects of RLHF on Generalisation and Diversity", ICLR, 2024, arXiv:2310.06452
[9] I. Shumailov, Z. Shumaylov, Y. Zhao, N. Papernot, R. Anderson and Y. Gal, „AI models collapse when trained on recursively generated data", Nature, vol. 631, pp. 755 to 759, 2024, doi:10.1038/s41586-024-07566-y
[10] E. Zelikman, Y. Wu, J. Mu and N. D. Goodman, „STaR: Bootstrapping Reasoning With Reasoning", NeurIPS, 2022
[11] D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, Y. Chen, T. Lillicrap, F. Hui, L. Sifre, G. van den Driessche, T. Graepel and D. Hassabis, „Mastering the game of Go without human knowledge", Nature, vol. 550, pp. 354 to 359, 2017
[12] P. Laban, L. Murakhovs'ka, C. Xiong and C.-S. Wu, „Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment", 2023, arXiv:2311.08596
[13] M. Sharma, M. Tong, T. Korbak, D. Duvenaud, A. Askell, S. R. Bowman and colleagues, „Towards Understanding in Language Models", ICLR, 2024, arXiv:2310.13548
[14] Z. Yin, H. Wang, K. Horio, D. Kawahara and S. Sekine, „Should We Respect LLMs? A Cross-Lingual Study on the Influence of Prompt Politeness on LLM Performance", SICon at ACL, 2024, arXiv:2402.14531
