The whack-a-mole game in prompt engineering: why fixing does not last
Dr. Aaron Hutzler · 14 August 2026 · 4 min

Hours on a until that one error finally went away. Then the next task fails somewhere else entirely.
Welcome to the whack-a-mole game of prompt engineering.
In classical programming a bug fix stays where you can follow it. Change a condition like if x < 0
and the effect on the rest of the application stays traceable. You can test it. Try to repair a
language model through its prompt and you hit a measured phenomenon. Models react extremely
sensitively to small changes in context. In the prompt the cause of that sensitivity is invisible.
1. What the research measures
That tweaking a prompt produces unpredictable results is documented.
Format beats content. Changing only the formatting of a prompt, without touching its meaning, shifts a model's accuracy by up to 76 percent. Larger models did not prevent it and more examples did not either. [1]
A life of its own under fixed settings. Five models, eight tasks, ten runs per task, all configured for output. Accuracy varied between runs by up to 15 percent. Not one model delivered the same value twice across all tasks. [2]
Order decides. The same examples in a swapped order move the hit rate between chance level and peak performance. [3]
The reason lies in the architecture. A model of this class runs the prompt through many layers of arithmetic and ends with a score for every word it might write next. [4] Small changes at the input move those scores. The prompt does not show by how much.
The arithmetic does not repeat either. Ten authors traced that to the way a processor handles decimal numbers. At limited precision addition is not associative. So the summing order decides the outcome. Batch size, number of processors and processor type all change that order. With sampling switched off, one reasoning model varied by up to 9 percent in accuracy. [5]
What happens between the prompt and the answer is not fully unravelled. What is measured is that small changes at the input produce large differences at the output. Adjust a prompt to fix one observed error and you turn a dial whose effect on every other document is unknown. You repair the error in the case in front of you and may silently change the behaviour on many others.
2. Trust through verification instead of hope through prompts
Since a model's reaction to a prompt stays unpredictable, patching prompts leads into an endless loop.
The engineering answer is not to make the prompt infallible. It is to separate drafting from checking:
-
The model drafts. It delivers its answer and must name a verbatim quote for every value.
-
A local checks. It searches for each quote in the source document on your own machine.
If the quote is in the source text and carries the value, the verdict is ACCEPTED. If it is not there, the verdict is REJECTED. What the check settles is where a number came from. Not what it means. And the origin is exactly what no prompt settles. How confident the answer sounded and which prompt produced it plays no part.
The tool runs on your machine and checks an AI summary against your source document
Open the tool3. References
[1] M. Sclar, Y. Choi, Y. Tsvetkov and A. Suhr, "Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design", ICLR 2024, arXiv:2310.11324.
[2] B. Atil, S. Aykent, A. Chittams, L. Fu, R. J. Passonneau, E. Radcliffe, G. R. Rajagopal, A. Sloan, T. Tudrej, F. Ture, Z. Wu, L. Xu and B. Baldwin, "Non-Determinism of 'Deterministic' Settings", arXiv:2408.04667, 2024.
[3] Y. Lu, M. Bartolo, A. Moore, S. Riedel and P. Stenetorp, "Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity", ACL 2022, arXiv:2104.08786.
[4] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser and I. Polosukhin, "Attention Is All You Need", NeurIPS 2017, arXiv:1706.03762.
[5] J. Yuan, H. Li, X. Ding, W. Xie, Y.-J. Li, W. Zhao, K. Wan, J. Shi, X. Hu and Z. Liu, "Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference", arXiv:2506.09501, 2025.
