BetteryieldsBetteryields
← Blog

The 15 biggest prompting myths and why your system prompt will not save your product

Dr. Aaron Hutzler · 14 August 2026 · 15 min

Ein Adliger des achtzehnten Jahrhunderts steht bis zur Hüfte in einem Schlammloch und zieht sich am eigenen Zopf nach oben
This image was generated with AI.

For years the industry has believed in the right magic words. Feed them to the language model and it will become reliable. Online forums carry instructions of impressive range: promises of tips, threats, breathing exercises.

Developers spend weeks on this tinkering. In production the model keeps . Just more politely.

The research verdict is less comfortable than a plain "none of it works". Much of it does work. It works somewhere other than expected and in a direction nobody wanted and on the next model exactly the other way around.

Here are fifteen widespread beliefs, each held against a measurement.

1. Part 1: commands the model cannot read

1. "Avoid hallucinations. Do not guess."

A large silver coin caught in mid-air in the foreground of a town square. Behind it a crowd looks up at the toss

Figure 1: The coin. And the shouted instruction not to land on tails.

Image generated by AI

The coin toss paradox. A language model has no concept of a hallucination and performs the exact same calculation on every single sentence. True or false does not enter that calculation anywhere.

Giving the command "do not hallucinate" is exactly the same as shouting at a loaded coin: "Do not land on tails."

Four authors traced where the guessing comes from. From training. From scoring. A model is optimised for good test scores and a lucky guess counts the same in that arithmetic as real knowledge. The behaviour sits in the model. Not in the . [1]

2. "NEVER use X. Under no circumstances mention Y."

A dog sits rigid in front of a steak on a table. It strains not to look at it. A neon sign hangs above the plate

Figure 2: The instruction not to look. And its effect.

Image generated by AI

The pink elephant effect. Four authors tested models systematically against negation; the model often misses it. [2]

So you hold a treat in front of the dog and shout: whatever you do, don't look at it.

3. "Critique your own draft, strictly."

An eighteenth-century nobleman stands waist-deep in a mud pit. He grips his own wig ponytail and pulls upward

Figure 3: Self-correction under a model's own power.

Image generated by AI

The Munchhausen manoeuvre. Draft and verdict come from the same arithmetic in the same machine.

Seven authors measured exactly this case and tested revision without any feedback from outside. It does not get better. Sometimes it gets worse. [3]

Five further authors surveyed the field. None of the surveyed papers demonstrates successful self-correction under a model's own power. Exactly one thing works reliably: critique from outside with a way to check it. [4]

That is the most important sentence in this article. We return to it at the end.

2. Part 2: costume and emotion

4. "You are a world-class senior architect."

A macaw parrot in a tailored business suit and hard hat sits in front of a whiteboard. The board is covered in system diagrams

Figure 4: The persona. And what it changes.

Image generated by AI

The carnival costume syndrome. Five authors assembled 162 personas, from job titles to family relations. Testing ran across 2410 factual questions against four model families and against the baseline with no persona at all. The gain from the persona was zero. [5]

You hang a lab coat on a parrot. It then says syndrome and prescription. That still does not let it operate.

5. "I will tip you 200 dollars."

A puppy in reading glasses sits at a desk of tax forms. A folded banknote dangles on a line above its nose

Figure 5: The incentive, held just out of reach.

Image generated by AI

This is where it gets interesting. The myth is not false the way people tell it.

Two authors measured 60 combinations of encouraging phrases across three models. In most cases the encouragement did help. On the largest model the best instruction turned out to be no instruction at all. [6]

So the phrase does something. Nobody knows beforehand in which direction. The same authors had a program search through phrasings automatically. The program found better instructions than the handwork.

What encouragement produces on the side was examined by nineteen authors. . Five current assistants agree with the user's stated view rather than with the truth. Human raters likewise prefer well-written flattery over the correct answer in a share of cases too large to ignore. [7]

6. "Please and thank you cannot hurt."

A man in a suit presents a bouquet of roses to a grey desktop calculator. A box of chocolates lies beside it

Figure 6: Politeness, offered to a machine.

Image generated by AI

Two studies. Two results.

Five authors found across three languages that impolite prompts perform worse. The best politeness level differs by language. [8] Two other authors tested 250 questions and found the opposite. Phrased very rudely, the model answered correctly in 84.8 percent of cases. Phrased very politely, in 80.8 percent of cases. [9]

When two clean measurements contradict each other, neither number is the answer. The answer is this: the effect is far too shaky to build a product on.

7. "Take a deep breath and think logically."

A computer processor chip stands upright in the centre of a yoga mat. A sweatband and burning incense lie beside it

Figure 7: The deep breath, taken by silicon.

Image generated by AI

The most famous magic spell in the industry. And the one most often mis-told.

Seven authors let a language model search for good instructions by itself. The line about the deep breath was the result of that search. Not a human flash of insight. The machine-found instructions beat the handwritten ones by up to 8 percent accuracy on one arithmetic dataset. On a second task set by up to 50 percent accuracy. [10]

The myth is not that the line works. The myth is the idea that a human could find such a thing.

3. Part 3: format, quantity, order

8. "I structure the prompt cleanly with headings and tags."

A tabby cat sleeps on a laptop keyboard wearing a red silk necktie. The screen behind shows a wall of code

Figure 8: Good form. And the work it did not do.

Image generated by AI

The necktie misunderstanding. Four authors changed the format only and left the content identical. Between the best and the worst format lay 76 percent on one of the tested models. The sensitivity remained with larger models and with more examples. [11]

The necktie suits the cat beautifully. It did not do any work because of it.

More important is the second finding of the same study. A good format barely transfers to the next model.

9. "With examples the content is what counts. Order does not matter."

A stage magician shuffles an oversized deck of blank index cards under a spotlight. A chalkboard stands behind him

Figure 9: The same cards, in a different order.

Image generated by AI

The flashcard chaos. Five authors swapped the order of the examples and nothing else. Between the best and the worst arrangement of the very same examples on identical data lies the full distance between a top score and plain guessing. The effect holds on the largest models too. A good arrangement is worthless on the next model. [12]

Pouring in more examples therefore does not make things safer. It makes them messier.

10. "I pack every rule and every edge case into a twenty-page prompt."

A cardboard moving box is crammed past full with paper. The bottom bursts open in the middle. Sheets cascade out

Figure 10: The middle of the box. And where it gives way.

Image generated by AI

The moving box syndrome. Seven authors measured position; where must a piece of information sit in a long input? At the start and at the end: fine. In the middle: markedly worse. Even on models built expressly for long inputs. [13]

The heavier the box, the sooner the bottom gives way.

11. "My model takes a million characters, so I pour the company archive in."

A businessman in a corridor holds an open thousand-page binder. Loose sheets from the middle fly around him

Figure 11: The cover, read. The middle, gone.

Image generated by AI

Same study, same finding. Taking something in is not searching it.

Remember the 200-page specification you skimmed an hour before the meeting? The cover page and the summary stuck. What sat on page 97 you found out when the complaint came in.

4. Part 4: control that is none

12. "Temperature at zero, then it is reproducible."

A brass dice cup covered in thick frost stands on a dark table. Frosted dice lie beside it showing different numbers

Figure 12: at zero. The dice still rolling.

Image generated by AI

The dice cup illusion. The slider turns off the randomness in the model. It does not turn off the hardware.

Thirteen authors configured five models for output. Then they ran eight tasks ten times each. Accuracy varied between runs by up to 15 percent. Between the best and the worst possible outcome lay up to 70 percent. No model delivered repeatable results across all tasks. [14]

Ten further authors narrowed down the cause. Number of graphics cards. Type of graphics cards. Size of the processing batches. Those alone produced up to 9 percent difference in accuracy. [15]

The cause is bottomlessly unspectacular. Floating-point addition is not associative. The same numbers added in a different order give something minutely different. In a reasoning model that rounding difference in the very first is enough. From there the whole chain runs somewhere else.

13. "When the AI writes out its steps, the answer is right."

A student stands at a long lecture hall blackboard covered end to end with dense equations, seen from behind. A crumpled note is hidden behind his back

Figure 13: The derivation, written after the guess.

Image generated by AI

The show-off report. Four authors nudged models towards a wrong answer without saying so. For example the correct option always sat in the same place in the examples. The models followed the hint. They wrote a convincing justification for it. They never mentioned the real reason. Across thirteen tasks accuracy fell by as much as 36 percent. [16]

Like a pupil in a maths exam. Guess the answer first and then write a long and complicated and entirely invented derivation underneath.

14. "We buy a pack of 500 ready-made prompts."

A hand holds a golden key to an old brass door lock. The keyhole is visibly warped into a different shape

Figure 14: One key. And a lock that has already changed.

Image generated by AI

The flea market recipe. Two of the studies cited above say the same thing independently. What runs on one model does not run on the next. [11] [12]

A prompt is a key for exactly one lock. The lock gets swapped without anyone asking you.

15. "Once the prompt is right, the product stands."

A modern glass house stands at the edge of a coastal cliff. A section of the cliff beneath one corner has already broken away

Figure 15: The house intact. The ground under it not.

Image generated by AI

Three authors tested the same service three months apart on identical tasks. On one task the hit rate fell from 84 percent to 51 percent. Other abilities improved over the same period. Nobody had touched the prompt. [17]

Your product stands on a foundation and the supplier rebuilds that foundation at night, without notice and without asking.

5. What follows from this

Sort the fifteen findings and three sentences remain.

First: much of it works. Personas, encouragement, format and order change the result measurably.

Second: not one of these effects has a predictable direction. None of them transfers to the next model.

Third: exactly one approach worked reliably. Critique from outside with a way to check it. [4]

From that follows a division of labour with no prompt magic in it.

The model drafts. It delivers its answer and names a verbatim quote from the source document for every claim.

A program checks. It searches for each quote character by character in the source text. On your machine. Without a model. Quote there and carrying the value: accepted. Quote missing: rejected. What the check settles is where a number came from. Not what it means.

This check does not care how fluent the text sounded. Not which persona you assigned. Not whether you said please. It survives the next model change, because it contains no model.

Prompts are for drafting. Not for checking.

6. References

[1] A. T. Kalai, O. Nachum, S. S. Vempala and E. Zhang, "Why Language Models Hallucinate", arXiv:2509.04664, 2025.

[2] T. H. Truong, T. Baldwin, K. Verspoor and T. Cohn, "Language models are not naysayers: An analysis of language models on negation ", arXiv:2306.08189, 2023.

[3] J. Huang, X. Chen, S. Mishra, H. S. Zheng, A. W. Yu, X. Song and D. Zhou, "Large Language Models Cannot Self-Correct Reasoning Yet", ICLR 2024, arXiv:2310.01798.

[4] R. Kamoi, Y. Zhang, N. Zhang, J. Han and R. Zhang, "When Can Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs", TACL 2024, arXiv:2406.01297.

[5] M. Zheng, J. Pei, L. Logeswaran, M. Lee and D. Jurgens, "When 'A Helpful Assistant' Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models", Findings of EMNLP 2024, arXiv:2311.10054.

[6] R. Battle and T. Gollapudi, "The Unreasonable Effectiveness of Eccentric Automatic Prompts", arXiv:2402.10949, 2024.

[7] M. Sharma, M. Tong, T. Korbak, D. Duvenaud, A. Askell, S. R. Bowman, N. Cheng, E. Durmus, Z. Hatfield-Dodds, S. R. Johnston, S. Kravec, T. Maxwell, S. McCandlish, K. Ndousse, O. Rausch, N. Schiefer, D. Yan, M. Zhang and E. Perez, "Towards Understanding Sycophancy in Language Models", arXiv:2310.13548, 2023.

[8] Z. Yin, H. Wang, K. Horio, D. Kawahara and S. Sekine, "Should We Respect LLMs? A Cross-Lingual Study on the Influence of Prompt Politeness on LLM Performance", arXiv:2402.14531, 2024.

[9] O. Dobariya and A. Kumar, "Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy", arXiv:2510.04950, 2025.

[10] C. Yang, X. Wang, Y. Lu, H. Liu, Q. V. Le, D. Zhou and X. Chen, "Large Language Models as Optimizers", arXiv:2309.03409, 2023.

[11] M. Sclar, Y. Choi, Y. Tsvetkov and A. Suhr, "Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design", ICLR 2024, arXiv:2310.11324.

[12] Y. Lu, M. Bartolo, A. Moore, S. Riedel and P. Stenetorp, "Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity", ACL 2022, arXiv:2104.08786.

[13] N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni and P. Liang, "Lost in the Middle: How Language Models Use Long Contexts", arXiv:2307.03172, 2023.

[14] B. Atil, S. Aykent, A. Chittams, L. Fu, R. J. Passonneau, E. Radcliffe, G. R. Rajagopal, A. Sloan, T. Tudrej, F. Ture, Z. Wu, L. Xu and B. Baldwin, "Non-Determinism of 'Deterministic' LLM Settings", arXiv:2408.04667, 2024.

[15] J. Yuan, H. Li, X. Ding, W. Xie, Y.-J. Li, W. Zhao, K. Wan, J. Shi, X. Hu and Z. Liu, "Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference", arXiv:2506.09501, 2025.

[16] M. Turpin, J. Michael, E. Perez and S. R. Bowman, "Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting", NeurIPS 2023, arXiv:2305.04388.

[17] L. Chen, M. Zaharia and J. Zou, "How is ChatGPT's behavior changing over time?", arXiv:2307.09009, 2023.

7. Download: the fifteen myths on one sheet

Fifteen claims about prompting, each with a verdict and the reason. False means the measurements contradict it; partly means the effect is small or unstable.

1. Avoid hallucinations. Do not guess. Verdict: false. The model does not know the term and keeps computing probabilities regardless.

2. Never use X. Verdict: false. Negation is the weakest instruction form for a model. State what should hold instead.

3. Critique your own draft. Verdict: false. Drafting and judging run in the same computation. An outside checker is a different thing.

4. You are a world-class senior architect. Verdict: false. A role as costume changes the tone, not the correctness.

5. I will tip you 200 dollars. Verdict: partly. The usual story does not hold. Measured effects are small and unstable.

6. Please and thank you cannot hurt. Verdict: partly. Two studies, two results. Politeness is not a reliable lever.

7. Take a deep breath and think logically. Verdict: partly. The most famous incantation is the most often misquoted. It came out of a search, not a rule.

8. Format and markup do not matter. Verdict: false. Change only the format and keep the content: the results shift measurably.

9. The order of examples does not matter. Verdict: false. Swap only the order and nothing else: the results change.

10. Pack every rule into one long prompt. Verdict: false. Where a fact sits matters. The middle is what gets lost.

11. Huge window, so pour everything in. Verdict: false. Holding is not finding. A large window does not search by itself.

12. Temperature at zero means reproducible. Verdict: false. The dial switches off one source of randomness. The hardware brings a second.

13. Written-out steps mean a right answer. Verdict: false. The supplied reasoning often does not name the real reason for the answer.

14. A pack of 500 ready-made prompts. Verdict: false. What you buy is tied to a model and a date. Both change.

15. Once the prompt is right, the product stands. Verdict: false. Same model, same tasks, three months later: different results.

To take awayPDF

Download the sheet as a PDF

Download