BetteryieldsBetteryields
← Blog

What a pack of 500 prompts really costs

Dr. Aaron Hutzler · 16 August 2026 · 5 min

Eine Person öffnet einen gläsernen Tresor mit sichtbarer Schliessmechanik; im Tresor liegt ein Stapel Papiere, obenauf eines mit der Aufschrift CONFIDENTIAL
This image was generated with AI.

Somebody is selling you 500 ready-made for 99 dollars. The marketing copy writes itself. Months of meticulous tuning, bundled together, in your possession by this afternoon.

The 99 dollars are not the price. They are merely the down payment.

Here is the rest of the bill, in four parts, each one measurable.

1. Part 1: the line item that repeats

A bought master prompt typically spans five to ten pages. Rules, examples, persona descriptions and a few pages of polite framing.

Ten pages of text is roughly 15,000 characters. As a rule of thumb four characters make one . That puts the instruction at about 3800 tokens. Providers bill input tokens.

At 10,000 customer requests a month that is 38 million tokens sent. Not one word of your actual document is in that number yet. Take your provider's price per million and multiply. The instruction is billed on request ten thousand exactly as heavily as on request one.

That is the part of the bill you can calculate today. The next three you cannot.

2. Part 2: the lock changes while you buy the key

A prompt is not a program. It works on one model in one version at one point in time.

A person stands small in front of an enormous polished steel vault door and holds up a golden key. The keyhole in the door runs and changes shape like liquid metal

Figure 1: The key you bought. The lock while it is being changed.

Image generated by AI

Three authors tested the same service three months apart on identical tasks. On one task accuracy fell from 84 percent to 51 percent. Nobody had touched the prompt. [1]

Two further studies say the same thing from two directions. A prompt format that performs well on one model correlates only weakly with the next. [2] A good ordering of examples for one model is worthless on another. [3]

So the pack you bought was tuned against a model that no longer exists in that form. It cannot be transferred to the one you would rather use either.

3. Part 3: what you bought, anyone can read out

Marketing materials call these prompts proprietary. Three authors measured whether that claim holds up. Across three sources of prompts and eleven models, simple text-based attacks recovered the hidden prompt with high probability. Their also separates a genuinely extracted prompt from an invented one. The result is therefore not an artefact. [4]

Nine authors went one step further and attacked real commercial prompt services. Their method infers what a prompt does from a very small number of inputs and outputs, then writes a prompt that reproduces the same behaviour. The work appeared at a security conference. [5]

A prompt is plain text sitting in front of a model that will talk about it. Whatever you paid for it, your users can obtain it. So can your competitor.

4. Part 4: the maintenance has no end

Every provider update, every migration to a cheaper model and every change to the surrounding application resets your calibration to zero. And a long prompt makes it worse in a way that is measured: seven authors showed that information in the middle of a long input is used markedly less reliably than information at the start or the end, even on models built for long inputs. [6]

The bigger the pack, the more there is to re-tune. And the more of it sits in the part of the input the model treats worst.

5. What to buy instead

Nothing. There is no artefact to buy here, only a way of working.

Keep the prompt short. Ask for one thing: the answer plus a verbatim quote from the source document for every claim.

Then let a program check it. It searches for each quote character by character in the source, on your own machine, without a model. Quote there and carrying the value: accepted. Quote missing: rejected. What the check settles is where a number came from. Not what it means.

This costs you a few hundred tokens per request instead of 3800. Which model produced the answer plays no part. You can switch providers on a Tuesday afternoon. And nobody can steal it from you, since nothing in it is secret. The check is twenty lines of code and the value sits in your documents.

A prompt is a key for exactly one lock. A check is a lock of your own.

6. References

[1] L. Chen, M. Zaharia and J. Zou, "How is ChatGPT's behavior changing over time?", arXiv:2307.09009, 2023.

[2] M. Sclar, Y. Choi, Y. Tsvetkov and A. Suhr, "Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design", ICLR 2024, arXiv:2310.11324.

[3] Y. Lu, M. Bartolo, A. Moore, S. Riedel and P. Stenetorp, "Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity", ACL 2022, arXiv:2104.08786.

[4] Y. Zhang, N. Carlini and D. Ippolito, "Effective Prompt Extraction from Language Models", arXiv:2307.06865, 2023.

[5] Y. Yang, C. Li, Q. Li, O. Ma, H. Wang, Z. Wang, Y. Gao, W. Chen and S. Ji, "PRSA: Prompt Stealing Attacks against Real-World Prompt Services", USENIX Security 2025, arXiv:2402.19200.

[6] N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni and P. Liang, "Lost in the Middle: How Language Models Use Long Contexts", arXiv:2307.03172, 2023.

Betteryields

Check AI answers against your own documents, with a source for every figure