Claude Code Prompt Caching: The 20x Rule That Decides Your Bill
Claude Code prompt caching decides if your next message costs cents or dollars. Here's the 20x math, the one-hour clock, and three ways to reset clean.

Your last Claude Code message may have cost 20x more than the one before it. Not because of what you typed — because of what expired while you stepped away.
Prompt caching is the single largest lever on a Claude Code bill, and it is almost entirely invisible. There is no warning when it lapses, no line in the interface that says the next message will be expensive, and no obvious relationship between what you did and what you paid.
Here is the full breakdown, concept by concept — what you are actually billed for, why the cache exists, and the three clean ways to reset when it is gone.
The short version
- Tokens are the bill. Every reply resends the whole thread.
- The cache is the discount — reading it costs 0.1x, rewriting it costs 2x. That gap is the 20x.
- The clock is the only real risk. Claude Code holds a one-hour cache; going quiet longer kills it.
1. Tokens are the currency
One word is roughly one token, and nothing you send is free — including the parts you did not consciously send.
The asymmetry that matters: output costs 5x input. On Opus that is $5 in and $25 out per million tokens. So a verbose response is five times more expensive per token than the context that produced it, which is a good reason to ask for concise output on routine work.
But output is the smaller problem, because output is bounded by what you asked for. Input is not.
2. Every reply resends the whole thread
This is the fact that surprises people, and everything downstream follows from it.
The API is stateless. It remembers nothing between messages. There is no session on the other end holding your conversation. Each request has to carry everything the model needs to answer.
Which means your 200th message re-sends messages 1 through 199 as fresh input. Not a summary of them — all of them, in full, every single time.
A long working session therefore has a cost curve that rises with every exchange, regardless of how short your latest message was. Typing “yes” at message 200 is not a cheap request.
3. Prompt caching is the discount
Caching exists precisely to defuse that. Rather than reprocessing the entire thread at full price, Claude reads a held copy of your thread at one tenth the input price.
Writing that copy is not free — it costs 1.25x the normal input rate, or 2x for the one-hour version. But you pay it once and then read cheaply for as long as the copy lives.
So the economics of a long session are: one expensive write, then many very cheap reads. That is a good trade, and it holds right up until the copy disappears.
4. The cache has a clock
Claude Code holds a one-hour cache. Go quiet longer than that and it dies, silently.
Time is not the only thing that kills it. Switching your model, your tools, or your settings kills it instantly — because the cached copy was written against a specific configuration, and changing the configuration invalidates it.
This is worth internalising, because it makes some very ordinary behaviour expensive. Swapping models mid-session to “save money” on a simple question throws away the cache and costs you the rebuild. So does adding a tool halfway through. So does lunch.
5. Why it is exactly 20x
The arithmetic is short and worth doing once.
- Cache alive: you pay 0.1x to read it.
- Cache dead: you pay 2x to write it again.
2 divided by 0.1 is 20. That gap is the 20x, and it applies to the entire thread, not to your last message.
Made concrete: a 500K-token thread costs about 50 cents to read while the cache is alive, and about 10 dollars to rebuild once it is not. Same conversation, same next question, twenty times the price — decided entirely by whether you were away for fifty minutes or seventy.
6. Three ways to reset clean
When the cache is gone, the cheapest move is usually not to resume the old thread but to start a clean one. Three ways to do that:
- /clear wipes the thread entirely. Best when your files already hold the context — if the state lives in the repository, the conversation history is redundant and you are paying to carry it.
- /compact writes a summary and restarts you from it. The middle option: you keep the conclusions and drop the transcript that produced them.
- A handoff file parks context outside the chat entirely, to be reloaded on demand. The most deliberate option, and the one that survives across days rather than hours.
The general principle: context that lives in files is free to re-read and can be loaded selectively. Context that lives in a thread has to be carried in full, forever.
7. Route the work, then maintain it
Two habits that keep the bill down structurally rather than reactively.
Let a cheap model execute and a smarter one advise. Each keeps its own cache, so the split does not cost you a rebuild — and the expensive model is only carrying the context it needs for the advice, not the full working transcript.
Run /doctor periodically to trim a bloated CLAUDE.md, plus stale skills and memory files. Everything in those files is prepended to every single request, which means a neglected CLAUDE.md is a permanent tax on every message you will ever send. It is the one place where a five-minute cleanup pays out indefinitely.
The master formula
- Tokens are the bill.
- Every reply resends the thread.
- The cache is the discount.
- The clock is the only real risk.
Frequently asked questions
What is prompt caching in Claude Code?
It is a held copy of your conversation that Claude can read at one tenth the normal input price. Writing the copy costs 1.25x, or 2x for the one-hour version, so a long session is one expensive write followed by many cheap reads.
Why does the cost jump 20x?
Because reading a live cache costs 0.1x and rewriting a dead one costs 2x, and 2 divided by 0.1 is 20. The gap applies to the whole thread: a 500K-token conversation is roughly 50 cents to read and 10 dollars to rebuild.
What kills the Claude Code cache?
Time and configuration. Claude Code holds a one-hour cache, so going quiet for longer kills it. Switching your model, your tools or your settings kills it instantly, because the cached copy was written against a specific configuration.
How do you reset a session cheaply?
Three ways. /clear wipes the thread, which is right when your files already hold the context. /compact writes a summary and restarts from it. A handoff file parks context outside the chat and reloads it on demand.
Does a long CLAUDE.md cost money?
Yes, on every request — its contents are prepended to everything you send. Running /doctor to trim a bloated CLAUDE.md along with stale skills and memory files is one of the few cleanups that keeps paying indefinitely.
Build it yourself
Everything written about here gets built in the open — the whole application, on camera, including the parts that did not work first time.
Keep reading
AI agent skillsAI Agent Skills: 9 Free Ones and the Step Most Skip
AI agent skills install from one pasted link — that part is easy. The step that makes them genuinely useful is the one almost nobody runs.
Read article→
Jev decision modelJev Decision Model: The AI That Can't Write, Only Decides
The Jev decision model can't write a word. Vercel and LangChain added it anyway. Why a "System One" classifier is 200x faster than the LLM in your agent.
Read article→
open source SaaS alternativesOpen Source SaaS Alternatives: 5 Repos Worth Selling
Semrush charges $250. This costs ten.
Read article→