Why Claude burns tokens so fast: the seven real reasons Claude uses so many tokens, from Fable 5.1's heavier draw and long conversations to always-on thinking, Claude Code exploration, cache misses, background agents and peak hours.

Why does Claude burn tokens so fast? The 7 real reasons

Claude burns tokens fast because every message re-reads the whole conversation, the newest models always think before they answer, and Fable 5.1 costs far more per token than Opus or Sonnet. Add Claude Code's file reading, cache misses after breaks, agents working in the background and peak hours, and a plan that looked generous empties in an afternoon.

The short answer

Direct answer

Your usage is measured in tokens processed, not messages sent. Anthropic lists the drivers itself: message length, attachments, how long the conversation already is, which model answers, the effort level, and tools such as Research and web search. A short question can be a very large request.

The seven that matter most in October 2026: Fable 5.1, long conversations, always-on thinking, Claude Code exploration, cache misses, background agents, and peak hours plus a smaller weekly pool. Each one has a fix, and most of the fixes are free.

How Claude counts what you use

Claude plans do not give you a fixed number of messages. They give you a budget that resets on two clocks: a session window of five hours and a weekly limit. Anthropic's article on understanding usage and length limits lists what moves the needle: conversation length and complexity, message length, the features you use, the model you pick, the effort level, and the files in your projects. It also says something people miss: your usage on claude.ai, Claude Code and Claude Desktop all counts towards the same limit.

So "Claude burns tokens fast" really means one of two things. Either each request is bigger than you think, or more requests are being sent than you think. Every reason below is one or the other. If you want the mechanics of the two clocks first, our guide to how Claude's usage limits really work covers them in detail.

The 7 reasons Claude uses so many tokens

# Reason Why it costs The fix
1Fable 5.1Highest price per token, High effort by default in Claude CodeSave it for the hard problems
2Long conversationsEvery turn re-reads the whole threadNew chat when the topic changes
3Always-on thinkingThinking is billed as output and cannot be turned off on the top modelsLower the effort level for routine work
4Claude Code explorationFile reads and tool calls, each a new requestSpecific prompts, plan mode
5Cache missesAfter a break the full context is processed againClear or start fresh instead of resuming cold
6Background workSubagents, agent teams and scheduled tasks spend while you watch something elseSmaller models for helpers, shut teams down
7Peak hours and a smaller weekFaster five-hour drain in chat on weekday mornings (Pacific), and a weekly pool cut on 14 SeptemberMove heavy chat work off-peak, watch the weekly number

1. Fable 5.1 spends your limit faster than any other model

This is the newest reason and, for many people this autumn, the biggest. Fable 5.1 shipped on 1 September 2026 as Anthropic's most capable generally available model. Anthropic's Help Center is blunt about the cost: Fable models draw from your plan's regular weekly usage limits and use them faster than other Claude models.

Anthropic does not publish how plan usage is weighted per model, but its API prices are the clearest public signal of relative cost, and the gap is wide:

Model Input, per million tokens Output, per million tokens Relative to Sonnet 5.5
Fable 5.1$10$505x
Opus 5.5$4$202x
Sonnet 5.5$2$101x

Prices are from Anthropic's launch pages for Fable 5.1, Opus 5.5 and Sonnet 5.5. The same Fable page adds a second multiplier: Fable 5.1 defaults to High effort in Claude Code, and to Medium on claude.ai. So the same task in Claude Code both costs more per token and tends to use more tokens.

How that lands depends on your plan. On Max, Fable can use up to 50% of your weekly limit at no extra cost, then switches to usage credits; we explain that ceiling in Fable's 50% weekly cap on Claude Max. On Pro, Anthropic's pricing page lists Fable under usage credits, so it never comes out of the plan at all; what Claude Pro gets from Fable 5.1 has the details.

The fix: treat Fable as the specialist, not the default. Sonnet 5.5 is, per Anthropic, 30% faster than Sonnet 5 and needs fewer tokens per task. Opus 5.5 got cheaper too, and its launch came with higher five-hour limits on Pro, Max, Team and seat-based Enterprise. Most everyday work does not need the most expensive model in the line-up.

2. Long conversations re-read themselves on every turn

This is the classic one, and it has not gone away. Claude has no memory between requests in the way people imagine. Each time you send a message, the conversation so far goes back to the model with it. Anthropic's Claude Code documentation spells it out: Claude Code sends your full conversation with every request, so a one-line question in a session that has been open all day still draws usage for the whole conversation.

Prompt caching softens this. Content that was just processed can be re-read at a much lower rate, and content stored in a Project is cached and, in Anthropic's words, counts less against your limits when reused. But cheaper is not free. A thread that started with a 40-page PDF is still carrying that PDF on message thirty.

Attachments, Research and web search make each turn heavier still. Anthropic's usage limit best practices name file size, tool use and artifact creation as factors alongside conversation length.

The fix: start a new conversation when the topic changes, and put documents you reuse into a Project rather than re-uploading them. In Claude Code, /clear between unrelated tasks costs nothing. Our list of 13 ways to make a Claude session last longer goes through the rest of these habits.

3. Thinking is always on, and it counts as output

The current top models reason before they reply, and that reasoning is not free. Anthropic's cost guide says thinking tokens are billed as output tokens, and the default budget can be tens of thousands of tokens per request depending on the model. Output is the expensive side of every price in the table above.

On older models you could switch thinking off. On the current ones you cannot: the same guide states that you can't turn off thinking on Opus 5.5, Sonnet 5.5 or the Fable models, which always use extended thinking. What you control is the effort level.

The fix: in Claude Code, lower the effort with /effort or in /model for simple tasks. Renaming a variable does not need the same depth of reasoning as designing a database schema. In the claude.ai app, keep the heavier effort settings for questions that need them.

4. Claude Code explores before it acts

If your limit vanishes faster in Claude Code than in chat, this is why. A chat turn is one request. A Claude Code turn can be a dozen: read a file, search the project, run a command, read the output, edit, run the tests, read the failures. Each tool call sends another request carrying the conversation plus the new results.

Anthropic's guidance for teams budgeting Claude Code puts it plainly: each Claude Code turn carries file contents, tool calls and multi-step reasoning, so one debugging session can consume more than a day of chat. And the size of the job depends heavily on how you ask. The docs contrast "improve this codebase", which triggers broad scanning, with "add input validation to the login function in auth.ts", which lets Claude work with minimal file reads.

The fix: name the file and the function. Use plan mode (Shift+Tab) for bigger changes so Claude proposes an approach before it starts editing, and press Escape early if it heads the wrong way. Keep your CLAUDE.md short, since it loads into every session, and move specialist instructions into skills that load only when needed. Disable MCP servers you are not using.

Catch the burn while it happens

See your Claude session and weekly limits drop in real time, from the Mac menu bar.

Most of these reasons are invisible until the limit message appears. tokn.watch shows your remaining Claude session and weekly percentages with exact reset timers, plus the Fable row on Max, so a runaway session shows up while you can still stop it. Same for ChatGPT if you use both. Free, no account needed.

GET IT FREE

5. Coming back after a break costs more than you expect

Prompt caching is what keeps long sessions affordable, and it expires. Per Anthropic's cost guide, the cache lifetime is one hour on a subscription and drops to five minutes once you are drawing on usage credits. Your first message after a longer break misses the cache and processes your full context again at the full rate.

That is why a session that felt cheap all morning can take a visible bite out of your limit with a single message after lunch. Compacting has a related cost: /compact reads the whole conversation it summarises, so compacting a large context is itself a large request. When you want a fresh start rather than continuity, /clear is the free option.

The fix: if you are returning to a big session after a long break, decide whether you really need all of that history. On Pro and Max, Claude Code offers to resume from a summary so later requests do not carry the full history. Accept it when the details no longer matter.

6. Work you are not watching is still spending

Some usage happens while your attention is elsewhere. Anthropic's list of reasons usage climbs in a long session includes several of these:

  • Subagents. Every subagent sends its own requests on top of the main conversation. They are worth it for keeping noisy output out of your main thread, but they are not free.
  • Agent teams. Anthropic's figure is that agent teams use approximately 7x more tokens than standard sessions when teammates run in plan mode, because each teammate keeps its own context window. Each active teammate keeps consuming until it exits.
  • Scheduled tasks and loops. A scheduled task fires on its interval even while the session is idle, sending your full context each time.
  • Other surfaces. Because claude.ai, Claude Code and Claude Desktop share one limit, a long chat on your phone and an agent run on your laptop drain the same number.

The fix: give helpers a smaller model (Claude Code lets you set model: haiku for a subagent), keep teams small and shut them down when their work is done. On a paid plan, the /usage screen in Claude Code now breaks recent usage down by skills, subagents, plugins and MCP servers, and flags any behaviour that accounts for 10% or more of it. It is the quickest way to find out which of these is yours.

7. Peak hours, and a smaller weekly pool

The last reason is not about how you work but about when, and about changes on Anthropic's side.

Peak hours. In late March 2026, Anthropic confirmed it had made five-hour session limits move faster during weekday peak hours, 5am to 11am Pacific, for Free, Pro and Max, while weekly limits stayed the same. It said about 7% of users would hit session limits they would not have hit before, particularly on Pro. On 6 May 2026, Anthropic removed the peak-hour reduction for Claude Code on Pro and Max, and doubled Claude Code's five-hour limits. That announcement covered Claude Code only. For heavy chat work, the weekday morning window in Pacific time (afternoon in Europe) is still the one to avoid.

A smaller week. If Claude has felt hungrier since mid-September, it is not your imagination. On 14 September 2026 Anthropic ended the temporary 50% weekly boost it had run since May and replaced it with a permanent 25% increase over the old baseline. Anthropic acknowledged the arithmetic: "Compared to today, this works out to a 17% reduction in weekly limits on Claude Code." The same work now takes a bigger share of the week. Our explainer on Claude's weekly limit covers how that pool resets.

The fix: schedule long agent runs and big document jobs outside the peak window if you can, and look at the weekly number, not just the session one. The five-hour session limit resets several times a day; the weekly one does not.

What is usually not the reason

Two explanations come up constantly and rarely hold up. The first is "Claude is broken". Sudden jumps almost always trace back to one of the seven above, most often a long thread, a cache miss or a model switch. The second is "a short message should cost less". As reason 2 shows, the size of a request is set mostly by everything that came before it, not by what you just typed.

One thing that is real but small: Claude Code uses a few tokens in the background even when idle, for summarising past sessions and similar housekeeping. Anthropic puts that at typically under $0.04 per session. It will not empty a plan.

How to see which reason is yours

On claude.ai, Settings then Usage shows your current session and weekly readings, and on Max a separate Fable row. In Claude Code, /usage adds the breakdown by subagent, plugin and MCP server, plus the behaviour flags for long context and cache misses. Between them they answer the question. The catch is that you only see them when you stop and look, which is usually after the expensive part.

That is the gap tokn.watch fills. A remaining percentage in the menu bar that drops 8 points on one message tells you, in the moment, that you just paid for a long context or a cold cache. Connecting takes about a minute, and the setup steps are on our homepage. If you are already out, what to do when you hit your Claude limit covers the next steps, and if you are weighing a second subscription to cover the gaps, should you use Claude or ChatGPT is the honest version of that decision.

The verdict

Claude does not burn tokens for no reason. It burns them because the expensive things are the default things: the newest model is the most expensive, the conversation you are already in is the one you keep typing into, thinking is always on, and Claude Code explores before it edits. None of that is a bug, and all of it is adjustable.

If you change only three habits, make them these: use Fable for the hard problems only, start a new conversation when the topic changes, and give Claude Code specific instructions. Then keep both numbers, session and weekly, where you can see them.

Frequently asked questions

Why does Claude use so many tokens?

Mostly because every new message sends the whole conversation back to the model, so a long thread costs more per turn than a fresh one. On top of that, Opus 5.5, Sonnet 5.5 and the Fable models always think before they answer, Fable 5.1 costs far more per token than the others, and Claude Code reads files and runs tools on every step. A short question late in a long session is never a short request.

Why does Fable 5.1 hit my Claude limit so fast?

Anthropic says Fable models draw from your plan's regular weekly usage limits and use them faster than other Claude models. At API list prices Fable 5.1 costs $10 per million input tokens and $50 per million output tokens, against $4 and $20 for Opus 5.5 and $2 and $10 for Sonnet 5.5. It also defaults to High effort in Claude Code. On Max it can take at most half your weekly limit; on Pro it bills from usage credits.

Does a short message cost less in a long Claude conversation?

Not much. Claude re-reads the conversation so far every time you send something, so a one-line question at the end of a long thread draws usage for the whole thread. Prompt caching makes the re-read cheaper, but it is never free. Starting a new conversation when the topic changes is the single cheapest habit there is.

Can I turn off thinking to save Claude usage?

Not on the current top models. Anthropic's Claude Code documentation says you cannot turn off thinking on Opus 5.5, Sonnet 5.5 or the Fable models, which always use extended thinking. What you can change is the effort level, with /effort or in /model in Claude Code. Thinking tokens are billed as output tokens, so a lower effort on routine tasks saves real usage.

Does Claude Code use more of my limit than chat?

Usually, yes. Both draw from the same usage limit, but each Claude Code turn carries file contents, tool calls and multi-step reasoning, and every tool call is another request carrying the conversation. Anthropic's own guidance is that one debugging session can consume more than a day of chat. Vague prompts make it worse because they trigger broad scanning of your code.

Do peak hours still make Claude run out faster?

In chat, they did from March 2026: Anthropic made five-hour session limits move faster on weekdays between 5am and 11am Pacific for Free, Pro and Max, with weekly limits unchanged. On 6 May 2026 it removed that peak-hour reduction for Claude Code on Pro and Max. The May announcement covered Claude Code only, so heavy chat work is still best moved outside that window.

Why did my Claude limit start running out faster in September 2026?

Because the weekly pool got smaller. On 14 September 2026 Anthropic ended the temporary 50% weekly boost it had run since May and replaced it with a permanent 25% increase over the old baseline, which it acknowledged works out to a 17% reduction in weekly limits on Claude Code compared with the boosted period. The same work now uses a bigger share of the week.

Know before the limit message

Your Claude session and weekly limits, always in your menu bar.

tokn.watch shows your remaining Claude session and weekly percentages with exact reset timers, plus a separate Fable row on Max, so you can spot a token-hungry session while it is still running.

GET IT FREE macOS 13+ · Free on the Mac App Store · or download directly

Last updated: 2026-10-05. Usage drivers and the shared limit across claude.ai, Claude Code and Claude Desktop per Anthropic Help Center - Understanding usage and length limits and Usage limit best practices. Fable's faster draw and the Max 50% ceiling per Claude Fable models on your plan; Fable on Pro via usage credits per the Claude pricing page. Model prices, Fable 5.1's default effort, Sonnet 5.5's speed and the Opus 5.5 five-hour limit increase per Anthropic's launch pages for Fable 5.1, Opus 5.5 and Sonnet 5.5. Full-context requests, always-on thinking, cache lifetimes, compaction, subagents, the 7x agent-team figure, background usage and the /usage breakdown per Claude Code docs - Manage costs effectively. Peak hours per PCWorld and Anthropic's 6 May 2026 announcement. The 14 September 2026 weekly change per BleepingComputer. Plans, prices and limits change; confirm current figures on Anthropic's own pages.

A note on numbers: Anthropic does not publish how plan usage is weighted per model. The per-token prices above are API list prices, used here as the best public signal of relative cost, not as a statement of what a plan charges.