Claude Code now shows exactly why it burned your tokens, and I finally stopped guessing

Claude Code now shows exactly why it burned your tokens, and I finally stopped guessing

Published Oct 5, 2026, 3:30 PM EDT Anurag is an experienced journalist and author who’s been covering tech for the past 5 years, with a focus on Windows, Android, and Apple. He’s written for sites like Android Police, Neowin, Dexerto, and MakeTechEasier. Anurag’s always pumped about tech and loves getting his hands on the latest gadgets. When he's not procrastinating, you’ll probably find him catching the newest movies in theaters or scrolling through Twitter from his bed. I leave Claude Code conversations open when I take a break and return to the same session when I’m ready to continue. The conversation still contains the files, instructions, and decisions relevant to the task, so I can pick up where I left off without explaining everything again. However, thanks to Claude Code’s new cost command, I found out how many tokens this one habit has been costing me. In case you don’t know, Claude Code’s /cost command now gives you more information about why it wasted your tokens. Alongside token usage, it shows cache misses, how much context needs to be cached again, and likely causes when it can identify them. For me, walking away is the biggest problem. A long enough break lets the prompt cache expire, making the next message more expensive even when I only ask Claude to continue. The cost command actually tells you a lot Including a breakdown by model covering input, output, cache reads, and cache writes Running the cost command brings up the session’s token usage, with a breakdown by model covering input, output, cache reads, and cache writes. It also shows a dollar estimate, but on a Pro or Max subscription, that figure doesn’t represent an additional charge. I’m more interested in how much work goes into a task, especially when the amount of code Claude produces doesn’t explain the usage. You can check the Prompt cache (main) line for more details. It shows the share of input served from cache, the number of cache misses, and how many tokens need to be cached again. It also reports whether the cache is warm or cold and, when possible, identifies a likely cause for the last miss. If you are a subscriber, the usage breakdown also attributes recent consumption to skills, subagents, plugins, and individual MCP servers. That’s useful when Claude delegates part of a task, since those agents make their own requests and consume usage alongside the main conversation. You can switch between the last day and week to inspect a longer period. Walking away costs you a lot of tokens Because Claude Code decides to go on a loop many times I give Claude Code a task and step away from my computer as it works through the changes. It can read files, edit code, and run tests without me directing every action, which is useful for time-consuming jobs. However, if it keeps trying fixes that don’t work or repeatedly searches through the same files, I’m not there to interrupt it and suggest another approach. Take a task like fixing TypeScript errors across an app. Claude can change one file, run the build, and discover errors elsewhere. Each attempt can legitimately uncover another issue. However, if the same error keeps returning and Claude keeps making similar changes, I need to question its approach. When I’m at the computer, I can interrupt, point it toward a relevant file, or ask it to explain its diagnosis before making another edit. Walking away removes that opportunity until I return and review what it does. Those repeated attempts also add more code and command output to the conversation. Claude Code sends that growing context with subsequent requests, so the cost includes processing earlier material as well as generating the next fix. Claude Code uses prompt caching to make repeated input cheaper, but cached reads still consume usage. Add clearer limits before stepping away Give Claude Code a more defined prompt If, like me, you also step away after giving Claude Code a task, give it a task with a clear stopping point. For example, “fix all the errors in this app” gives it room to keep investigating whatever the next build throws up. I’d rather use a more specific request that names the affected feature and the test that should pass. You can also ask it to stop and explain its findings if the same error survives two attempts. For larger changes, I like to use plan mode, which lets you review the proposed approach before Claude starts editing. You can catch an unnecessary refactor or point out a requirement it misses before implementation consumes more usage. You can also give Claude Code an expected result for a better output. It could be a particular test that needs to pass, or a bug needs to stop occurring under the conditions you describe. If you keep returning to repeated unsuccessful attempts, interrupt and ask Claude to explain what each attempt establishes. You can use the rewind command to return to an earlier conversation or code checkpoint when you want to abandon that approach. I’d rather not keep building on edits that leave the original problem unresolved. You also have the accumulated context to consider before starting another task. Use the clear command to start a fresh conversation for unrelated work, and the compact command to summarize a session you need to continue. But note that compaction itself consumes tokens, so running it repeatedly adds overhead. Make Claude Code more efficient While the cost command tells you where Claude Code has wasted your tokens, you can start solving the problem by avoiding that waste in the first place. What really helped me was to stop walking away and start watching what Claude was doing so I could interrupt it when it got stuck in a loop. I’ve also had success by setting a clear goal for Claude and being precise with my prompt from the start. There are also several settings you can change to make Claude work more efficiently.

Original Source

Read the full article at Xda-developers →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.