I set up Claude Code the way Spotify does to save tokens, and it cost me more

I set up Claude Code the way Spotify does to save tokens, and it cost me more

Published Oct 3, 2026, 7:00 PM EDT Maker, meme-r, and unabashed geek, Joe has been writing about technology since starting his career in 2018 at KnowTechie. He's covered everything from Apple to apps and crowdfunding and loves getting to the bottom of complicated topics. In that time, he's also written for SlashGear and numerous corporate clients before finding his home at XDA in the spring of 2023. He was the kid who took apart every toy to see how it worked, even if it didn't exactly go back together afterward. That's given him a solid background for explaining how complex systems work together, and he promises he's gotten better at the putting things back together stage since then. One thing I'm always aware of when handling agentic tasks is my token burn. Simple tasks like generating test files from a template or skimming through a folder of PDF documents to answer a question are almost insulting to the power of a frontier model. That's less of a concern with subscription pricing, but with API billing, token costs add up. Unnecessarily, when a cheaper model (or even a local one) could handle the task. One of Spotify's engineers fixed this issue by creating two new AiKA Modes in Portal, the managed developer platform the company built internally and later released as a product. Now, I don't have access to that platform, but what I do have is Claude Code and access to a variety of other LLMs. I rebuilt the two modes in my own setup to see what it did to my token burn, but it didn't go exactly as I expected. Spotify's shunt plugin hands Claude Code's grunt work to a cheaper model Getting it running broke a few things Shunt's pitch is simple: most of what a coding agent does isn't thinking; it's reading. When Claude Code tries to read a file over 350 lines, a hook blocks it and points Claude at a bulk-reader skill, which sends the files to a cheaper worker model and returns a short summary. A second skill, code-writer, does the same for boilerplate like test files, writing the result straight to disk, so Claude never generates it. Spotify's numbers are impressive. Against a 162,000-line Java monorepo, the shunt README reports 82% to 94% savings on large reads, with a mean of 90%, using Gemini 2.5 Flash as the worker. The catch is that the worker lives behind Spotify Portal, the company's commercial developer platform, and getting access means talking to sales. Luckily, the plugin is open source, and the Portal dependency lives in exactly one file. I replaced it with a version that sends the same requests to Lemonade Server running Qwen3-Coder-30B-A3B or to Claude Haiku through a headless Claude Code session. Haiku is the closest match to Spotify's setup, and the local model is the one you can run for free.​​​​​​​ Three things broke before I measured a single token, and each is worth knowing. First, shunt's hooks don't work on the current Claude Code. They tell Claude Code to "allow" a command, a value the hook format no longer accepts, so every allowed command threw an error until I made them return nothing instead. Second, my local model was quietly reading garbage. On Lemonade's ROCm backend, Qwen3-Coder invented function names that don't exist, and when I asked it to copy the first five lines of a file, it returned an opening code fence and stopped. The same model on the Vulkan backend copied those lines character for character, so test whether your worker can copy its input before you blame the model. Third, headless Claude Code is an agent, not a text box. When asked to write a test file, Haiku tried to write it itself, hit a permission prompt it couldn't answer, and the plugin dutifully saved its permission request as the "code." Launching it with --tools "" strips its tools, and that fixed it. Opus 5.5 already reads files the way that shunt wants it to And the hooks shunt uses aren't watching for the right thing on my PC My test bed was Agency, an open-source dashboard for managing a team of AI agents that I've written about before. Claude Code ran Opus 5.5 throughout, so the shunt-on and shunt-off runs differed in exactly one way. I measured the Messages figure from Claude Code's /context command, because the total includes about 37,500 tokens of fixed overhead every session pays before you type a word. The big target was agency/app.py, all 3,079 lines of it. My first task asked Claude to list every HTTP route in that file, and that's where the plan started to wobble. Opus answered with two grep calls and never opened the file, spending 17,300 tokens on a correct table of 62 routes (before cheerfully summarizing it as 76, but I'll let that slide). A question grep can answer gives shunt nothing to do.​​​​​​​ So I asked for something grep can't answer: how the file is organized, which helpers the routes depend on, and how they find their data on disk. With shunt off, that took 36,700 tokens; with shunt on, 35,500. That difference is noise, because Opus did exactly the same thing both times: it mapped the file with grep, then pulled about 1,000 of the 3,079 lines in targeted chunks with sed.​​​​​​​ Shunt's hooks only intercept Claude Code's Read tool and the cat, head, tail, less, and more commands. Opus didn't use any of them, so the bulk-reader skill sat in its context the whole time without ever being called. The plugin was installed, working, and completely ignored.​​​​​​​ The guard has its own gap, too. A chained command that read a small file and then a 439-line one slipped straight through, because the hook only checked the first file. Neither is a deal-breaker, but Spotify's 90% clearly depends on how your Claude reads, and mine had already taken the hint. Delegating the writing cost more than doing it myself Opus wrote great specs, and the workers still broke them Reading didn't work out so well, so the real test was writing. I asked Claude Code to generate tests for the CLI commands the project's existing test file doesn't cover, matching its style, and to use a delegation skill if one was available. Opus loaded shunt's code-writer skill on its own, which felt like a win for about ten minutes. Setup Messages tokens Time Worker's draft as delivered Final tests passing Opus 5.5 alone 31,900 About 1 minute Not applicable 16 of 16 Opus 5.5 with Qwen3-Coder (local) 39,300 10 minutes, 3 seconds 2 of 21 passing 21 of 21 Opus 5.5 with Claude Haiku 42,800 7 minutes, 14 seconds 19 of 23 passing 23 of 23 Both delegated runs used more tokens than Opus working alone, 23% more with Qwen3-Coder and 34% more with Haiku, and took seven to ten times as long. Sure, both ended up with a few more tests, but those came from Opus's planning, and Opus could have written them itself in a fraction of the time. The frustrating part is that Opus did its half of the job brilliantly. Before delegating, it worked out that the CLI looks for its config in the current working directory, that test data needed recent dates to survive a cleanup rule, and that one command takes its answers from standard input. Then it handed the local worker a 635-word spec spelling all of that out. Qwen3-Coder ignored the important parts. At least one of its tests ran the CLI from the project folder instead of the temporary one, the exact trap the spec warned about, and 19 of 21 tests failed. Opus read all 339 lines of the draft and rewrote the file from scratch, which meant it generated the whole thing anyway. Haiku came much closer, with 19 of 23 tests passing on delivery, but it hard-coded a fixed date into its test data, the cleanup-rule trap the spec had flagged. Opus read nearly 400 lines of the draft and patched it rather than starting over. That's the cost shunt can't remove: Opus won't ship code it hasn't checked, and checking a draft means reading it. Waiting on a worker at 12 tokens per second Then there's the waiting. On Vulkan, Qwen3-Coder managed about 12 tokens per second, so a test file of a few thousand tokens takes four minutes or more. The first delegation timed out at the three-minute mark while the model was still writing, and Opus had to retry with a longer timeout. Haiku isn't instant either, because each delegation starts a whole Claude Code session in the background. On a tiny file-reading test, it took 15.8 seconds, compared with the local model's 3.2, though it got every line number right where Qwen3-Coder got one of three. Spotify warns about latency in its own post, and on this writing task it turned a one-minute job into a ten-minute one. Spotify's idea is sound, but my Claude was already doing the cheap part I don't think Spotify's numbers are wrong. Its benchmarks came from a big Java monorepo and a setup where Claude was evidently reading large files whole, and the post doesn't say which Claude model it measured. If your sessions work like that, shunt's hooks will fire constantly, and the savings will be real. But Opus 5.5 already reads selectively, and when it delegates writing, it checks the work, which costs more than writing it in the first place. To find out which camp you're in, run /context after a big task and note the Messages line, then check whether Claude actually called Read on your large files. If it didn't, the cheapest token-saving plugin is the one you don't install. Claude is a suite of LLMs and tooling to help you get work done.

Original Source

Read the full article at Xda-developers →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.