Artificial intelligence (AI) is turning sharply cheaper to run on a per-token basis. Yet, the amount companies spend on AI is rising as they move from experiments and chatbots to more complex, multi-step applications.During a recent executive briefing, Coforge management said that over the previous seven months, the unit cost of tokens needed to deliver the same level of AI intelligence had fallen about 100-fold, while token consumption had grown about 8,000-fold.The company also said open-weight models accounted for about 35% of token consumption on the platforms it was observing, compared with about 11–12% in early 2025.There are broadly two types of models. Closed or proprietary models such as OpenAI’s GPT, Anthropic’s Claude and Google’s Gemini, keep their trained parameters under the developer’s control and are typically accessed through a hosted service or API.Open-weight models, such as Meta’s Llama, Google’s Gemma and Mistral’s open-weight models, make their trained parameters available for organisations to download and run themselves, subject to the model’s licence. Open-weight does not necessarily mean fully open-source. Their training data, training process and other components may remain undisclosed.Coforge also argued that the number of tokens used is not, by itself, the right way to assess the economics of AI. The company said the more meaningful measure is the “cost per unit task”. This means what it costs to complete a defined piece of work and what value that work creates.What is getting cheaper?Stanford’s 2025 AI Index, an annual report tracking major developments and trends in artificial intelligence, found that the cost of querying a model with GPT-3.5-level performance on a standard AI knowledge and reasoning benchmark fell from $20 per million tokens in November 2022 to $0.07 per million tokens by October 2024, a reduction of more than 280 times in about 18 months. It, drawing on analysis by Epoch AI, found that LLM inference prices had fallen between nine and 900 times per year depending on the task.The OECD (Organisation for Economic Co-operation and Development), an international policy and economic research organisation, found a similar trend and said that its July 2026 analysis estimated that quality-adjusted prices for text-to-text AI models fell by nearly 80% between January 2024 and April 2026.Why are AI workloads growing?The nature of the work being given to AI is changing at the same time. A chatbot might receive a question, process it and return an answer. An AI agent can break a task into several steps, reason through them, call external tools or systems, inspect the results, revise its approach and continue until the task is completed. Each additional step can involve further model calls and more tokens.The OECD notes that AI agents consume substantially more tokens per task and can increase model-use intensity by orders of magnitude. As a result, falling prices per token do not necessarily translate into lower costs for complex AI applications.Gartner has described this as the “Inference Paradox”. Its August 2026 analysis says routing a task to an agentic reasoning model can increase provider inference costs by at least five times compared with a basic chatbot interaction, and often by much more as complexity rises. Gartner forecasts that inference costs per agentic workflow will increase more than fivefold through 2028.How are AI agents changing the equation?A 2026 working paper by researchers from the University of Michigan, Stanford University, All Hands AI, Google DeepMind, Microsoft AI and MIT found that agentic coding tasks consumed about 1,000 times more tokens than the code-reasoning and code-chat tasks used for comparison. It also found that runs of the same task could differ by as much as 30 times in total token consumption.This greater intensity of use is showing up in enterprise spending. McKinsey’s May 2026 Enterprise AI FinOps survey, covering 75 qualified respondents across five industries, found that 93% reported exceeding their AI budgets. McKinsey also cited Menlo Ventures data showing that enterprise spending on large language models roughly tripled over the 12 months to the end of 2025.What are open-weight models changing?The growing use of open-weight models is changing the economics further. Instead of relying entirely on a frontier model accessed through a provider, companies can use different models for different tasks. Smaller or open-weight models can handle some workloads at lower cost, while more expensive frontier models can be reserved for tasks that require greater capability.Coforge said its clients were increasingly moving towards a mixture of models rather than relying on a single model provider. It also identified seven billion tokens of consumption as a “tipping point” in its experience at which clients start considering whether to consume and operate more of the AI stack themselves rather than simply subscribe to a frontier model. The company said that threshold could fall as open-weight and sovereign AI develop further.The broader competitive effect is also reflected in OECD research, which finds that competition and the availability of alternative models have contributed to falling AI prices.How are firms measuring AI costs?For businesses, the implication is that token price alone tells only part of the story. A useful measure is the economics of the complete task, like the model calls, tokens, computing resources, orchestration and human oversight required to produce a result, set against the value of that result.That is the logic behind Coforge’s emphasis on “cost per unit task”. Gartner’s research similarly points towards measuring token use against business value and the economics of the overall workflow rather than treating token prices as a proxy for the cost of AI.The result is a rapidly changing cost structure, like cheaper models are making more applications viable, while agents are making individual applications more computationally intensive.
The price of AI is falling; why are enterprises still spending more? Explained
Full Article
Original Source
Read the full article at Thehindu →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.