Zuzanna Stamirowska at Pathway discuss the causes of the explosion in AI token and explores what solutions are available to businesses to keep AI overheads down

The public discourse around AI is moving so fast that relatively recent headline discussions are already painfully outdated. Around May, reports that Meta employees were competing to be the company’s number one AI token user via a gamified leaderboard caught global attention. Amazon employees admitted to the Financial Times that colleagues were using AI needlessly to increase their consumption of tokens amid manager pressure to use AI frequently.
Good luck finding similar examples of ‘Tokenmaxxing’ behaviour in any business today. The costs associated with AI are growing faster than many organisations ever anticipated. That’s been down to a foundational misunderstanding: the assumption that falling raw token unit prices would decrease overall AI costs over time. In reality, the upward pressure on token consumption as AI systems have pivoted from simple text generation to both more expensive reasoning models, and agentic workflows have totally overwhelmed any downward pressure on token costs.
There’s no doubt the businesses are recognising these costs and mounting a response (see the new wave of enthusiasm for open weight models and model routing). What is in question is whether that response goes far enough. The hard truth is we’re locked into an unsuitable model of AI economics which sees any gains in intelligence come at a higher price. That model isn’t inevitable, if we’re brave enough to seek out another way forward.
One of the basic explanations as to why AI is getting more expensive is the aforementioned pivot to more token-consuming models and agentic systems. This speaks to something more systemic about the architecture of current AI models. They use dense, Transformer-based architectures that break problems down into language, generating ever-longer chains of text to reason their way to an answer. In modern reasoning models, this generation of text happens behind the scenes – a technique called chain-of-thought (CoT) reasoning where models generate masses of ‘thinking’ tokens before committing to a result.
Higher token generation naturally sets the stage for AI cost spikes. But the accelerant has been the introduction of new features and engineering requirements to deliver capability gains - notably, the recent shift to agentic systems.
But as there’s been no step change in the underlying technology (the transformer), these innovations have amounted to additional layers that increase overall token usage. In the case of agentic workflows, for example, the workaround to give agents context memory - crucial to deliver on their autonomous promise but not compatible with the disconnection between short-term and long-term memory has been to externalise memory management. This, in essence, requires agents to constantly re-read past interactions and summaries (burning tokens each time) rather than learning once from them.
It’s perhaps obvious that this situation isn’t sustainable, especially for enterprises looking to scale the deployment of agents beyond limited pilots. There’s been a subsequent doubling down on strategies to control AI expenditure recently, but I fear they don’t go far enough.
Specifically, model routing has become a major point of discussion. The concept of directing workloads to different models based on complexity has major appeal. Especially given another recent development – the launch of cutting-edge open-weight models out of China with cheaper inference and hosting costs, including much-discussed Kimi K3 – arguably further commoditising AI access, with many workloads now able to run on cheaper open-weight models with limited performance trade-offs.
A selective approach to how models are used for work is a sensible development in enterprise AI. But, in and of itself, it does little to address the underlying challenges if the models in rotation remain dependent on the transformer architecture and its deficiencies that drive up token consumption.
The other problem is that enterprise AI use cases can only become commoditised to a point. There’s always going to be a grouping of high-value use cases, related to critical workflows, that require deep, proprietary context and continuous learning. I’d argue that current frontier models aren’t capable of this work. So there’s a clear rationale to strive for more.
The good news for businesses is that work to build the successor to the transformer architecture is underway. One of the most significant areas of focus is reducing the need for AI systems to repeatedly process and regenerate context. Rather than effectively starting from scratch during each interaction cycle, researchers are exploring architectures that allow models to retain and adapt knowledge more efficiently over time. This was our focus in developing the BDH architecture.
The objective is to build models that are flexible and able to retain information across interactions, as opposed to relearning each time – similar to how intelligence develops in the human brain. The goal is for AI to build understanding over time, adapting its behaviour based on experience rather than repeatedly reconstructing context.
In practical terms, this would move AI away from systems that primarily process information and towards systems that accumulate and internalise experience and expertise over time. The ambition is not simply to produce better responses, but to create models that become more capable as they interact with the world.
For businesses, this could have important implications. More efficient architectures have the potential to reduce operating costs and enable more sophisticated AI applications that demand continuity and long-term reasoning.
Growing token consumption is placing unsustainable pressure on the compute, infrastructure and resources that underpin today’s AI models. Workarounds without rethinking the fundamental, architectural drivers of that pressure will prove toothless. Business leaders would be savvy to track the development of post-transformer architectures and think again about how intelligence is delivered and at what cost.
Brute-force scaling is no longer the answer, and improved financial return on intelligence is within reach – if we’re brave enough to demand it.
Zuzanna Stamirowska is CEO at Pathway
Main image courtesy of iStockPhoto.com and AndreyPopov
Winston House, 3rd Floor,
Units 306-309, 2-4 Dollis park,
London, N3 1HF
020 8349 4363
© 2026, Lyonsdown Limited. teiss® is a registered trademark of Lyonsdown Ltd. VAT registration number: 830519543