When research shows $30-40 billion is being collectively poured into AI and only 5% of that investment generates any sort of return, that’s not just a problem – it’s a crisis.
A lot of it comes down to AI tokens: how many are being used, by who, and for what. In fact, better visibility into token spend is the one thing IT and finance teams say they want most right now. Sound familiar? Keep reading for strategies that control token spend and deliver measurable business value.
Let’s start with the fundamentals.
In cloud computing, you pay for servers, storage, or software licenses. With AI, you pay for tokens – small, model-readable pieces of text (or data) that AI processes to understand requests and generate responses.
The word “unbelievable,” for example, might be split into three tokens while a short word like “cat” is just one. This process is called tokenization, and it’s been standard practice in natural language processing for years. That changed in 2020 when OpenAI launched the GPT-3 API and began charging developers per token. An internal engineering concept became commercialized, and every major AI provider has billed that way since.
This is the model most companies use: pay a provider per token and let them worry about the infrastructure. Some companies, however, skip the provider entirely and run AI models on their own infrastructure – self-hosted models, private deployments, or systems that need tighter data control. When that’s the setup, there’s no per-token bill from OpenAI or Anthropic to review. The cost shows up differently, in GPU and cluster spend instead of a token invoice. Same economics, just one layer further down.
Every AI interaction costs money, and the unit that ties cost to value is the token. The task itself doesn’t determine the cost but rather the amount of information the AI has to process.
This is your everyday interaction with a conversational AI tool. “Rewrite this paragraph.” “Summarize this email.” “Give me five campaign ideas.” These requests are relatively inexpensive because the AI is processing a small amount of information and sending short responses.
That helps explain the whiplash in different projections we’re seeing. Within two months, Gartner separately predicted that 40% of enterprise applications would be integrated with task-specific AI agents by the end of 2026, and more than 40% of AI agent projects would be canceled by the end of 2027. The reasons? Unmanaged costs and unclear business value.
The writing’s on the wall. The organizations that succeed with AI will know exactly what they’re spending, on what, and why.
Not all AI agents are approved by IT. Research shows about 30% of employees create cost and security risks by using unsanctioned AI agents for work.
There are two kinds of tokens, each with its own cost: input tokens and output tokens. Input tokens represent the charge for what you send to AI: your prompt, your uploaded documents, your images or audio. Output tokens represent the charge for everything AI generates back to you: the 1,500-word article, the redesigned presentation, the webinar transcription. Output tokens tend to cost more because generating content is more computationally expensive than reading it.
Pricing depends on the provider.
Generally, input runs $1-$5 per 1M tokens and output runs $10-$30 per 1M tokens.
You ask AI to draft a proposal. It ends up being 1,500 words, equaling 2,000 output tokens. If you were using a model priced around $2.50 per million input tokens and $10 per million output tokens, that entire back-and-forth would cost roughly 2 cents total.
Try it, learn from it, see what works. But the numbers on where that experimentation lands are stark. All those pilots being pushed? Ninety-five percent are unsuccessful. Token competitions? A fast track to budget burnout. It doesn’t matter how many tokens you use or how much you track their usage if you can’t defend that usage with measurable, bottom-line impact.
As mentioned, for companies running models on their own infrastructure, the math moves. You’re the one paying for the GPUs directly, and how efficiently that hardware runs becomes the new version of cost per token.
First and foremost, cost management is a visibility problem. You know the money’s going out – you just don’t know where. What you need is hard data: exactly where your tokens are going, who’s spending them, and what you’re getting back for it.
If you go the manual route, you’ll have to pull invoices from each AI provider, do your best to map API keys or project IDs back to teams, tag spend by best guess (if you didn’t tag on day one), and reconcile it all in a spreadsheet. This repeats monthly, by hand, because none of it updates automatically.
The other option is to use a platform that sits between your teams and your AI providers, capturing and categorizing token spend as it happens. Every request gets tagged automatically – by team, model, and workflow – the moment it’s made, which is critical for controlling cost creep.
Usage from different providers gets normalized into one view, so you’re not stuck manually reconciling how OpenAI, Anthropic, and Azure each report cost differently. And because the platform watches in real-time, it can flag unsanctioned tools and runaway usage before they show up as a surprise on your next invoice.
Tagging is the practice of labeling every piece of cloud or AI usage with metadata – like team, project, or environment – the moment it’s created so costs can be allocated accurately. In our experience, a lot of companies don’t tag AI from the start. They don’t have time later, so the bill gets paid and the budget bloat grows.
It’s possible to go back and retroactively tag AI usage, but there are limitations and the time it would take is extensive. A more realistic option is a platform that reconstructs the best available picture of historical token spend by cross-referencing your existing billing data, API keys, and usage logs. It would also flag where shared keys or missing tags cause gaps and set up proper tagging for the future.
You can also have experts come in and do a deep cleanup. Tangoe’s team, for example, will go through months of untagged spend and set up tagging and allocation rules so you don’t end up back at square one.
Companies can achieve anywhere from 40-90% cost savings by taking steps like this as part of a comprehensive cost reduction strategy. Tangoe delivers both the platform and the support to make this happen through our technology consulting and advisory services.
Here’s a sobering number: 85% of employees don’t use AI for business value. In fact, less than 3% use it in ways that create meaningful ROI.
Having AI organize your inbox, summarize long email threads, draft replies, and flag the messages that need your attention.
It’s easy to see where the value of AI lives – the second use case – but research shows up to 35% of enterprise AI tokens are wasted on inefficient tasks like the first. Instead of asking how much AI was used, ask questions that draw a clear link to ROI. How much time was saved? How much work was accelerated? How much rework was avoided? Make it even easier with a simple litmus test: time saved vs. tokens spent.
But don’t forget that your metrics should evolve. Early on, when you’re just trying to prove that a tool works, hours saved is a fine place to start. As adoption scales, “hours saved” doesn’t say whether that time translated into something valuable. The metric gradually needs to go from hours to dollars to bottom-line impact.
This is exactly what leaders are doing. Research shows that enterprises reporting “productivity gains” as their primary AI ROI metric fell from 23% in 2025 to 18% in 2026 while “direct financial impact on revenue and profit” nearly doubled.
More than two-thirds are leaning on estimates like time saved or projected cost reductions instead of real, measured financial results.
Platforms purpose-built for AI cost complexity, like Tangoe’s, have an AI cost visibility dashboard plus reporting on AI-driven metrics like cost per inference, cost per business transaction influenced by AI, GPU utilization rates (for teams running their own infrastructure), and more.
Pro tip: Give each employee a monthly budget for AI usage (tokens/cost), establish a defined spending threshold, and use what transpires as a diagnostic tool. When someone hits their cap early, look at why. If they’re doing genuinely high-value work and need more budget, give it to them. If they’re burning through tokens inefficiently (bloated prompts, unnecessary back-and-forth), coach them on how to prompt better.
The hard truth is that most companies know their data isn’t ready to be used by AI, and 60% of AI projects with bad data end up being abandoned. That’s money paid with absolutely nothing to show.
Here are two scenarios where messy data tacks on tokens and leaves employees thinking, “I could’ve just done this myself faster.”
They ask AI to analyze sales data, but half the rows use “N/A,” some use blank cells, and others use “0” to mean missing data. The model can’t tell what’s real and what’s a placeholder, so it either produces a wrong analysis (requiring a correction round) or asks several clarifying questions back.
They ask AI to “pull up everything on this customer,” but the data lives across multiple systems that each have a different spelling of the customer’s name or a different account ID. Instead of one clean pull, they need to spoon-feed the model all the sources and tell it how to reconcile the mismatches.
Every clarifying question and reconciliation attempt is a full round of tokens – input and output – spent compensating for a data problem instead of focusing on the task at hand. Multiply that across dozens of interactions a day and bad data stops being an inconvenience and becomes a real line item on the IT bill.
In a way, it’s like strategy No. 1. To know where your money goes, you need clean AI tags. To know where your AI’s answer comes from, you need clean, accurate data. Being disorganized with either will cost you.
Focus on strategy over speed. Start with focused teams, proven use cases, and – critically – clean, well-scoped data. Measure the results, learn what works, and expand from there.
You can’t tell 10,000 employees to go use AI and hope it works out. Every employee could be spending a few dollars – or a few hundred dollars – every day. The bigger question is: what are you asking them to accomplish?
Every initiative should have a purpose. Are you trying to reduce customer response times? Speed up software development? Automate invoice processing? Whatever the goal is, define it and decide how you’ll measure success before tokens get consumed.
The same goes for AI pilots. Treat them like formal business projects: assign an owner, set a budget, establish success metrics, and review costs regularly. If the pilot is delivering measurable value, scale it. If
it isn’t, adjust your approach or move on.
Arguably the most important thing you can do is establish metrics. McKinsey recently surveyed companies about how they’re managing their AI rollouts and less than 20% could say they have actual, defined success metrics they’re tracking for their initiatives. That means 80% of companies using AI have no real way to answer the question “is this working?” because they never set up a way to measure it in the first place.
Only about one in three companies has governance mature enough to match the AI they’re running day-to-day.
The organizations seeing the greatest return from AI are continuously measuring what’s working, nixing what isn’t, and investing more in what delivers demonstratable value.
Whether your company buys AI tokens from a provider or runs models on its own infrastructure, the same problem shows up: token spend that’s hard to see and harder to attribute.
Tangoe’s platform brings visibility to both sides of that equation: tracking and attributing token spend from providers like OpenAI and Anthropic, and, for companies running their own AI infrastructure, extending that same visibility to the GPU and Kubernetes layer through our partnership with Kubex.
See how much you could be saving with a demo of our AI cost management platform.