Agentic AI’s hangover
“How's your token spend these days?”
This past Sunday I ran into a friend I hadn't seen in months, fresh back from Iceland. Before I could ask about his trip, he shook my hand and said, “How’s your token spend these days?”
Not the glaciers. Not the northern lights. Tokens. When a guy skips his own vacation stories to ask about your Anthropic invoice, there has been a seismic shift.
And here’s the shift: agents don’t need bigger engines; they need the right engine for the job, and most companies are burning real money finding that out the hard way. I’ve felt it for weeks. We rode off the top of the AI euphoria mountain into the trough of agentic disillusionment, and the fall came faster than anyone anticipated. DHH said it well on a podcast recently — he called the dopamine loop of shipping with agents “intoxicating,” and he’s right. A lot of companies have been partying like drunken sailors, and now the hangover has arrived.
Not the glaciers. Not the northern lights. Tokens. When a guy skips his own vacation stories to ask about your Anthropic invoice, there has been a seismic shift.
And here’s the shift: agents don’t need bigger engines; they need the right engine for the job, and most companies are burning real money finding that out the hard way. I’ve felt it for weeks. We rode off the top of the AI euphoria mountain into the trough of agentic disillusionment, and the fall came faster than anyone anticipated. DHH said it well on a podcast recently — he called the dopamine loop of shipping with agents “intoxicating,” and he’s right. A lot of companies have been partying like drunken sailors, and now the hangover has arrived.
The token problem
Somebody has to own the token risk
Over the past six weeks, I’ve sat across the table from CEOs, CIOs, CTOs, and directors at some of the best-run companies in the post-acute healthcare industry. They all keep telling me a version of the same thing: consumption-based AI pricing does not work for them.
One CIO put it bluntly: “I want you guys to take the risk of token cost,” he told me. “Give me a fixed monthly rate for a guaranteed amount of agent work completed.” I pushed back and said I wasn’t sure we could, that doing it wrong meant either losing our shirts or throttling the product until it was useless. He didn’t blink. “Look, I view you guys as our outsourced AI systems engineers,” he said. “If the people building the agents and workflows aren’t motivated toward token efficiency, who is?”
One CIO put it bluntly: “I want you guys to take the risk of token cost,” he told me. “Give me a fixed monthly rate for a guaranteed amount of agent work completed.” I pushed back and said I wasn’t sure we could, that doing it wrong meant either losing our shirts or throttling the product until it was useless. He didn’t blink. “Look, I view you guys as our outsourced AI systems engineers,” he said. “If the people building the agents and workflows aren’t motivated toward token efficiency, who is?”
“I want you guys to take the risk of token cost.” – Healthcare CIO
I hated it, mostly because he’s right. Nobody outside the team building these systems can actually solve tokenomics. The customer and end user can’t see where the tokens go. The finance team can’t either, and the model vendors (OpenAI, Anthropic, Google) are all delighted to sell you more. The only people who can make an agent both cheap and accurate are the people who build the agent systems, and that responsibility isn’t moving anywhere else.
At Carebility, we run everything against a single rule we call Voss’s Law: AI accuracy has to climb while AI token cost falls, or there’s no value. By accuracy, I mean output that consistently meets or beats the eval tests for highly specific clinical admin workflows, which is a far higher bar than it sounds. Miss on either side and you don’t have a product. You have a science experiment with an invoice.
At Carebility, we run everything against a single rule we call Voss’s Law: AI accuracy has to climb while AI token cost falls, or there’s no value. By accuracy, I mean output that consistently meets or beats the eval tests for highly specific clinical admin workflows, which is a far higher bar than it sounds. Miss on either side and you don’t have a product. You have a science experiment with an invoice.
Right-sized AI
Match the engine to the job
So the answer for us is not necessarily faster AI models. It’s better AI systems. And I explain that with engines.
The latest and greatest frontier LLMs from OpenAI, Anthropic, and Google are like the Ferrari V12s of AI. Insanely powerful, beautifully engineered, expensive, and capable of things that feel impossible. But a lot of AI work doesn’t need a Ferrari engine. Sometimes you need a Toyota RAV4, reliable and cheap and safe, the thing that gets the job done every day. Sometimes you need a Ford F-150, durable and practical, built to haul tools and do work.
And sometimes you just need plain old-fashioned tools in that F-150. Tools that are simple but get the job done precisely. In our line of work, that translates into developing classic, deterministic logic tools tailored for smaller-parameter model agents. Tools that successfully complete the job at a fraction of the expense of high-priced frontier models.
That’s the lesson for agentic AI. Stop putting a Ferrari engine in every workflow. Better agents aren’t always built by throwing bigger models at the problem; they’re built by matching the engine to the job, with LLM routing, smaller open models, rules, tools, memory, retrieval, and human review, and by calling the frontier model only when the road actually demands it.
Use the Ferrari when you need genius. Use the truck when you need work done.
The latest and greatest frontier LLMs from OpenAI, Anthropic, and Google are like the Ferrari V12s of AI. Insanely powerful, beautifully engineered, expensive, and capable of things that feel impossible. But a lot of AI work doesn’t need a Ferrari engine. Sometimes you need a Toyota RAV4, reliable and cheap and safe, the thing that gets the job done every day. Sometimes you need a Ford F-150, durable and practical, built to haul tools and do work.
And sometimes you just need plain old-fashioned tools in that F-150. Tools that are simple but get the job done precisely. In our line of work, that translates into developing classic, deterministic logic tools tailored for smaller-parameter model agents. Tools that successfully complete the job at a fraction of the expense of high-priced frontier models.
That’s the lesson for agentic AI. Stop putting a Ferrari engine in every workflow. Better agents aren’t always built by throwing bigger models at the problem; they’re built by matching the engine to the job, with LLM routing, smaller open models, rules, tools, memory, retrieval, and human review, and by calling the frontier model only when the road actually demands it.
Use the Ferrari when you need genius. Use the truck when you need work done.
Post-acute care reality
Providers who need AI most can’t afford waste
In my world, it isn’t even optional. We serve post-acute healthcare — skilled nursing, home health, and hospice — where the economics are brutal: capitated rates, Medicare Advantage and insurance plans squeezing reimbursement year after year, regulation stacked on regulation, and a staffing shortage that isn’t a blip but a structural, societal crisis. Operators like these can’t and won’t pay Ferrari prices to run software.
But here’s the twist: post-acute care providers need agents more than anyone, because the very economics that make them refuse to overpay are the same economics that make agents transformational in the first place. For the provider drowning in critical, reimbursement-impacting paperwork with nobody qualified to do it, an agent that actually works changes the whole equation. The financial discipline isn’t a constraint on the opportunity. It is the opportunity.
But here’s the twist: post-acute care providers need agents more than anyone, because the very economics that make them refuse to overpay are the same economics that make agents transformational in the first place. For the provider drowning in critical, reimbursement-impacting paperwork with nobody qualified to do it, an agent that actually works changes the whole equation. The financial discipline isn’t a constraint on the opportunity. It is the opportunity.
“The financial discipline isn’t a constraint on the opportunity. It IS the opportunity.”
The cost reckoning
The fountain of unlimited token spending is drying up
None of this is theoretical, and the signals are everywhere. The first wave of enterprise AI got sold like software: pay per seat, roll it out, let people use it. But agents don’t behave like traditional software. They read, reason, retry, summarize, call tools, inspect files, write code, and check their own work, and every one of those steps burns tokens that aren’t free. So the budget crisis this time isn’t cloud sprawl or SaaS bloat. It’s token burn.
Let’s start with GitHub, which moved Copilot to usage-based billing on June 1. Every plan now meters input, output, and cached tokens at API rates, and GitHub’s own stated reason was that the flat fee had become “no longer sustainable.” Remember that GitHub is Microsoft, and Microsoft is pushing AI into daily work harder than almost anyone, and Microsoft is one of the richest companies on the planet. So when they blink on pricing, it’s worth paying attention.
Let’s start with GitHub, which moved Copilot to usage-based billing on June 1. Every plan now meters input, output, and cached tokens at API rates, and GitHub’s own stated reason was that the flat fee had become “no longer sustainable.” Remember that GitHub is Microsoft, and Microsoft is pushing AI into daily work harder than almost anyone, and Microsoft is one of the richest companies on the planet. So when they blink on pricing, it’s worth paying attention.
A cautionary tale
Uber learned the cost of ungoverned AI
Uber is the louder warning, and it’s worth reading carefully, because the lesson isn’t “AI doesn’t work” — it’s “nobody was governing the spend.” Forbes reported the company burned through its entire 2026 AI budget in four months, with five thousand engineers on Claude Code and Cursor and some of them spending two grand a month each.
Uber had even been running internal leaderboards that ranked teams by AI usage, gamifying the burn. When its COO was finally asked whether all that spend was producing better products for riders, he admitted the link “is not there yet.” That’s not a verdict on agents.
That’s a verdict on running agents with no engine-matching discipline at all — exactly the gap this piece is about.
Uber had even been running internal leaderboards that ranked teams by AI usage, gamifying the burn. When its COO was finally asked whether all that spend was producing better products for riders, he admitted the link “is not there yet.” That’s not a verdict on agents.
That’s a verdict on running agents with no engine-matching discipline at all — exactly the gap this piece is about.
The market response
FinOps is coming for AI agents
The same scramble is showing up across the market, and the Wall Street Journal has traced it through Priceline, Qualcomm, and Bristol Myers Squibb, all of them reaching for dashboards, spend caps, showback, and chargeback. You could call it FinOps for tokens. The financial discipline we spent a decade building for cloud is being reinvented overnight for agents.
The research backs it, too. A Stanford Digital Economy Lab study found that agentic coding tasks can consume a thousand times more tokens than a normal chat, with most of that usage on the input side as the agent re-reads its own context on every step, and that running the same task twice can swing the cost thirtyfold. But the most critical part that every CFO should tattoo somewhere: spending more tokens did not buy more accuracy.
It peaked in the middle and then flatlined. In other words, you can pay double and get worse. This is why we created Voss’s Law in the first place.
The research backs it, too. A Stanford Digital Economy Lab study found that agentic coding tasks can consume a thousand times more tokens than a normal chat, with most of that usage on the input side as the agent re-reads its own context on every step, and that running the same task twice can swing the cost thirtyfold. But the most critical part that every CFO should tattoo somewhere: spending more tokens did not buy more accuracy.
It peaked in the middle and then flatlined. In other words, you can pay double and get worse. This is why we created Voss’s Law in the first place.
The new discipline
Govern AI as a cost center before expecting profit
That’s the whole game now: AI has to be governed like a cost center before it can ever become a profit center.
Which means that for every task, you need to know what it’s doing, what it’s worth to the business, what model it actually requires, how many tokens it’s allowed to spend, how it is secured, when it should stop and hand off to a human, and where the audit trail lives.
Which means that for every task, you need to know what it’s doing, what it’s worth to the business, what model it actually requires, how many tokens it’s allowed to spend, how it is secured, when it should stop and hand off to a human, and where the audit trail lives.
“AI has to be governed like a cost center before it can ever become a profit center.”
Nvidia, Microsoft, Meta, Google, Amazon, and the frontier labs like OpenAI and Anthropic can absorb a monstrous compute bill because they’re building the infrastructure and fighting for a once-in-a-generation platform shift. But a nursing home operator can’t, and neither can a regional bank or a manufacturer. They don’t get rewarded for burning tokens; they get rewarded for bringing value, reliability, compliance, and outcomes you can actually measure.
The euphoria is over and the hangover is real, and the builders who sober up first, the ones who make agents cheap and accurate and governed, are the ones who’ll own the next decade.
The euphoria is over and the hangover is real, and the builders who sober up first, the ones who make agents cheap and accurate and governed, are the ones who’ll own the next decade.
Practical AI for post-acute care
See what right-sized AI looks like in practice.
Carebility helps post-acute organizations deploy AI that is accurate, governed, and built for real workflow economics — not maximum model spend.
See Carebility in action