11 Comments
User's avatar
Louis Hemming's avatar

I know the cheap model is 90% as good and I still reach for the flagship out of pure just-in-case, so multiply that by a whole org and the 45-day doubling stops being mysterious.

Ben Botes | GP & 4x Founder's avatar

<p>A year ago every CEO had twenty questions about AI. Now it's two. That's not because they figured it out — it's because they picked a vendor and stopped thinking about it. Different thing.</p>

RenderRepo's avatar

I wonder if we're optimizing the wrong metric. Token costs will come down over time, but the expensive part is still human verification. The companies that win may not be the ones with the cheapest AI, but the ones that can produce trustworthy outputs with the least amount of human review. That's why I'm trying to build a tool that builds trust among developer and product to shorten this loop before the AI workflow gets started.

David Alpert's avatar

Your two questions are one question, and it is the one you asked on X earlier this week without naming it.

Start inside the numbers. Token costs doubling every 45 days against 5-10% productivity gains is not a pricing failure, and cheaper models won't fix it: routing fifty-cent tokens toward unspecified work just buys the same drift at a discount. The problem is that most enterprise AI spend is demand without a valuer. Companies went all in because the demand was installed by board pressure, competitor FOMO, and the narrative, not because anyone wrote down what they were trying to achieve. When demand is manufactured rather than expressed, consumption compounds and output doesn't. That is what your CTO's numbers are measuring.

The meter makes it worse. A vendor paid per token gets richer as your usage grows, whether or not you do. Those are the economics of the engagement feed applied to enterprise software: the seller's revenue is your consumption, so the product is tuned for faster, longer, more. Tokenmaxxing isn't indiscipline. It's the meter working as designed. It is also why you and others spend so much time debating whether Anthropic wins or loses to the cheaper model. That is a debate about which seller captures your consumption, and the customer's objective appears nowhere in it. Their business models are priced on what you use, not on what you achieve.

So the question before "which model" is the one almost no company can answer in writing: what are we trying to achieve? Not a use-case list but a specification: purpose, objectives, constraints, and the measures of fulfillment, durable enough that providers can compete against it. The moment that document exists, everything you recommend becomes executable. Model choice gets obvious, because the spec defines what each job actually needs. The control plane finally has something to route toward. And payment can migrate from consumption to fulfillment, which aligns the vendor's income statement with your outcome for the first time.

That is also the answer to your second question. The $1.4T gets paid for by whoever captures fulfilled intent, and Baker and Masad are both describing the migration: value moving from general tokens to systems that complete defined jobs. The demand signal durable enough to fund the buildout isn't appetite for intelligence. It is specifications: the first true reading of real demand ever taken.

Which brings back your NPC post. Directionless people cycling through contradictory health protocols and directionless enterprises watching token bills compound are the same phenomenon at two scales: consumption standing in for purpose. You made the diagnosis twice in one week. The common cause is that expressed intent has never had infrastructure: no instrument by which a person or a company defines its objectives, states them durably, and makes the market compete to fulfill them. Manufactured demand filled the vacuum because installation was the only technology available.

The CEO test, one quarter out: are your people more capable, your intent more specific, your cost per outcome falling? If yes, you found a multiplier. If usage grew while capability didn't, you weren't being served. You were being harvested.

You told your readers not to take your answers or anyone else's, but to think for themselves. That is the whole game. Commerce should follow expressed intent. It never fully has. For the first time, it can.

Hunter Hastings's avatar

ROI is the wrong way to look at corporate investment in AI. They’re thinking cost control. They should be thinking about how much additional value they are creating for their customers. Value creation for customers - human flourishing - is the purpose of the firm, not efficiency. Management and administration (to achieve ROI) are restrictive growth killers. Imaginative value creation utilizing the new availability of intelligence is the growth driver.

dylan2045ad's avatar

Now there’s Qwen 3.8 too. It’s gonna be another magic Monday on the trading desk. Opus 5 has potential for a redemption arch on Weds or Thurs, if they master tokenomics and take one for Team America.. f yeah. But that loss from low prices would have to last months, but would lock 🔐 in Enterprise Suite flight to Deepseek v4 pro this week.

Jeff's avatar

Chamath correctly identifies that enterprises will optimize AI spending and increasingly mix frontier, open-weight, and distilled models. I think he’s much less convincing in implying this threatens AI infrastructure. Cheaper models typically expand the set of economically viable AI workloads. If inference demand continues to compound faster than cost per token falls—as it has so far—the total market for compute can continue growing even while individual models become commoditized. The key variable for AI infrastructure isn’t who owns the model; it’s how much inference the world consumes. Hat trick AI

Steven Shapiro's avatar

At the end of the day, the ROI from AI will come from two possible sources: 1) lower costs from higher employee productivity and/or fewer employees, and 2) incremental profitable products and services. All other metrics are intermediaries. So the corporate executives will need to track incremental AI expenditures against the marginal impact of those expenditures.

David Turner's avatar

The timing of ROI coming to roost isn’t about model progression as they already outpace 95% of “workers”. It’s about a judgement call as to when the worker learning curve reaches a point that both the worker and model should be able to produce quantifiable AND meaningful ROI (don’t forget meaningful). So baton officially passed to CFO and CEO to make that call. There is a cost to now, later and never as I recently wrote about. Make the call now, could miss the opportunity later. Make the call later and you waste money now. Make the call never and you’re really rolling the dice.

Shantanu Uniyal's avatar

Everyone’s obsessed with cutting costs, but AI’s true ROI is the compounding value of smarter, data-backed decisions, in the future - may be during the post AGI phase. How can we calculate that?

Dr. Mohammed Nadeem's avatar

Great question — and the tokenmaxxing framing is exactly right. I'd add one layer: token cost is really an answerability problem wearing a finance costume.

It compounds invisibly because no one owns it at the board level yet. Ask a CFO for cloud spend and you get a number to the dollar, because a decade of governance sits behind it: named owners, budgets, variance reviews. AI spend doesn't have that scaffolding yet, so it hides in OpEx until it clips an EPS number.

Routing to cheaper models is the right tactical fix. The durable one is governance: put a single owner on model spend, make "cost per decision" a reported metric, and treat the control plane as a board-visible artifact, not an eng detail.

The companies that miss the quarter won't be the ones paying $56 a million. They'll be the ones who couldn't say where the money went.