← Insights

Enjoy the Subscription While It Lasts

Clessidra con lettere che scorrono verso l'AI

$725 billion of capex, a marginal cost that finally became visible, and why the flat AI fee is living on borrowed time — including yours.

If you work in tech — or you just love following it — there's one number you're certainly aware of: $725 billion.

That's the combined 2026 capital expenditure of the four hyperscalers — Amazon, Google, Meta, Microsoft — on AI datacenters, silicon, shells and power. It's up roughly 77% on 2025. Amazon alone is around $200 billion.

To put that in perspective: it equals roughly 28% of Italy's entire 2025 GDP (~$2.55 trillion — a G7 economy). And in absolute terms, $725 billion is larger than the entire annual GDP of most countries on earth — bigger than the full-year output of all but about twenty of them. More than a quarter of what a G7 nation produces in a year, or the whole economy of a mid-sized country — spent on datacenters, in twelve months, by four companies, on one bet.

For a long time this was the question nobody asked out loud — impolite, at a party where everyone was drunk on the future. That's over. Slowly, then all at once, everyone is starting to ask it:

Who pays that back?

Not "Who benefits?" Not "Will it change the world?" Those are the fun questions. The boring one — the one that decides who's still standing in 2030 — is the repayment question. $725 billion is not a grant and it's not philanthropy. It sits on balance sheets as an asset that depreciates on a brutal schedule, because GPUs are not cathedrals; they're closer to milk than to marble. Somebody, somewhere, through some pricing mechanism, has to turn that capex into cash faster than it rots.

For two years, the answer was supposed to be you.

The subscription has never been a product — it's a customer-acquisition subsidy

Here is the playbook, and it's an old one. You get people addicted to a flat monthly fee. Twenty dollars, two hundred dollars, whatever — the point is that it's predictable and it feels generous. "All you can eat." You stop counting. You build workflows, habits, entire teams around a model. Meanwhile, switching costs quietly accumulate. And then, once you're inside, the pricing quietly migrates from flat to metered.

This isn't a theory, it's already happening in the open. In 2026 Anthropic moved its enterprise billing toward per-token pricing, unbundling model usage from seat fees. GitHub Copilot shifted its plans to usage-based AI credits. OpenAI has been dancing around the same fault line with credit packs bolted onto "included" tiers. The flat-rate, all-you-can-eat era is being dismantled in real time — and the reason is arithmetic. The subsidy was always going to end, because the subsidy was funding a rounding error against $725 billion.

Except something happened on the way to the metered future that the spreadsheet didn't model.

The moment cost became visible, customers did the math — and ran

Here's the mechanism that the strategy decks missed. A flat fee doesn't just simplify billing. It hides the marginal cost of serving you. When it's twenty dollars a month no matter what, you never ask what a single query actually costs. You can't optimize what you can't see.

The instant you switch to a meter, that changes. Now every request has a price tag. Now finance is in the room. Now someone builds a dashboard, sees the number ballooning past the payroll line, and asks the only question that matters: do we actually need the expensive model for this?

And the answer, embarrassingly often, is no.

The Financial Times laid it out plainly: companies from Silicon Valley to Europe are rerouting workloads onto Chinese open-weight models to cut costs and reduce dependence on US frontier labs. This is not a fringe of cost-desperate startups. DoorDash delegates lower-level work to Moonshot's Kimi and reserves the US frontier model only for the hardest tasks — a combination its co-founder says outperforms the all-frontier setup at lower cost. Siemens runs a broad mix including DeepSeek and Z ai alongside US and French models, explicitly for "flexibility." Airbnb uses China-origin models routed through approved US providers. Lindy, the AI agent startup, moved off Anthropic entirely to DeepSeek's V4 and called the result "transformative" — millions saved, performance up on many core use cases.

The numbers underneath are the story. The best open-weight models are, per the people selling access to them, 10 to 60 times cheaper than their proprietary equivalents. And the capability gap that once justified the premium is closing fast: when Z ai shipped GLM-5.2 in June, even Marc Andreessen — not a man who undersells American technology — noted that insiders were calling it the first Chinese model to match and often beat the big US labs "with no compromises."

Ten to sixty times cheaper. Read that again with a capex balance sheet in your other hand.

The scissor no one is drawing

Put the two halves together, because separately they're just news and together they're a thesis.

On one side: fixed costs of historic size. $725 billion this year, and the depreciation clock started the day the GPUs were racked. This is committed capital. You can't un-spend it, and you can't wait — the assets decay whether or not the revenue shows up.

On the other side: pricing power on inference that is evaporating. Every quarter, an open-weight model gets good enough to absorb another slice of the workload that used to require the premium tier. Inference is commoditizing in front of us, and commoditization doesn't destroy margin so much as relocate it — away from whoever trained the model, toward whoever is smart enough to orchestrate a fleet of cheap ones.

That's the scissor. Enormous, sunk, depreciating fixed costs meeting a variable-revenue line that the market is actively trying to drive to zero. Every business-school reflex says: when that happens, you protect the fixed-cost recovery by locking in pricing. Meter everything. Make it sticky. But the very act of metering is what made the cost visible, which is what triggered the exodus. The move that's supposed to save the model is the move that detonates it.

Furthermore there's a geopolitical accelerant on top, and it's worth naming because it makes the trend one-directional. When the US briefly imposed export controls on Anthropic's frontier models, European enterprises got a cold look at a risk they'd been ignoring: single-vendor, single-jurisdiction dependency for a system now wired into their operations. The ban was reversed. It didn't matter. As one European AI officer put it, you can put the model back on the market, but you can't put the genie back in the bottle. A venture investor quoted in the same coverage said the quiet part with startling clarity: two years ago the fear in Europe was China; now the bigger fear is the US. That is a staggering reversal, and it means the flight to open, self-hostable models isn't only a cost trade. It's a sovereignty trade. Cost and control now point in the same direction, and that's the most dangerous kind of trend for an incumbent — the kind where your customer saves money and sleeps better by leaving.

Why the flat fee is structurally doomed — a first-principles detour

Strip away the news and go to the physics of the business model, because that's where the real answer lives.

A flat subscription is an honest, durable product under one specific condition: the marginal cost of serving one more unit of you is either near zero or stable and predictable. That's why it works for Netflix and Spotify — the marginal cost of streaming one more film is a few cents of bandwidth and a fixed royalty, and it doesn't move whether you watch one film or a hundred. The provider can pool risk across millions of users and price the average with confidence. The heavy user is subsidized by the light user, everyone gets predictability, and the math holds.

AI inference violates that condition at the root. The marginal cost of one more query is real — measurable compute and measurable energy, dollars not cents — and it's volatile, swinging by orders of magnitude between a one-line answer and an agent that thinks for twenty minutes and calls forty tools. When marginal cost is real and volatile, flat pricing stops being a pricing strategy and becomes a bet — the provider is gambling that the heavy users don't cluster and the average holds. But agents break the average. Autonomous, always-on workloads are precisely the use case everyone is racing toward, and they are the use case that turns a flat-fee cohort into a portfolio of uncapped liabilities.

A business whose marginal cost is real and now transparent is not a subscription business. It's a utility. And utilities meter. The flat AI fee isn't a product decision that can be defended with better retention tactics. It's a category error that the market is in the process of correcting.

So on the enterprise side, the direction is set. The meter is coming, the workloads are already leaking to models 10-60x cheaper, and the smartest buyers are building routing layers that treat the frontier model as an expensive specialist, not a default.

The interesting question — the one I actually want to leave you with — is what happens to you. The consumer. The person paying twenty dollars a month and not counting.

The consumer subscription doesn't have one future. It has five.

Here's where I part company with the confident takes. Everyone wants a single prediction. I don't think there is one. I think the flat consumer AI subscription fractures into several distinct futures, and which one you end up living in depends on who you are, what you use, and who owns the rail underneath you. Let me lay out the branches, roughly in order from most brutal to most disguised.

Path 1 — The pure meter (the utility endgame). The honest one. You pay for what you use, per token, per task, like electricity and water. Transparent, fair, and psychologically miserable — because metering breeds anxiety, and anxiety kills usage. This is the destination the economics want, and precisely for that reason it's the one providers will resist offering nakedly, because a visible meter makes people use less, and they need you to use more. Expect this to arrive first for power users and developers, dressed up as "credits."

Path 2 — The hybrid: flat base plus overage (the most likely). A modest flat floor for predictability, then the meter kicks in above a threshold. This is already the shape of OpenAI's credit packs and Copilot's AI credits, just not yet marketed to consumers in those words. It preserves the comforting monthly number while quietly reintroducing marginal pricing at the top of the distribution — which is exactly where the money bleeds. My bet is that this is where the middle of the market lands. "The flat fee disappears entirely" is the provocative version; "the flat fee survives as a shrinking floor under a meter" is the version I'd actually defend.

Path 3 — The bundle absorption (the meter you never see). The AI subscription stops being a line item and disappears into something bigger — Prime, Apple One, your telco plan, your bank, your operating system. It's "free," which means it's cross-subsidized, which means you are the instrument of subsidy: your data, your attention, your lock-in to the parent ecosystem. The meter still runs. You just never see it, because someone with a strategic reason to keep you inside is paying it on your behalf and booking the cost against a different P&L. This is the Apple/Amazon/Google-shaped future, and it's the one I'd watch most closely, because it's how a flat fee survives a metered reality.

Path 4 — The capability ladder (flat, but capped by quality). Flat pricing survives — for the cheap models. You get unlimited access to a good-enough commodity tier, and the frontier capability sits behind a meter or a much higher tier. This is the airline model applied to intelligence: economy is all-you-can-eat and forgettable, and the good seats are priced by the trip. It works precisely because open-weight models made "good enough" nearly free to serve, so the provider can afford to give it away flat and monetize the delta.

Path 5 — The sovereign split (bring your own model). The most radical, and initially the most niche. As open-weight models get good enough and local hardware gets cheap enough, a slice of consumers — the technical, the privacy-obsessed, the geopolitically nervous — stop renting intelligence altogether and run it themselves. No subscription, no meter, no vendor. Today this is a rounding error. But every enterprise trend in this essay started as a rounding error eighteen months ago, and the same forces — cost, control, sovereignty — point the same way for the individual. Don't bet the house on it. Don't ignore it either.

Notice what these five paths share. In every single one, the flat fee as you know it today — pay once, forget the cost, use without limit — is gone. It either becomes a meter, hides a meter, floors a meter, tiers around a meter, or gets replaced by ownership. The unlimited buffet doesn't survive contact with a marginal cost this real and this visible. It only survives as an illusion someone else is funding for a reason.

So: enjoy it while it lasts!

The flat AI subscription is a subsidy paid for out of $725 billion of capital that has to be repaid on a depreciation schedule that doesn't care about your feelings. Subsidies end. The all-you-can-eat pricing you're enjoying right now is not a business model reaching maturity — it's a customer-acquisition promotion in its final, generous act, right before the meter gets switched on. Enjoy it. Genuinely. Just don't build your life or your company on the assumption that the number on the invoice stays flat and stays small.

And here's the question I'd actually sit with, because it's the one that separates the people who'll be fine from the people who'll get a very unpleasant invoice one quarter:

Look at your stack — personal or corporate. If tomorrow every single request were billed at full price, meter running, no flat fee to hide behind — how much of what you do today would you still do? And how much of it were you only doing because, for a little while, it felt free?

The uncomfortable follow-up — which of the five paths wins for the mass consumer, and what would have to be true for the flat fee to actually survive — is the one worth disagreeing with me about. If you have an answer, I'd like to hear it: db@ogeno.it.

  • Financial Times, Companies turn to Chinese AI models to cut costs and reduce US dependenceft.com
  • Hyperscaler AI capex 2026 (~$725B, +77% YoY) — Value Add VC tracker · Tom's Hardware
  • Anthropic's shift to usage-based enterprise billing — implicator.ai
  • GitHub Copilot moves to usage-based AI credits (June 2026) — reported alongside the broader metering shift
  • Lindy's migration from Claude to DeepSeek — The Decoder

The five-path framework and the capital thesis are my own analysis. Facts are sourced above; the forecast is opinion — argue with it.