Journal / Startups

Your AI feature has a gross-margin problem.

The demo was magic and the launch went well and then the first full month's model bill arrived, and someone in finance did the division and went quiet. AI-first software is running gross margins that would have gotten a normal SaaS company laughed out of a board meeting. This is why, and what to do before it is your problem.

RL
RBB LAB
Studio
Published 17 Jul 2026 8 min read
$/tok RBB/LAB STARTUPS RBB LAB · JOURNAL 8 MIN READ

Traditional software has a lovely economic shape. You pay to build it once, and every copy after that costs you almost nothing to serve. That single fact, near-zero marginal cost, is why software gross margins sit at seventy to ninety percent, why the whole venture model works, and why "just add another customer" is a sentence a SaaS founder can say without flinching. It is the best business anyone ever invented, and a generation of company building was quietly organised around it being true.

AI features do not have that shape. Every time a user does the magic thing, the thing you put in the demo, you make a call to a model and you are charged for it, by the token, in proportion to how much they used. The marginal cost of serving one more request stopped being approximately zero. It became a real, recurring number that grows exactly as fast as your product succeeds. AI-first software is now running gross margins in the twenty to sixty percent range, and a lot of founders have not yet looked hard enough to know which end of that range they are on.

Software's superpower was near-zero marginal cost. AI quietly gave it back a cost of goods sold, priced per request, that scales with usage instead of with your bank balance. That is not a rounding error. It is a different kind of company.

Why the margin is structurally worse

The reassuring thing people say is that inference is getting cheaper, and it is, dramatically so. Per-token costs have fallen something like ten to a hundred times since 2023, and a large chunk of that just in the last year. So the problem solves itself, the story goes. It does not, for two reasons that pull in the opposite direction from the price cuts.

First, consumption grows faster than unit price falls. As the model gets cheaper you do not use the saving to bank margin; you use it to do more per request, longer context, multiple calls, an agent that loops. The industry has spent every efficiency gain on ambition, not on profit, and that is a rational thing to do when your competitor is doing it too. Second, you do not control the input price. Your single largest variable cost is a line item on someone else's price list. Frontier models still run somewhere around two to fifteen dollars per million input tokens and ten to seventy-five per million output, and a pricing change you did not vote on lands straight on your margin.

The result is a cost of goods sold that behaves nothing like traditional software hosting. It is not a small fixed platform cost you amortise across everyone. It is a per-use meter running underneath your most-loved feature, and your heaviest, happiest users, the ones a normal SaaS would treasure, are the ones costing you the most. This is the same tension we flagged from the inside in how an AI-native team actually operates: you can buy more generation cheaply, but the economics only work if someone is watching what it costs.

In classic SaaS your best customer is nearly free to serve. In an AI product your best customer is your biggest cost line. Nobody's spreadsheet was built for that inversion.

The per-seat trap

Here is where good companies walk into the wall. They price the AI product the way they have always priced software: a flat fee per user, per month. It is familiar, it is easy to sell, and it is quietly lethal, because it severs the one link that has to stay connected, the link between what a customer pays you and what that customer costs you.

Consider two customers on the same plan. One runs a few hundred model calls a month. The other runs millions. Under per-seat pricing they pay you the identical amount, and you are comfortably profitable on the first and losing money on every login from the second. Worse, the heavy user is the one who loves the product most, tells their friends, and expands. You have built a machine that grows fastest exactly where it loses the most money, and the better your product is, the faster the trap closes.

20–60%
Gross margin range for AI-first software
70–90%
Gross margin for traditional SaaS, for comparison
Of top AI companies now run two or three pricing models at once

That last number is the tell. Nearly half the leading AI companies have abandoned a single, clean price and now run two or three models side by side, a subscription for predictability, usage-based for the API, a free tier to acquire, precisely because no single lever both sells cleanly and protects the margin. Pure-play, one-line pricing is dying in AI for a structural reason, not a fashionable one.

The payback math quietly broke

The damage does not stop at gross margin; it flows straight into the growth engine. Every venture-scale software company runs on the same underlying bet: spend to acquire a customer now, earn it back over the following months, then profit. That bet is only as good as your margin, because you pay back acquisition cost out of gross profit, not revenue.

Cut the margin from eighty percent to forty and you have not trimmed the model, you have roughly doubled how long it takes to recover the cost of every customer you buy. A payback period that was a comfortable twelve months slides toward two years, and a growth plan financed on the old assumption is now underwater without a single thing changing in sales or marketing. This is exactly the class of error we catalogued in the three decisions that kill technically sound startups: the product works, the team is good, and the company still fails because a number nobody was watching moved underneath everything else.

You repay customer acquisition out of gross margin. Halve the margin and you have doubled your payback period, for free, without noticing, until the quarter the growth model stops adding up.

What actually works

None of this means AI features are a bad business. It means they are a different business, and the companies that treat them that way are fine. A few things separate them.

Price along the axis that drives your cost. If cost scales with usage, some part of the price has to scale with usage too. The winning shape in 2026 is rarely pure usage billing, which terrifies buyers who cannot predict their invoice, and rarely a flat seat, which bankrupts you on the heavy users. It is a hybrid: a base subscription for the predictability customers want, plus metered usage or generous-but-real included limits above which the meter starts. The subscription sells; the meter survives contact with a power user.

Know your unit cost before you set the price, not after. The single most common failure we see is a team that cannot tell you the model cost of one typical customer action. You cannot price what you have not measured. Instrument the cost per request, per feature, per customer, from the first week, and treat that number as a first-class product metric, not a finance afterthought.

Engineer the margin back, deliberately. Most of the cost is recoverable by people who bother: route the easy eighty percent of requests to a smaller, cheaper model and reserve the frontier model for the fifth that needs it; cache aggressively; cap runaway agent loops; trim context that is padding rather than signal. This is unglamorous margin engineering, and it is the closest thing to free money in an AI product, the same way we treat a deployment pipeline as infrastructure you build on day one rather than repair under duress.

What we tell founders

When a founder shows us an AI product, the gross-margin question comes before the feature demo, because it is the one that decides whether there is a company here or just an expensive hobby that charges money. We ask three things, in order.

What does one unit of the magic cost you today, in cents, at the model. How does that number move as a customer goes from casual to power user, is it flat, linear, or does it explode. And does your pricing touch that curve anywhere, or have you priced on a flat seat and simply hoped the average holds. If the answer to the last one is a seat and a hope, we stop and fix the pricing before we build another feature, because more features on a broken unit economic just means losing money more efficiently. It is the pricing equivalent of the compounding cost of technical debt: cheap to fix at the start, ruinous to fix once the whole business is built on top of it.

The bottom line

The uncomfortable truth of AI software is that the technology gave you a feature customers love and, in the same motion, handed your product a cost of goods sold that classic software had spent forty years engineering away. That is not a reason to avoid building; it is a reason to build with your eyes open. The winners of this cycle will not only be the teams with the best model output. They will be the teams who knew their unit cost cold, priced against the curve that actually drives it, and engineered their margin back on purpose while their competitors were still admiring the demo. Everyone can ship the magic now. Far fewer can afford to serve it, and that, not the model, is where this round gets decided.

RL
RBB LAB
Studio · San Marino
A small team of senior engineers building production software for businesses and founders. We ship, hand off, and disappear cleanly.
Stay in the loop

One email when we publish. Nothing else.

About once a month. Sometimes less. No funnels, no drip campaigns.

Or grab the RSS · Follow on LinkedIn / X