AI inference costs are eroding software's 75-85% margins

An essay published on nicolo.xyz, whose author the text does not name, argues that AI is dismantling one of software's oldest advantages. For about fifteen years, software ran on a genuine anomaly: build a product once and distribute it to a million users for roughly the cost of distributing it to one. Every additional customer flowed largely to the bottom line, producing gross margins of 75-85%, a level the essay says made software unlike any other industry in history. That underwrote the standard SaaS playbook: land customers aggressively now and worry about profitability later, because unit economics only improved with scale, not the reverse.
Once a product answers a user by calling a large language model, the essay argues, every interaction carries a real, scaling compute cost, an inference call, as distinct from the one-off cost of training a model. That creates a tradeoff SaaS never had: a cheaper model protects margin but risks losing users to a better-performing competitor, while a frontier model wins on quality but bleeds on unit economics. The essay likens it to the component-sourcing tradeoffs manufacturers have always faced, arriving on a software founder's desk for the first time. Most AI founders already manage it in practice, running several models at once and routing routine requests to smaller or fine-tuned models while reserving frontier models for hard cases, but the essay argues that as products scale, inference tends to become the dominant cost line, displacing even talent.
The essay's central evidence is a survey from ICONIQ, 'State of AI: Bi-Annual Snapshot,' published in January 2026 and covering roughly 300 executives at software companies building AI products. It puts the average gross margin on AI products at about 52% in 2026, up from 41% in 2024, but still well below the 75-85% range that defined the SaaS era. The essay cites Legora, described as built on foundation models from Anthropic or OpenAI, only as an example of what it calls the 'application layer', not as a specific case of margin compression.
A second dimension of the squeeze, per the essay, is variability in cost to serve: a power user making a hundred inference calls costs a company dramatically more than a casual user making five, where in classic SaaS both cost almost nothing. Flat-rate subscriptions spread that gap across the whole user base, which works only until the heaviest users become expensive enough to blow up a company's margins, and the essay argues that is why AI-native companies are steadily moving toward usage-based pricing that ties what customers pay to what they actually cost to serve.
On the hope that falling per-call inference prices will solve the problem by themselves, the essay invokes Jevons paradox, named for William Stanley Jevons, who observed in 1865 that more efficient coal engines led to more coal being consumed, not less. Its argument: cheaper inference will not be banked as margin but spent on deeper integration, more calls per interaction, and more autonomous agents running in the background, so the cost per call falls while the calls per user multiply. The essay concludes that the resulting playbook looks less like the growth-at-any-cost SaaS era and more like traditional capital-efficient business building: unit economics considered from day one, pricing that reflects real cost structure, and growth funded by the business itself rather than by investors betting on margin expansion that may never arrive. It frames the shift not as a eulogy for software but as one that could produce more resilient, more honest companies than the SaaS generation built.
Key facts
- Traditional SaaS ran on 75-85% gross margins because an extra user cost almost nothing to serve; the essay argues AI ties compute cost directly to every interaction, breaking that pattern for the first time.
- ICONIQ's January 2026 'State of AI' survey of roughly 300 software executives puts average AI product gross margins at 52% in 2026, up from 41% in 2024, still well below the old 75-85% SaaS baseline.
- A power user making a hundred inference calls costs a company dramatically more than a casual user making five, unlike classic SaaS where both cost almost nothing, a gap the essay says is pushing AI-native companies toward usage-based pricing.
- The essay argues falling per-call inference costs will not be banked as margin: invoking Jevons paradox (from William Stanley Jevons' 1865 observation on coal use), it says cheaper inference gets consumed by more calls, deeper integration, and more autonomous background agents.
- As products scale, the essay argues, inference tends to overtake even talent as the dominant cost, pushing founders toward unit-economics-first, capital-efficient business building instead of the old growth-at-any-cost SaaS playbook.
Why it matters
For about fifteen years, software's near-zero cost of serving an additional user underwrote the entire SaaS growth playbook: acquire customers aggressively now, because unit economics only improve with scale. The essay argues AI breaks that assumption for the first time in the industry's history, since every LLM call carries a real, per-user cost that scales with usage, putting growth and product quality into direct competition with margin. It frames this as a hardware-style tradeoff, the kind of component and input-cost decision manufacturers have always faced, landing on a software founder's desk for the first time.
Who it affects
Founders and investors in AI-native software, especially companies in what the essay calls the 'application layer', those building on top of foundation models such as Anthropic's or OpenAI's rather than owning the underlying infrastructure (the essay's example of this category is Legora). It also concerns anyone who priced a software business on the old 75-85% gross margin assumption, since the ICONIQ data the essay cites shows AI products currently averaging well below that.
How to use it
The essay's practical suggestions: run multiple models rather than one, routing everyday requests to smaller or fine-tuned models and reserving frontier models for the hardest cases; build unit economics into a product from day one instead of assuming margin will improve automatically with scale; and shift pricing away from flat-rate subscriptions toward usage-based pricing that tracks what each customer actually costs to serve, since a small share of heavy users can otherwise erase the margin flat pricing depends on.
How solid is it
This is a single essay published on nicolo.xyz and discussed on Hacker News, where it drew 50 points and 36 comments; it is argument and analysis, not a study the author ran, and the text does not name its author. Its central quantitative claim rests on one third-party source, ICONIQ's 'State of AI: Bi-Annual Snapshot' (January 2026, roughly 300 executives surveyed), which the essay cites but did not produce. The rest of the piece, including the hardware analogy and the Jevons paradox argument, is the author's own interpretive framework rather than separately sourced data.
Risks and caveats
The essay names no company as a concrete case of margin compression from inference costs and gives no model names or per-call dollar figures, so its mechanism stays qualitative even where the ICONIQ percentages are precise. Those percentages are self-reported survey averages across roughly 300 executives, not audited financials, and will mask wide variation between individual companies. The essay does not say when the market 'equilibrium' it predicts will actually arrive. And its answer to falling inference prices, that Jevons paradox will absorb any savings through higher usage, is an analogy drawn from nineteenth-century coal economics rather than a measurement of how software costs are actually behaving.
“The rules of software are being rewritten. The best founders are already playing by the new ones.”
— the essay's conclusion