SaaSletter - Brute-Force AI + Gross Margins
Why exponential inference breaks software gross margins - plus token pricing data + takes from Carmen Li, CEO of Silicon Data.
Inference Goes Exponential → Brute Force Implications
Our recent editions have covered the exponential curve for inference needs (h/t Siddharth Ramakrishnan of Scale Venture Partners) that sets up…
Dave Kellogg’s post on the "Brute-Force Era of AI". He rightly points out that while applying massive compute to architectures works (Sutton's Bitter Lesson), it doesn't inherently solve the efficiency problem.
We are officially entering Phase 2:
Phase 1 was about scaling to discover what works; Phase 2 is about efficiency to make those discoveries economically viable.
Why Does This Matter For Software?
While AI NRR might offset weak gross margins… those NRR scenarios rely on world-class, possibly unsustainable longer-term rates:
High gross margins have underpinned the business quality and investment case for software:
ICONIQ (h/t Vivian Guo) recently released their “State of AI” report (47 slides, ungated).
Key call-outs: while AI product gross margins are projected to increase by ~1,400 bps (‘25: 45% → ‘26e: 53% → 2027e: 59%):
… the fabulously granular cost structure breakdown makes those margin increases look optimistic. “Talent” - a largely fixed cost - is only 28%-34% of the cost structure for the GA to Scaling stages. All other costs are highly variable, such as inference and cloud costs. Even for model training, there is little fixed-cost leverage, since it is an ongoing process.
Key message: There seem to be too few fixed costs to generate software economics.
Inference is the #1 non-talent cost at all stages (Beta, GA, Scaling). Ahead of infrastructure & cloud, data storage, and model training.
That said, the ICONIQ “State of AI” did include encouraging unit cost mitigants:
And as inference increases … in a volatile fashion, managing the cost of inference becomes even more critical.
As we researched this trend, we connected with Carmen Li, Dual CEO of Compute Exchange and Silicon Data, publisher of multiple AI market cost trend indices, including the Silicon Data LLM Token Price Index.
Silicon Data - Token Index + CEO Commentary
The above LLM Token Price Index highlights:
Increasing Costs: up 50%+ over past 6 months
Volatility: like a ~34% decline in ~1 month seen in mid-January to February. With a similar rapid decline in June as companies take steps to mitigate the “tokenmaxxing” trend.
Key points from our research exchange with Carmen Li:
1. The Subscription Illusion vs. Commodity Reality “Revenue may look subscription-like on the surface, but underneath, the cost structure increasingly behaves more like a dynamic commodity market... Token expenditure volatility increasingly resembles the pricing behavior of energy, bandwidth, or other critical industrial inputs rather than conventional cloud software pricing.”
2. The Shift to Deployment Economics “One of the biggest misconceptions in AI right now is that model capability alone determines competitive advantage. In reality, deployment economics are becoming just as important. Two companies may offer similar AI functionality while operating under completely different infrastructure economics depending on how efficiently they route workloads, what compute environments they access, and how exposed they are to pricing volatility underneath the stack.”
3. Infrastructure Arbitrage & Regional Disparity “The same inference workload can carry materially different economics depending on where it runs, what hardware it runs on, and whether compute capacity was secured through long-term agreements or acquired dynamically in tighter market conditions.”
Morgan Stanley On Gross Margins + “Tokenomics”
Morgan Stanley’s very recent + excellent “The Moat & The Journey – Assessing Software Durability + Growth in the Age of AI” (170 pages; our selected excerpts here) covered both Gross Margins and “Tokenmaxxing” Mitigation:
MS View Recap: Some gross margin dilution offset at EBIT level; levers (like model routing or open-source) help mitigate AI COGS impact.
MS View Recap: “Token scrutiny is real, but not uniform.” + “Workflow-rich incumbents become the optimization layer: routing simple work to cheap models, reserving frontier for risk, capping low-value usage.”
With a useful - at least for our hyper-granular readers - sensitivity table wherein gross margin declines are/can be mitigated at the EBIT level:
Lazy AI: A Pet Theory
On our recent podcast with Tim Sanders of G2, he emphasized why AI engines are incentivized to provide high-quality answers (their own brand, trust, and competition).
However, a pet theory of mine relates to the AI engines’ strong incentives to be lazy. See these inference costs. Think through the costs of scaling web search and crawling. Especially into higher-inference knowledge sources like audio and video.
After all, what do I know if they missed the equivalent of page 16 of Google Search results?
AI engine outbound web traffic data, in absolute terms, suggests very few are clicking on the provided source links. If not even checking sources, are users exploring what was missed?
Meaning AI engines can (and should?) be lazy.
For our institutional investor audience, there are read-throughs to the respective moats and unit economics of Google, OpenAI, and Anthropic if this theory is right.
SPOTIFY | APPLE | OTHER PODCAST PLATFORMS| VIDEO
About Cloud Ratings
In mid-2024, we announced a research partnership with G2 - more here:
with this slide showing how our G2-enhanced Quadrants (like our recent Sales Compensation Software) release, this business of software newsletter you are reading, our podcasts, and our True ROI practice area all fit within our modern analyst firm:













