The Inference Tax How AI Feature Debt Is Destroying Your Portfolio's Gross Margins

Runway Intelligence is OpenMetal’s executive insight series for late-stage startups and their investors, exploring how cloud economics, infrastructure design, and operational strategy shape valuation, margins, and time to exit. 

AI-first SaaS companies are reporting 52% gross margins, more than 20 points below the 75-85% range that built the SaaS valuation premium. The cause isn’t just market pressure. It’s an infrastructure decision being made by default.

Key Takeaways

  • AI-first SaaS gross margins have compressed to 52% as inference costs consume 23% of revenue at scale, the infrastructure cost of shipping AI features is now a primary financial risk variable for PE portfolios
  • Inference COGS is capitalized at exit, so restructuring hyperscaler GPU spend into a fixed-cost alternative compounds directly into enterprise value at the multiple the business trades on
  • AWS raised reserved GPU capacity prices ~15% in early 2026, portfolio companies on on-demand contracts have no mechanism to hedge against further increases
  • The inference tax is not a market condition. It is a procurement decision; dedicated bare metal GPU replaces the variable per-GPU-hour meter of hyperscaler on-demand with fixed monthly pricing
  • Operating partners who include infrastructure cost structure in portfolio reviews 18–24 months before a process show buyers clean gross margins; those who don’t show buyers a COGS explanation

AI-first SaaS companies are reporting gross margins of 52%, more than 20 points below the 75-85% range that traditional SaaS has maintained for years (ICONIQ State of AI: Bi-Annual Snapshot, January 2026). The cause isn’t market pressure, pricing strategy, or customer acquisition cost. It’s the inference tax: the compounding COGS burden of running AI features on hyperscaler GPU infrastructure at rates the margin model was never built to absorb.

At the portfolio level, this is not an operational footnote. It is the most concrete mechanism by which AI-augmented SaaS companies are repricing themselves out of premium exit multiples, and most PE operating partners are watching it happen in the gross margin line without a framework for attributing it to a controllable infrastructure decision.

The math is direct: ICONIQ’s analysis found that inference costs now represent 23% of revenue at scale for AI-first SaaS companies. For a company at $30M ARR, that’s $6.9M annually in compute spend running at hyperscaler markup, and at 24-25x gross profit multiples, every dollar of that spend that could be structured differently represents $24-25 of exit market cap that isn’t being captured.

Why the Math Gets Worse at Scale

Every AI feature that ships is a commitment to ongoing inference compute. Unlike traditional SaaS costs, which tend to grow sub-linearly with revenue as companies achieve economies of scale, inference costs scale with usage, and usage scales with product success. A company that builds a successful AI feature on hyperscaler on-demand GPU is committing to an indefinite, variable, and upward-trending COGS line that is directly correlated with the product engagement it most wants to grow.

AWS raised prices on its reserved GPU capacity (EC2 Capacity Blocks for ML) by approximately 15% in early 2026, driven by enterprise AI demand and constrained supply (Techzine, January 2026). There is no mechanism within a hyperscaler contract that isolates a portfolio company from those market price movements. The company pays the prevailing rate. And the rate is moving against them precisely when they are most successful, when inference volume is highest.

For a PE firm benchmarking portfolio companies on gross margin trajectory, the question is not whether a company has AI features. The question is whether those AI features are built on an infrastructure cost structure that holds under buyer scrutiny at exit.

The Infrastructure Decision That Isn’t Being Treated as a Financial One

The root cause of the inference tax is a decision that rarely involves a CFO or operating partner: which cloud infrastructure runs the AI inference layer. Engineering teams default to hyperscaler on-demand GPU because it is immediately available, requires no capital commitment, and is operationally familiar. The cost structure that decision creates (variable, usage-correlated, hyperscaler-priced) is an afterthought.

In the traditional SaaS world, that decision had limited gross margin consequence. Compute was cheap relative to revenue, and the optimization opportunity was a rounding error. In an AI-augmented SaaS world, where inference costs represent 23% of revenue at scale, the decision is a core financial decision being made by engineers without financial input. That structural gap is the inference tax.

PE operating partners who close that gap (by including infrastructure cost structure in portfolio reviews, benchmarking inference COGS against dedicated alternatives, and creating a governance path for infrastructure cost decisions) are the ones whose companies show up to a sale process with defensible gross margins rather than a COGS explanation.

OpenMetal: A Fixed-Cost Alternative to the Inference Tax

For portfolio companies running AI inference on hyperscaler on-demand GPU, the practical alternative is dedicated bare metal GPU infrastructure at fixed monthly pricing, the kind of cost structure that moves inference costs out of variable COGS and makes them predictable, stable, and auditable.

OpenMetal’s dedicated GPU bare metal infrastructure, available as single-tenant dedicated GPU servers with fixed monthly billing, runs at fixed monthly pricing rather than a variable per-GPU-hour rate. That cost-structure difference flows entirely to gross margin. For a portfolio company running AI inference spend on hyperscaler infrastructure, migrating to dedicated bare metal moves recoverable gross profit onto the margin line that buyers capitalize at exit, so the benefit compounds into enterprise value at the multiple the business trades on.

Single-tenant bare metal gives operators full visibility into what they’re running, eliminates noisy-neighbor interference on inference latency, and carries no egress fee structure that compounds with usage volume. For portfolio companies where inference reliability is a product quality requirement, not just a cost concern, dedicated bare metal resolves both problems simultaneously.

For operating partners addressing inference cost exposure across a vintage, the migration pattern is repeatable: identify inference-heavy workloads, benchmark current hyperscaler spend against dedicated alternatives, and scope a 90–180 day migration that produces at least two clean quarters before any sale process begins. Contact the OpenMetal team to start with a cost audit against your current inference spend.

“The go-to option for battling the high costs of public clouds.”

Chris Ueland, Hunt Intelligence

What This Means for PE Analysts

The inference tax is a controllable cost. It is not a market condition, a competitive dynamic, or a product architecture constraint. It is an infrastructure procurement decision that is currently being made by default, and the default is the most expensive option.

For every portfolio company with AI features in market, the gross margin analysis should include an inference cost review: what is the current spend, what infrastructure is it running on, and what is the fully loaded cost of that choice at current exit multiples. Where the math reveals exposure, the intervention is straightforward, the timeline is predictable, and the gross margin recovery is quantifiable before the first investor meeting.

Portfolio companies that address the inference tax 18–24 months before a sale process show buyers clean gross margins with a documented, stable cost structure. Those that don’t show buyers a COGS line that correlates with product success, and buyers price that risk accordingly.


Sources

  • ICONIQ Capital, “State of AI: Bi-Annual Snapshot, The Execution Era of AI,” January 2026. iconiq.com
  • SaaStr, “Inference Costs Average 23% of Revenue at AI B2B Companies,” 2026. saastr.com
  • Techzine, “AWS increases EC2 Capacity Block prices by 15 percent,” January 2026. techzine.eu