
Runway Intelligence is OpenMetal’s executive insight series for late-stage startups and their investors, exploring how cloud economics, infrastructure design, and operational strategy shape valuation, margins, and time to exit.
The AI infrastructure market is splitting into hyperscale capital at one end and commodity rental at the other. Sustained inference, the workload that now drives the spend, belongs cleanly in neither.
Key Takeaways
- Neocloud revenue is on track to reach roughly $180B by 2030 (Synergy Research), growing more than 200% a year from about $23B in 2025, and inference is projected to be roughly 80% of the market by 2030 (ABI Research). The center of gravity is moving from training bursts to steady-state serving.
- The contest at the top is now capital and power, not chip count. Combined 2026 capital spending by the four largest hyperscalers is running near $725B, up roughly 77% on a base of about $410B in 2025, and individual neocloud buildouts are contracting power by the gigawatt.
- The spend itself is shifting from training to inference: Deloitte estimates inference was about half of AI compute in 2025 and rises toward two-thirds in 2026. The workload that dominates spend is the one both market poles serve worst.
- Memory, not chips, is the binding supply constraint. Micron’s high-bandwidth memory is sold out through 2026, industry leaders now expect no meaningful relief before 2028, and HBM is 30 to 40% of an accelerator’s cost, so access to scale is gated by allocation, not order size.
- For diligence, exposure to metered GPU-hour variance is a forecastable margin risk, not a convenience, and capital markets have begun to treat utilized inference capacity as a financeable asset in its own right (the first inference-chip-backed loan closed in July 2026). Predictable dollars per sustained-inference-hour is the defensible position through 2030.
The most valuable position in AI infrastructure through 2030 will not belong to whoever owns the most GPUs. It will belong to whoever can price sustained inference in dollars a finance team can forecast. That is a different race than the one making headlines, and it is being run in a part of the market the headlines ignore.
Watch where the money and the power are going. Combined 2026 capital spending by the four largest hyperscalers is running near $725B, up roughly 77% on a base of about $410B in 2025, and the specialist neoclouds are matching the posture with balance-sheet commitments and gigawatt power contracts rather than incremental hardware orders. On their most recent earnings calls, CoreWeave guided 2026 spending to $31 to $35B against a 1.7GW power target, Nebius to $20 to $25B with 4GW contracted, and Meta to $125 to $145B. The moat at the top of this market is no longer access to chips. It is access to capital and to megawatts.
The race at the top is capital and power, not chips
Once the constraint becomes power and financing, the winners are decided by who can underwrite multi-year, multi-gigawatt commitments. That is a game for a handful of players, and it is increasingly a credit story: when Meta signaled the scale of its own compute buildout, it moved a competitor’s bonds, and a chipmaker took a direct equity stake in another neocloud to secure the relationship. This is what a capital arms race looks like when it reaches the financing markets.
It is also a race being run ahead of profits, and against the grid. The largest specialists are not yet profitable on a GAAP basis: CoreWeave’s depreciation and amortization alone exceeded half its revenue in the first quarter of 2026, and Nebius is not expected to post a GAAP profit before 2028. That gap is why the neocloud trade wobbled in mid-July, with Nebius falling about 13% on July 16. Money is not the only gate. Gartner projects power availability will operationally constrain 40% of AI data centers by 2027, and in the busiest hubs a new grid connection can now carry a four-to-seven-year wait. Entry at this end is gated by megawatts and balance-sheet endurance, not by the ability to place a hardware order.
The supply chain reinforces the concentration. The bottleneck has moved from the GPUs themselves, whose lead times have largely normalized, to the memory stacked on them. Micron has said its high-bandwidth memory is sold out through 2026 under multi-year hyperscaler contracts, HBM is now 30 to 40% of an accelerator’s cost, and HBM3E stacks carry lead times on the order of 20 to 26 weeks. With new fabrication capacity not arriving in volume before late 2027, industry leaders now expect no meaningful relief until 2028, which makes this a structural reallocation of supply rather than a passing spike. Capacity is allocated to whoever committed earliest and largest. For a company that is not spending tens of billions a year, the practical takeaway is simple: you will not win by out-buying the hyperscalers, and you should not build a plan that assumes you can.
Inference is where the spend is going, and it behaves differently
The second shift matters more for most operators. Budgets are moving decisively from training to inference. Deloitte estimates that inference accounted for about half of AI compute in 2025 and rises toward two-thirds in 2026, and a DigitalOcean survey early in 2026 found 44% of respondents already directing three quarters or more of their AI budget to inference. Looking further out, ABI Research projects inference will be roughly 80% of the neocloud market by 2030.
Inference is not a bigger version of training. Training is bursty, tolerant of interruption, and a natural fit for spot-style rental where the buyer trades predictability for a low headline rate. Sustained inference is the opposite. It runs continuously, it sits directly in the path of a product’s gross margin, and its cost has to be forecastable a quarter or a year out. A workload that runs every hour of every day is precisely the one you least want priced on a meter that moves. This is the mechanism behind what we have called the inference tax: AI features that ship on variable-cost infrastructure quietly convert into a permanent drag on margin.
The middle the market forgot
Put the two shifts together and the gap is obvious. At one pole sits hyperscale, where you rent into a capital-and-power machine you cannot match and accept its pricing model. At the other sits commodity GPU-hour rental, cheap at the headline but metered, variable, and subject to the same allocation crunch when supply tightens. Neither pole is built for the company running steady inference at real but not hyperscale volume: the mid-market SaaS business, the AI-native startup past its first product, the PE-backed portfolio company with a defined inference footprint and a board that wants a number it can trust.
That the neocloud market as a whole is growing more than 200% a year (Synergy Research) tells you the demand is not theoretical, and the fastest-growing part of it is the steady inference the middle runs. What the middle needs is not more raw scale. It is dedicated capacity priced in predictable dollars, from a provider whose economics it can actually see.
How OpenMetal Fits Into the Neocloud Middle
For a company whose inference footprint is steady and whose finance team needs to forecast it, the practical alternative to both poles is dedicated infrastructure billed at a fixed monthly rate. OpenMetal runs on single-tenant bare metal with an OpenStack control plane and Ceph storage the operator can actually inspect, so there is no shared-tenancy variance and no per-GPU-hour meter to model. A production-ready private cloud can be provisioned in under a minute, as little as 45 seconds, which gives a team room to add capacity on its own terms rather than renting elasticity by the hour. Block and object storage and private inter-server networking are included rather than itemized, public egress is included rather than billed as a separate line, and the invoice is a flat monthly figure a controller can drop into a model without a variance assumption. GPU bare metal is available for inference at dedicated-hardware economics rather than hyperscaler markup; contact the OpenMetal team for a configuration and quote.
When capacity is a fixed monthly line rather than a moving meter, utilization is yours to capture: every additional inference served on hardware you already pay for lowers the cost per request instead of adding to a bill. That is the inverse of the idle-silicon tax. Renting a GPU by the hour leaves the utilization upside with the landlord; owning the box, on predictable terms, turns high utilization from an operating metric into a balance-sheet one.
The number that matters here is not a discount on a chip, it is the variance you remove from the model. Teams moving off VMware or hyperscaler setups to dedicated OpenStack often report meaningfully lower infrastructure cost, and, more to the point for inference, they trade a moving rate for a fixed one. Because the platform is OpenStack and Ceph rather than a proprietary black box, that fixed cost also comes with economics the operator can audit line by line, which is the transparency a diligence process actually wants. When the binding question is “what will this workload cost next quarter,” a fixed monthly line is worth more than a low but variable one. That is the same predictability argument behind two-tier portfolios, applied to the workload now consuming the budget.
What This Means for PE Analysts
Treat infrastructure posture as a market-structure bet, not just a cost line. A portfolio company that has anchored its sustained inference on metered GPU-hour rental has taken on an unpriced exposure: its unit economics move with a market that is currently supply-constrained and capital-driven. That is a forecastable risk you can raise in diligence and a lever you can pull post-close. Ask where each company’s inference runs, on what pricing model, and what a supply crunch or a rate change does to gross margin at scale.
There is a mirror image to that exposure worth watching. In mid-July a lender extended a facility of up to $400 million to an inference-cloud operator, secured by the inference silicon itself, in what was reported as the first financing of its kind. The signal for diligence is that capital markets are beginning to treat sustained, well-utilized inference capacity as a financeable, cash-generating asset rather than a pure cost center. That cuts both ways: a saturated dedicated fleet on predictable economics behaves like an asset you can underwrite, while an inference footprint on a moving meter is the harder cash flow to model and the weaker thing to lend against. Predictability is no longer only a margin question. It is becoming a financing one.
The strategic read is that the AI infrastructure market is not a single race to the largest cluster. It is bifurcating, and the underserved middle, dedicated and cost-transparent, is where most portfolio companies actually live. The winning move is not to chase scale you cannot afford. It is to buy predictability where the spend is heading, and to make that posture a standard part of how you evaluate and improve the companies you own.
For teams whose inference footprint is steady enough to be an asset rather than a variable line item, OpenMetal’s Bare Metal Servers and Hosted Private Cloud turn high utilization into your margin instead of the meter’s. Talk to an OpenMetal infrastructure engineer when you’re ready to move off the hourly clock.
Sources
- Synergy Research Group, “Neoclouds Currently Growing by Over 200% per Year; Will Reach $180 Billion in Revenues by 2030,” October 2025. https://www.srgresearch.com/articles/neoclouds-currently-growing-by-over-200-per-year-will-reach-180-billion-in-revenues-by-2030
- ABI Research, “The State of Neocloud: Four Trends for 2026,” 2026. https://www.abiresearch.com/blog/neocloud-market-trends
- Deloitte, “2026 Technology, Media and Telecommunications Predictions: Why AI’s Next Phase Will Likely Demand More Computational Power, Not Less,” November 2025. https://www.deloitte.com/us/en/insights/industry/technology/technology-media-and-telecom-predictions/2026/compute-power-ai.html
- DigitalOcean, “Currents Research, February 2026: AI Agents, Inference and the Widening Adoption Gap,” February 2026. https://www.digitalocean.com/currents/february-2026
- Investing.com, “Micron Faces a Rerating Moment as Sold-Out HBM Supply Reshapes the Earnings Story,” 2026. https://www.investing.com/analysis/micron-faces-a-rerating-moment-as-soldout-hbm-supply-reshapes-the-earnings-story-200676155
- CoreWeave Inc., “CoreWeave Reports Strong First Quarter 2026 Results,” U.S. Securities and Exchange Commission, May 2026. https://www.sec.gov/Archives/edgar/data/1769628/000176962826000220/coreweave1q26earningspress.htm
- CNBC, “CoreWeave stock sinks 10% on weak revenue guidance, increased spending forecast,” May 7, 2026. https://www.cnbc.com/2026/05/07/coreweave-crwv-q1-earnings-report-2026.html
- Yahoo Finance, “Nebius Raises Capex to $20-$25B: A Bold Growth Move or Risky Bet?” May 2026. https://finance.yahoo.com/markets/stocks/articles/nebius-raises-capex-20-25b-125500935.html
- Yahoo Finance, “Meta stock sinks after Q1 earnings as company raises 2026 AI spending forecast to $125 billion-$145 billion,” April 2026. https://finance.yahoo.com/sectors/technology/article/meta-stock-sinks-after-q1-earnings-as-company-raises-2026-ai-spending-forecast-to-125-billion-145-billion-160136308.html
- The Motley Fool, “Better Datacenter Stock: CoreWeave (CRWV) or Nebius (NBIS)?” May 14, 2026. https://www.fool.com/investing/2026/05/14/better-datacenter-stock-coreweave-or-nebius/
- TechCrunch, “Why the first GPU financiers are turning to inference chips in a $400 million deal,” July 17, 2026. https://techcrunch.com/2026/07/17/why-the-first-gpu-financiers-are-turning-to-inference-chips-in-a-400-million-deal/
- 24/7 Wall St., “Nebius Sinks 13% as the Neocloud Trade Unravels,” July 16, 2026. https://247wallst.com/investing/2026/07/16/nebius-sinks-13-as-the-neocloud-trade-unravels-how-coreweave-iren-and-the-ai-data-center-stocks-stack-up/
- IEEE Spectrum, “AI Boom Fuels DRAM Shortage and Price Surge,” 2026. https://spectrum.ieee.org/dram-shortage
- Gartner, “Gartner Predicts Power Shortages Will Restrict 40% of AI Data Centers by 2027,” Nov. 12, 2024. https://www.gartner.com/en/newsroom/press-releases/2024-11-12-gartner-predicts-power-shortages-will-restrict-40-percent-of-ai-data-centers-by-20270




































