The trend: commercial use dictates the silicon
Roughly 90% of leading GPU revenue now comes from AI, not traditional HPC. That mix, not the needs of science, is what the newest chips are designed around — enormous low-precision throughput for AI, at the direct expense of the double-precision (FP64) math that simulation and modeling depend on. The evidence is blunt: on the most recent top-end AI part, peak FP64 fell to about 1.2 teraflops (from roughly 37 on the prior generation), and the newest AI flagship's FP64 actually trails the previous generation — so, for HPC, the older silicon is now the better buy. In the same chips, low-precision (FP4) throughput has climbed past 14 petaflops. The market builds the die; HPC rides along.
FP64 is a high cost for low commercial return
Dedicated double-precision units are die area and power spent on a small market. Commercially that is irrational, so vendors are cutting native FP64 hardware and pushing emulation instead — reconstructing high-precision results from low-precision tensor cores. Call it the post-FP64 posture: keep the accuracy available, but stop paying silicon for it.
But accuracy is still necessary — for reproducible, predictive science
Here is the tension. Simulation, climate and fluid dynamics, fusion, inverse problems, and any model whose answers must be predictive and reproducible still need genuine numerical accuracy — native FP64, or carefully controlled mixed precision with error bounds you can defend. "Good enough for a plausible token" is not "good enough for a result someone will build on." The commercial trend leaves a widening gap: the hardware the market optimizes for is optimized away from the precision reproducible science requires — and emulation shifts the burden onto whoever can prove the emulated answer is actually correct.
So how does the GPU cloud make money — on trailing technology?
Mostly by inference. Inference is now about two-thirds of all AI compute (up from a third a few years ago) and is heading toward 80–90% of an AI system's lifetime cost, because it runs continuously. Crucially, inference lives on the trailing edge: it prizes tokens-per-dollar, so it runs happily on last-generation, cheaper GPUs at low precision. The newest silicon and its frontier training runs are the capital-intensive tip; the recurring margin is inference on depreciating hardware. That is how the depreciation clock slows — an old card that no longer trains competitively keeps earning on inference for years. A GPU cloud is, in large part, a business that buys the frontier and monetizes the trailing edge.
Where that leaves the field (by archetype, not by name)
Sort the providers by capital access, customer diversification, leverage, and differentiation, and they fall out into: a public bellwether (largest, but most leveraged and most concentrated); a deep-pocketed spinout (sturdiest balance sheet); an energy-edge builder (a power moat rather than a pure GPU lease); a platform play (sells inference and tooling, so it is less capex-heavy and less exposed to the clock); and a long tail of sub-scale renters most exposed when supply, price, and depreciation are tested together. The survivors are the ones riding the low-precision inference wave on trailing hardware with capital and a moat behind them; the fragile ones are betting on the frontier alone.
The honest takeaway
The market optimizes silicon for cheap, high-volume, low-precision inference, and the GPU cloud monetizes the trailing edge of exactly that. Meanwhile the precision and reproducibility that predictive science needs is increasingly emulated, not built. In both worlds the scarce, valuable thing is the same: not peak FLOPS, but knowing what accuracy you are actually getting, at what cost — the accuracy-versus-cost frontier — and extracting honest, useful, reproducible work from hardware the market is optimizing for something else. Utilization and goodput decide who is running a business; measured accuracy-per-dollar decides whose results you can trust.
Market and vendor figures are third-party/analyst- and vendor-reported and treated as claims, not audited facts; company-specific characterizations are deliberately generalized. SV Advanced Computing LLC builds honest, vendor-neutral performance, accuracy, and reproducibility measurement for HPC and AI infrastructure — that's the work we do.