// HACKER NEWS — CYBERSECURITY
The Economics of Open-Weight Inference
How open-weight demand can support the useful life of NVIDIA GPU families
GPUs are commonly depreciated on the assumption that each new NVIDIA generation renders the previous one obsolete. In this paper, we provide a counter to this thesis by examining the effect of open-weight demand on the economic usefulness of older GPU families. Closed-model access runs through subscription allowances that the provider is able to reset, so the posted token rate card represents the marginal price of additional usage. Across eleven open-weight and eight closed models on the Artificial Analysis Intelligence Index, the cheapest qualifying open-weight model, standardized by intelligence, completes a task at roughly one fifth of the cost of a comparable closed model. Self-hosting on rented hardware lowers this to $0.12 to $0.35 per million output tokens at full utilization and reverses the hardware ranking: on gpt-oss-120b, a sparse model with 5.1 billion active parameters, the A100 produces output more cheaply than the H100 at spot and at the three- and five-year term prices. Ornn’s rental data show the market reflecting this utility. The five-year A100 rental price maintains 80 percent of its one-month term price (vs 44 to 60 percent for the Hopper and Blackwell families) for a contract ending when the Ampere family is more than eleven years old. We show that today’s compute-intensive workloads—long-running agents, batch evaluation, and reinforcement learning—tolerate latency and are hardware agnostic, which incentivizes price-elastic demand to route to any cost-efficient hardware. See NVIDIA’s acquisition of Hugging Face on 3 September 2026 (NVIDIA, 2026c). These findings challenge forecasts that newer hardware eliminates the earning capacity of older GPUs. Instead, they suggest that older NVIDIA generations retain a multi-year earning life so long as they serve suitable workloads competitively and operators remain free to deploy those workloads on them.
Source: Ornn Data calculation from published throughput and 1 September 2026 spot rents (paper Table 7). Full use takes Offline throughput with no headroom. The base case takes Server throughput where available and otherwise rescales the GPUStack baseline; the latter does not establish a Server service level. Dense A100 is estimated. The A100/H100 sparse rows are third-party measurements, not MLPerf results. Costs are compute-only USD per million output tokens.
Source: Ornn occupancy and listed-capacity series and the settled spot index (paper Table 8), retrieved 2 September 2026. Occupancy is rented capacity divided by listed capacity across tracked on-demand providers in the global region. Listed capacity measures tracked on-demand supply, not the installed base.
Source: Ornn reported term marks (paper Table 9). A100, H100, H200, and B200 marks were published 13 August 2026; B300 on 1 September 2026. Retention and implied-forward ratios are calculations from those marks. Ages use an illustrative 1 September 2026 start measured from family announcement dates. Forward marks are analyst-produced indicators, not executable quotes.
Closed models are available only through deployments authorized by their developers. Open weights allow independent deployment and, where compatible endpoints exist, a choice of providers serving the same checkpoint. Reported September 2026 OpenRouter snapshots listed eighteen to twenty-two providers for several widely served open models, with highest-to-lowest output-price ratios between 1.8 and 5.6. Those figures illustrate price variation, not interchangeable service: quantization, capacity, reliability, and migration costs can limit substitution.
The same choice extends to hardware. Memory, numerical format, software, and licensing determine which deployments are feasible. gpt-oss-120b uses MXFP4 expert weights and has been served on a single 80 GB A100. Operators can therefore consider older hardware where the model fits, the software supports it efficient