The AI value chain, measured from the silicon up
Almost everything published on AI demand starts with a growth rate and works forward. This model works the other way. It starts with the chips companies already reported to the SEC, then derives the tokens, the memory, the gigawatts and the money from there. Every number on this page traces back to a public filing or a calculation you can follow. If a term is new to you, it's explained the first time it appears.
All three scenarios run in parallel on the same formulas. The only thing that changes is 43 assumption levers. The shaded bands on the charts are always the full bear-bull range, whichever scenario you pick.
Short scale: billion = 10⁹ · trillion = 10¹² · quadrillion = 10¹⁵. All figures in US dollars.
Three things the model says and the consensus doesn't
If you only read one section, make it this one. Three conclusions, each derived — not assumed — from the chain below.
The model runs in one direction only: from audited silicon to disputed dollars. Chips sold set the fleet, the fleet sets the tokens, and the tokens set the memory, the gigawatts and finally the money. Run it forward on 43 labeled assumptions and the base case reaches US$1.75 trillion of total AI revenue in 2030. Three conclusions fall out of that arithmetic that the consensus does not carry. Each one can be checked against a public number.
Energy binds first. The silicon these companies already bought still has to be plugged into a grid. Plug all of it in and, in the base case, AI datacenters in 2030 need about three times the electricity the International Energy Agency (IEA) budgets for them. They also need more than the agency budgets for every datacenter on Earth, AI or not. Even the bear case, on a much smaller fleet, does not fit. So the ceiling is physical, and it arrives before the financial one. What binds is not capital. It is the plug, and the rent goes to whoever controls it.
S30S32Then the unit economics. About 45% of the tokens the fleet processes are never invoiced. They are internal company use, free tiers and test traffic. That share is measured where it can be measured: only 26% of Google's tokens are billed through its API, the paid interface other companies plug into. Spread the full yearly cost of the infrastructure across only the tokens that do get billed, and in 2030 the base case still loses money on each one. The price it sells for does not cover what it costs to make. Selling tokens, in other words, does not repay the machines that produce them. If this build-out earns a return, it is earned downstream, in the products the tokens power. Not at the token counter.
And then the variable that decides everything. Move from the bear scenario to the bull and the tokens processed rise ninefold. AI revenue rises a hundred and forty-fold. That gap is not volume — it is price. It turns on whether the price per token keeps falling about 30% a year, as it has, or stops falling altogether (lever B5). I think price per token is the input that moves this model most, and the one least discussed. The public debate stays fixated on how many tokens the world will run. The number that moves the money is what each one sells for.
Push the model to its bull scenario and AI revenue in 2030 lands at nearly one dollar of every eight the world produces. That is measured against the IMF's projection of roughly US$151 trillion of global GDP for that year. The number is absurd, and it is built to be read that way. An economy does not reorganise an eighth of everything it makes around a single input in five years. The bull case earns its place by pricing the dream, not by defending it. It marks where the arithmetic stops being physical and starts being fantasy.
S54The base case is not the cautious case, and the adoption it assumes makes that plain. MIT finds 95% of company AI pilots still show no measurable return. Goldman counts barely 2% of S&P 500 companies putting a number on what AI adds to earnings. Yet the base case needs US company adoption to climb to about 35% by 2030 (lever B6) for its revenue to hold. I leave that tension in view on purpose. The model is not smuggling in caution. It needs adoption to reach a place no measurement has reached yet, and it asks you to judge whether it gets there.
The value chain
Each link is derived from the one before it; the only assumptions sit at the tail, and they are labeled. Read top to bottom, this is the whole derivation, from audited silicon to disputed dollars.
The fact
Start where the numbers are audited and the first thing you find is that the forecasts are already behind. Omdia projected a US$286bn market for AI chips in 2030. Reported sales passed that figure in 2026, four years early. That is the argument for deriving this chain from filings instead of from anyone's growth rate.
S14Nobody reports chips, they report dollars, so units are solved from revenue. NVIDIA's datacenter segment billed US$162.4bn in its fiscal 2025. At about US$42,000 per chip, that is 3.87M chips sold in the year. For 2026 the four buckets (NVIDIA, AMD, the custom chips cloud firms design for themselves, and China) sum to US$381bn of accelerators. Then comes networking: the switches, cables and cards that wire thousands of chips into one machine. NVIDIA reports it at 24.5% of accelerator revenue, which lifts the total to US$474bn, within a few percent of what the four report to the SEC.
The industry frames 2026 as a growth story. It is a price story. AI-server shipments grow 28% while silicon revenue climbs about 90%, because the average rack price tripled in two generations (GB200 near US$3.0M, GB300 near US$4.0M, Vera Rubin up to US$8.8M). NVIDIA is already touting US$1 trillion of Blackwell and Rubin orders through 2027, more than Switzerland's GDP.
S10So 2026 is as much a price cycle as a unit cycle. The number to watch each quarter is datacenter revenue against the price-per-chip assumption, not the unit-count headlines that double-count multi-die packages.
Nobody reports chips; they report dollars. Units are solved from revenue.
units = reported revenue / price per chip -- NVIDIA 2025: US$162.4bn / US$42,000 = 3.87M chips
2026, the four buckets, and the sum that validates against the SEC within -1%.
NVIDIA 4.83M x $60k + AMD 1.40M x $12.5k + ASIC 5.10M x $11k + China 1.75M x $10k = US$381bn -- + networking 24.5% = US$474bn
| Metric | 2023 | 2024 | 2025 | 2026 | 2027 | 2028 | 2029 | 2030 |
|---|---|---|---|---|---|---|---|---|
| Accelerator units shippedmillion unitsS01S02S03S04S05S06S07S08S09 | 3.18 | 6.17 | 8.95 | 13.1 | 15.6 | 18.6 | 22.2 | 26.6 |
| Accelerator revenueUS$ bnS01S02S03S04 | 49.0 | 118 | 198 | 381 | 435 | 498 | 570 | 653 |
| AI datacenter networkingUS$ bnS01S02 | 9.10 | 15.0 | 38.3 | 93.3 | 95.8 | 110 | 125 | 144 |
| Total AI siliconUS$ bnS01S02S03S04S05 | 58.1 | 133 | 237 | 474 | 531 | 608 | 695 | 797 |
| Installed fleet, year-endmillion H100-eqS06S31 | 2.06 | 7.49 | 23.7 | 58.7 | 116.5 | 210 | 362.4 | 607.6 |
| Average operating fleetmillion H100-eqS31 | 1.03 | 4.78 | 15.6 | 41.2 | 87.6 | 164.3 | 288.9 | 493.1 |
Reported ≤ 2025 · Projected ≥ 2026
A common currency for silicon
A 2023 GPU and a 2026 rack are not the same unit, so every chip generation is converted into H100-equivalents using public benchmarks. One H100-equivalent means as much computing as a single 2023 H100 chip, which lets silicon from different years add up in one honest unit.
The conversion indices are measured, not assumed. The H100 sits at 1.00. The B200 at 2.90 (SemiAnalysis clocks its cost from US$0.49 to US$0.17 per million tokens). The GB300 at 4.50 (MLPerf shows 45% more than the GB200). The unit tracks work actually delivered, not the peak speed on the spec sheet.
This is also where the loudest counting error hides. Headlines count 6M GPUs where the filings describe 3M packages, because the newer parts carry two chips each. Halving the count halves the implied price per chip, and that error then runs through every estimate below it. The model counts packages, exactly as the filings do.
Each generation converts into a measured, not assumed, common currency.
H100-eq = units x speed index -- H100 = 1.00 · B200 = 2.90 (SemiAnalysis: US$0.49 -> US$0.17 per M tokens) · GB300 = 4.50 (MLPerf: +45% over GB200)
The living fleet
Chips sold are not chips running, so the fleet is built as a sum with memory. Each year's purchases come in; silicon past its service life goes out. What comes in is the shipment count from link 01, converted at link 02's indices. The 13.08M packages sold in 2026 are worth 35.0M H100-equivalents, because the average part of that year does 2.7 times the work of a 2023 H100. The second chart below shows those annual additions. Add the cohorts (2.1M from 2023, 5.4M from 2024, 16.2M from 2025 and 35.0M from 2026) and 58.7M H100-equivalents are standing by the end of 2026.
Not all of that runs a full year. A chip delivered in March cannot work twelve months, so new silicon is assumed to run half the year in its first year. The working fleet in 2026 is therefore 23.7M carried from end-2025 plus half of the 35.0M just delivered: 41.2M H100-equivalents actually running. How long a chip lasts is lever A7.
That lever is not a detail. Service life is the number-one variable of total cost, and the hyperscalers (Microsoft, Alphabet, Amazon and Meta, who rent out most of the world's computing) have already stretched it from 3 to 5-6 years in their accounts. Every extra year cuts the annual cost without moving a dollar of spending. It belongs on the dashboard as a lever, not buried in a footnote.
S31The fleet is a sum with memory: purchases come in, silicon past its service life goes out.
fleet 2026 = 2.1M ('23) + 5.4M ('24) + 16.2M ('25) + 35.0M ('26) = 58.7M H100-eq
A chip delivered in March doesn't work the whole year: new silicon runs at 50% duty its first year.
operating fleet 2026 = 23.7M (end of 2025) + 50% x 35.0M = 41.2M H100-eq
Each year's additions are the shipment count solved in link 01: reported revenue divided by price per chip. Four buckets add up here: NVIDIA, AMD, the custom chips cloud firms design for themselves, and China. 2023 to 2025 are reported quarters, 2026 is reported through June plus guidance, and 2027 onward grows at lever A1. Link 02 then converts the packages into H100-equivalents. That is why 13.08M packages in 2026 become 35.0M units of 2023-grade computing in the fleet above.
Tokens
A token is AI's unit of work, roughly a word read or written. Rather than trust what the chips could do on paper, the fleet's output is calibrated backwards against the one volume anyone has disclosed: Google's 3.2 quadrillion tokens per month. The model solves how much each chip produces instead of assuming it.
The arithmetic is direct. Take the global rate for 2026, 6.0 quadrillion tokens a month, multiply by twelve, and divide by the 41.2M working fleet. That gives 1.75 billion tokens per chip per year, solved rather than assumed. But the 6.0 is my own estimate, not a disclosure. It is built bottom up from what is public, mostly Google's 3.2 quadrillion a month and OpenAI's reported throughput, and that bottom-up work gives a range of 5.5 to 7.7 quadrillion. Choosing 6.0 is the same as saying Google runs 53% of the world's tokens. That one share is what the whole step rests on.
S40The sanity check is what makes the calibration believable. 1.75 billion tokens a year is 55 tokens per second, sustained. That is about 3% of what an H100 can do at full tilt. Low, and correctly so, because part of the fleet is training models rather than answering users. Measured efficiency gains run at 41% a year (SemiAnalysis) against the 60-70% sometimes claimed, while MLPerf shows 45% per hardware generation. I believe the measurement, not the press release.
S11S12Two conclusions follow, and they matter more than the calibration itself. Chips do keep getting cheaper per token, at a measured 41% a year, but the fleet grows faster than that. So almost every extra token comes from buying more silicon, not from better silicon. And at 3% of its paper peak, the fleet is nowhere near the ceiling of what it could print. Token supply is not what limits this industry. What each token sells for is.
The one genuine cross-check is a rival estimate, and it does not settle the question. Goldman puts the global rate at 5.0 quadrillion a month, about 20% below mine. On their number Google runs 64% of the world's tokens; on mine it runs 53%. Neither share is absurd, so I carry the gap as uncertainty rather than argue it away. A second lab disclosing its volume would settle it. What is not in doubt is the direction: OpenAI's API went from 6 to over 15 billion tokens per minute in five months. Agents change the scale outright, since one coding-agent session burns 100 to 3,500 times the tokens of a chat, already about US$13 per developer per active day.
S40S43We don't assume token volume: we solve per-chip yield from the one observed figure.
tokens per chip = (6.0 quadrillion/mo x 12) / 41.2M fleet = 1.75bn per chip per year -- solved, not assumed
Sanity check: that implies a sustained 55 tokens per second — ~3% of an H100's theoretical peak. Low, and correctly so: part of the fleet trains rather than serves.
1.75bn / 31.5M seconds = 55 tok/s
| Metric | 2023 | 2024 | 2025 | 2026 | 2027 | 2028 | 2029 | 2030 |
|---|---|---|---|---|---|---|---|---|
| Tokens processed per monthquadrillionS40 | 0.11 | 0.57 | 2.06 | 6.00 | 14.0 | 29.0 | 56.0 | 105.2 |
| Tokens processed per yearquadrillionS40 | 1.40 | 6.90 | 24.8 | 72.0 | 168.5 | 347.5 | 672.2 | 1,262 |
| Blended realized priceUS$ / M tokensThe EFFECTIVE price charged, blending expensive and cheap models, volume discounts and caching. Not list price.S46S43S49S50 | 3.44 | 1.60 | 1.62 | 1.40 | 1.23 | 1.08 | 0.95 | 0.84 |
| Model-layer revenueUS$ bnS46 | 2.60 | 6.10 | 22.1 | 55.4 | 114 | 207 | 353 | 583 |
| Total AI revenueUS$ bnModels + cloud + applications. Applies the downstream multiplier (lever B8).S46S44S48 | 7.70 | 18.2 | 66.2 | 166 | 343 | 622 | 1,058 | 1,748 |
| US enterprise adoption% of firmsS45 | 3.8% | 5.5% | 9.2% | 19.8% | 23.6% | 27.4% | 31.2% | 35% |
| Full-time agents required per workeragentsPlausibility check: to absorb every token the fleet produces, each knowledge worker would need to run this many coding agents full-time, all day.S43S45 | 0.06 | 0.22 | 0.46 | 0.61 | 1.20 | 2.12 | 3.58 | 5.96 |
Reported ≤ 2025 · Projected ≥ 2026
Billable tokens and their price
Not every token that runs is a token that sells. The 55% that does (lever B3) is not a guess, it is a blend of two measured extremes. Google bills only 26% of its tokens through its API, the paid interface other companies plug into. The rest of its inference feeds its own products, free tiers and test traffic. OpenAI and Anthropic sit near the other end, close to 100%, since almost everything they run is sold. Weighted by size, the two ends give about 55%.
Price is anchored to booked revenue, not to a run-rate (an annualised figure built by multiplying one good month by twelve). US$55bn of booked revenue in 2026 over 39.6 quadrillion billable tokens gives US$1.40 per million tokens. That price (lever B4) is the single most uncertain input in the chain, for a plain reason: nobody publishes it. No lab reports revenue divided by tokens, so the figure has to be assembled from two estimates at once, and both of them move. It is also not a list price. It blends expensive reasoning models with cheap ones, volume discounts and cached traffic, and that mix shifts every quarter.
Two forces pull it down, and the market routinely misreads both. Run-rate flatters the number. Anthropic closed 2025 at a 'US$9bn run-rate' but booked US$4.5bn of actual revenue for the year, a 2x overstatement that inflates the implied price with it. And the cheap end of the market is no longer marginal. OpenRouter is a service that routes developer traffic to whichever model the developer picks, and it publishes its mix. Open-weight and Chinese models went from under 2% of that traffic to over 45% in twelve months. Every user who moves there keeps working and pays less, which drags the blended price down without any lab cutting a published price.
S46S50It does not only fall, which is what makes it so hard to pin down. Anthropic raised Sonnet 5 by 50%, and its new tokenizer emits about 30% more tokens for the same text. That is a price increase that appears on no price list. Net price per token is the hardest number in the model to forecast, and the one that moves the answer most.
S43Only a fraction of tokens gets billed; price anchors to ACCOUNTING revenue, not run-rate.
billable = tokens x 55% (B3) -- 2026 price = US$55bn accounting revenue / 39.6 quadrillion = US$1.40 per million tokens
Revenue
Billable tokens multiplied by price gives the revenue the AI labs themselves earn, and the arithmetic here is still clean. 694 quadrillion billable tokens in 2030 at US$0.84 per million works out to US$583bn, up from US$55bn in 2026.
A multiplier then stacks cloud and applications on top of the labs (lever B8), set at 3.0x in the base case. That figure is not picked from the air. It is the ratio observed in 2026: US$166bn of AI revenue across the whole stack against US$55bn earned at the model layer. The base case holds today's ratio flat to 2030, which carries US$583bn of model revenue to US$1.75 trillion of total AI revenue. Holding it flat is itself the assumption. It says the industry keeps splitting its dollars between labs, clouds and software exactly as it splits them now. It also means both layers compound at the same rate by construction, so in the chart below they differ in level, never in slope.
The multiplier deserves to be named, but not oversold. Swing B8 alone from 2.2x to 4.0x, leaving every other lever at base, and 2030 revenue moves from US$1.28 trillion to US$2.33 trillion. The US$17.5 trillion of the bull case comes from moving all 43 levers at once, and mostly from price. Where B8 does decide the answer is the verdict: at 2.2x the chain still fails to cover its annual cost in 2030 (0.99x), at 4.0x it covers it comfortably (1.79x). One judgment about who captures the value flips whether this build-out ever pays for itself.
Two bookends show how wide that judgment can run. On one side, money already under contract but not yet earned: Google Cloud and AWS together hold about US$826bn of it, up 131% in two quarters. On the other, companies paid just US$12.5bn for language-model access in all of 2025, against roughly US$410bn of spending on the infrastructure. That is 33 to 1. The lab layer is solved. The layers above it are the real debate: the US$1.75 trillion of 2030 is a claim about who captures value downstream, not about how many tokens the fleet can print.
S47S44From tokens to dollars in two multiplications — the second is the assumption that closes the whole account.
model-layer rev 2030 = 694 quadrillion x US$0.84 = US$583bn -- total AI rev = US$583bn x 3.0 (B8) = US$1,748bn
Same formulas, same fleet. Tokens multiply by ~9 across scenarios; revenue by ~140. The difference is price per token — lever B5.
Memory
Chips carry memory by geometry, not by assumption. 13.08M packages in 2026, each holding about 255 GB of high-bandwidth memory (HBM), demand 3.34bn GB in total. That sits 11% below TrendForce's independent estimate of 3.75bn, which is the cross-check. It pulls the whole memory market from US$900bn in 2026 toward US$1.61 trillion in 2030, of which the AI share climbs from US$125bn to US$453bn.
The memory the chain drags along is no cheap accessory. DRAM contract prices multiplied 5.35x in four quarters, and Micron booked an 84.6% gross margin. That is not a typo. What drove it is a squeeze, not a fashion. HBM consumes between two and three times the silicon wafer of ordinary memory for the same gigabyte, so every stack built for AI removes capacity that would have made ordinary DRAM. Supply went short while phones, cars and ordinary servers still wanted it, and the price of everything repriced upward. The HBM market alone goes from US$42.7bn in 2026 to US$307bn in 2030. For once the bottleneck earns more than the computing that created it.
S23S25S26Here the consensus is exactly backwards, and the reason is the clock each market runs on. HBM sells on contracts fixed up to a year ahead; ordinary DRAM reprices every quarter. NVIDIA locked its 2026 HBM prices in May 2025, before the shock, so HBM has been less profitable than ordinary DDR5 memory since early 2026. Samsung's own results confirm it: ordinary DRAM now out-earns HBM even as HBM4 ramps for Vera Rubin. So the makers earning most from AI are not the HBM champions. They are Samsung and Micron, whose large ordinary-memory books reprice into the shock every quarter, while SK hynix, the nominal HBM leader, sits inside contracts signed before it. I think the profit in AI memory is not where the roadmaps point.
S22S28No opinion here: it's geometry. Every chip carries a known amount of memory stacked on it.
bits = 13.08M chips x ~255 GB average = 3.34bn GB of HBM -- TrendForce estimates 3.75: -11%
| Metric | 2023 | 2024 | 2025 | 2026 | 2027 | 2028 | 2029 | 2030 |
|---|---|---|---|---|---|---|---|---|
| HBM bits demandedbn GBS24S21 | 0.18 | 0.70 | 1.75 | 3.34 | 4.58 | 6.29 | 8.66 | 11.9 |
| HBM marketUS$ bnS27S24S26 | 1.60 | 6.70 | 25.2 | 42.7 | 147 | 187 | 240 | 307 |
| AI-attributable memoryUS$ bnS24S25S27 | 3.40 | 11.2 | 36.9 | 125 | 253 | 305 | 371 | 453 |
| Total memory marketUS$ bnS20S23S27 | 76.1 | 115 | 202 | 900 | 1,244 | 1,348 | 1,468 | 1,608 |
Reported ≤ 2025 · Projected ≥ 2026
Energy
Chips draw power by spec sheet, so this step has almost no assumption in it. Only physics and the fleet from step 3. A GB200 rack pulls 132 kW across 72 GPUs, which is 1.83 kW per chip. Multiply that by each year's shipments and the capacity added follows: 21.0 GW of new computing load in 2026, on top of everything already plugged in. The additions grow every year, because the fleet grows and each new generation draws more. Racks go from 132 kW today toward 600 kW on the roadmap.
Turning gigawatts into a yearly bill takes three steps. Multiply the running gigawatts by the 8,760 hours in a year. Add 15% for cooling and losses, which is what a PUE of 1.15 means. The letters stand for power usage effectiveness: the total power the building draws, divided by the part that actually reaches the chips. Then assume the fleet is busy 71% of the time, the utilization Epoch AI measures on real fleets (lever D2). Chips idle between jobs, and assuming they run flat out would inflate the answer by 41%. The result is 207 TWh in 2026, rising to 1,238 TWh in 2030. A number that large deserves a second route, so here is one that never touches the chip count. Take the US$757bn the four largest cloud firms spend in 2026. Divide it by the roughly US$40M it costs to build a megawatt of datacenter. That gives 18.9 GW of new capacity, against the 21.0 GW the chips imply. Two independent paths, 11% apart.
That 1,238 TWh is where the physics turns into a wall. It is 3.4x the entire electricity budget the IEA sets aside for AI datacenters in 2030, and it clears the budget the agency sets for every datacenter on Earth, AI or not. Even the bear case, at 591 TWh, does not fit. The strain is already visible in hardware. A single 2027 rack will draw the peak power of 65 homes and shed the heat of 30 gas boilers. Gas-turbine orders rose 70% in 2025, with Mitsubishi sold out to 2028. And the big five now spend more on AI in 2026 than the world invests in producing oil and gas.
S30S32Two readings, one conclusion. Either the IEA's datacenter budget is wrong by a large factor, or the industry's chip growth is physically impossible. Both make the grid the binding constraint, and both hand the rent to whoever controls the connection to it. The US queue of projects waiting to plug in already exceeds 1 terawatt, and its own operators call the 'vast majority highly speculative'. In Texas alone 410 GW sit queued, 87% of it datacenters. Reports of '50% of projects cancelled' refer to exactly that unfunded queue. Measured North-American capacity forecasts moved only about 1% in six months. The ceiling is physical. And it arrives ahead of the financial one.
S39Physics again: watts per chip, chips per year.
GW = units x kW per chip (GB200: 132 kW / 72 GPUs = 1.83 kW) -- 2026: 21.0 GW added
And from gigawatts to annual consumption, with measured utilization — never 100%.
TWh = operating GW x 8,760 h x PUE 1.15 x utilization 71% -- 2030: 1,238 TWh
Cross-check via an independent path: capex divided by cost per megawatt converges with the chip count.
US$757bn capex / US$40M per MW = 18.9 GW vs the model's 21.0 GW: +11%
| Metric | 2023 | 2024 | 2025 | 2026 | 2027 | 2028 | 2029 | 2030 |
|---|---|---|---|---|---|---|---|---|
| IT GW added in yearGWS34S32 | 1.93 | 5.26 | 11.3 | 21.0 | 27.8 | 36.9 | 49.1 | 65.3 |
| Average operating IT GWGWS31 | 0.97 | 4.56 | 12.8 | 29.0 | 53.4 | 84.8 | 124.2 | 173.1 |
| AI datacenter electricity useTWh / yrS30S31S34 | 7.00 | 33.0 | 92.0 | 207 | 382 | 607 | 888 | 1,238 |
| Gap vs. IEA electricity budgetx1.0x = the fleet fits exactly within what the IEA projects for AI datacenters. Above 1.0x, the model is asking for more electricity than is expected to exist.S30S32 | 0.13× | 0.41× | 0.77× | 1.30× | 1.82× | 2.29× | 2.82× | 3.39× |
| Annual electricity costUS$ bnS38S37 | 0.50 | 2.60 | 7.50 | 17.6 | 33.8 | 55.8 | 84.9 | 123 |
| Physical infrastructure capexUS$ bnS36S35 | 19.3 | 57.8 | 141 | 294 | 417 | 592 | 841 | 1,197 |
Reported ≤ 2025 · Projected ≥ 2026
Cost and coverage
The public debate compares one year of spending against one year of revenue and pronounces the industry underwater. That is an accounting error. Equipment is paid for once but earns for years, so its cost is spread across those years, not charged all at once. Three different numbers get confused here, and 2026 separates them cleanly. The chain pays out US$911bn of cash that year, which is revenue for the chipmakers, the memory makers, the builders and the utilities. Spread over the years the equipment earns, that same outlay is US$270bn of annual cost. That is US$216bn of silicon over five years, US$37bn of buildings over fourteen, and US$18bn of electricity, consumed in the year it is bought. And US$166bn is what the AI layer itself sells that year.
Coverage compares the third number against the second, never against the first. It is AI revenue divided by annual cost, and it says when the chain begins to pay for itself. In 2026 it is 0.62x: US$166bn of AI revenue against US$270bn of cost, or 62 cents earned per dollar of cost. It crosses 1.0x in 2029 and reaches 1.35x in 2030, as revenue of US$1.75 trillion outruns cost of US$1.30 trillion. That crossing is the moment the arithmetic starts working in the industry's favour. Cash out the door is a separate question, and it never turns positive inside this window.
Two outside markers show how fragile that crossing is. Bain sets a stiffer bar. It puts the revenue needed by 2030 at about US$2 trillion a year, and counts US$800bn of that as still missing even after efficiency savings. Measured against Bain's bar, the base case falls 13% short. The second marker points the other way. For every US$1 of annual cost their AI infrastructure carries, the hyperscalers already book US$1.19 of revenue, above 1.00 for the first time. So the crossing is real, and narrow. It rests on the price assumption that is the least certain input in the chain, and in the bear case it never happens at all: coverage ends 2030 at 0.17x.
S48The public debate compares a year's capex against a year's revenue. That's an accounting error: capex depreciates, it isn't expensed at once.
TCO 2026 = US$216bn silicon depreciation (/5 yrs) + US$37bn build (/14) + US$18bn electricity = US$270bn -- not US$911bn
With the accounting done right, coverage says when the industry starts paying for itself.
coverage = AI rev / TCO -- 2026: 166/270 = 0.62x · crosses 1.0x in 2029 · 2030: 1.35x
| Metric | 2023 | 2024 | 2025 | 2026 | 2027 | 2028 | 2029 | 2030 |
|---|---|---|---|---|---|---|---|---|
| Total value-chain spendUS$ bnS01S02S03S24S36S38 | 81.4 | 205 | 422 | 911 | 1,235 | 1,561 | 1,992 | 2,570 |
| Annual economic cost (TCO)US$ bnS31S33 | 14.2 | 49.3 | 119 | 270 | 473 | 707 | 981 | 1,300 |
| Economic coverage ratioxAI revenue ÷ annual economic cost. Above 1.0x, the industry covers its cost of capital and operations.S48S46 | 0.54× | 0.37× | 0.56× | 0.62× | 0.72× | 0.88× | 1.08× | 1.35× |
| Cost per million tokens processedUS$ / M tokS31S33 | 10.5 | 7.14 | 4.81 | 3.75 | 2.81 | 2.04 | 1.46 | 1.03 |
| Margin on selling tokens%Realized price against economic cost per BILLABLE token. Negative means selling tokens directly doesn't pay for the infrastructure — the return has to be captured downstream.S46S43 | -455% | -711.7% | -439.5% | -387% | -314.1% | -241.4% | -178.1% | -123% |
Reported ≤ 2025 · Projected ≥ 2026
The chain in numbers
The five links side by side, in the scenario you picked above. This is where the chain's tensions become visible: what each link earns, what the whole thing costs, and who is supposed to pay for it.
Read the bars as compound annual growth rates: the steady yearly pace that carries each link's 2026 value to its 2030 value. The fastest-growing links are not the computing at the head of the chain. They are the physical resources downstream of it. HBM memory and electricity compound far above the roughly 14% a year of AI silicon, because every extra chip drags an outsized load of memory and power behind it. That is the tell of where the bottlenecks sit, and with them the pricing power: downstream of the silicon, not in it. Memory appears three times on purpose, because the three lines say different things. The total memory market compounds at 15.6%, but most of it is memory for phones, cars and ordinary servers, which grows slowly. The slice the AI fleet consumes compounds at 37.9%, and HBM, the stacked memory bolted onto the chips themselves, at 63.7%. Reading the 15.6% as 'AI memory grows slowly' is reading the wrong line.
- Accelerator revenue
- Units grow near 19% a year while the average price per chip slips about 2%. Revenue is what is left of the two pulling against each other.
- Total AI silicon
- Accelerators plus the networking that wires them together, so it tracks the accelerators closely. This is the base rate the rest of the chain is measured against.
- AI datacenter networking
- Priced off accelerator revenue at the ratio NVIDIA reports, drifting slightly below it as racks get denser and fewer cables serve more chips.
- Total memory market
- The whole memory industry, phones and cars included. Most of it has nothing to do with AI, which is why it compounds slowest on the chart.
- AI-attributable memory
- The slice of that market the AI fleet consumes. It grows with chips shipped and with the gigabytes each one carries, both rising at once.
- HBM market
- Stacked memory that only goes onto AI chips. Gigabytes per chip and price per gigabyte are climbing together, so the two compound on top of each other.
- Physical infrastructure capex
- Buildings, land and power delivery. It follows megawatts rather than chips, and megawatts grow faster because each new generation draws more.
- Annual electricity cost
- The bill for running the fleet. It tracks everything ever installed and still alive, which compounds far faster than any single year's purchases.
- Total AI revenue
- The fastest line here, and the least anchored. It compounds with tokens, which compound with the fleet, at a price the base case assumes falls 12% a year.
Three cash lines sit side by side. Total AI revenue. Chain spend, the cash actually out the door each year. And annual cost, that same spend spread across the years the equipment earns. Two crossings matter. Revenue overtakes cost in 2029 in the base case, as US$166bn against US$270bn in 2026 becomes US$1.75 trillion against US$1.30 trillion in 2030. That is the moment the arithmetic starts to pay. But revenue never catches spend inside this window. The chain lays out US$911bn in 2026 and US$2.57 trillion in 2030, and the cash bill runs ahead of revenue throughout. That distinction is the whole point. Spend is what the industry writes in cheques each year; cost is that outlay spread over time.
Stack the four spending links (silicon, AI memory, physical infrastructure and electricity) and the mix tells a story the headline totals hide. Base-case figures throughout. Silicon leads early, at roughly US$474bn of the US$911bn spent in 2026. But by 2030 physical infrastructure, meaning the buildings, land and power delivery around the chips, becomes the single largest line at about US$1.2 trillion of US$2.57 trillion. Underneath, the fastest-climbing slices are AI memory (US$125bn to US$453bn) and electricity (US$18bn to US$123bn), the bottleneck links from earlier in the chain. That memory line is the AI-attributable slice, compounding at 37.9%, not the 15.6% of the whole memory market in the chart above. The drift is plain. A rising share of every dollar goes not to the computing itself but to the memory and the megawatts it cannot run without.
Two spending lines sit here that look comparable and are not, which is exactly why they are plotted together. The solid line is the model's own AI-silicon demand across every buyer: GPU rental firms, governments, ordinary companies and China, not just the usual names. In the base case it runs from US$474bn in 2026 to about US$796bn in 2030. The dashed line is Wall Street's spending forecast for the big four hyperscalers alone, Microsoft, Alphabet, Amazon and Meta, at US$757bn rising to US$1.25 trillion. They are not meant to match. One counts silicon bought by the whole world; the other counts total spend by four companies. Read the distance between them as a measure of how much AI-silicon demand comes from outside the big four's budgets, not as an error in either series.
The nine variables that decide the outcome
Everything above resolves to a single headline; everything below is where that headline can break. This is the machine room, the part of the model most publications leave out. It holds four things, so you can audit the number instead of trusting it. The levers that move the answer, and by how much. The outside figures the model is checked against, none of which it used as inputs. The limits of what this exercise can and cannot claim. And the source behind every input that enters it. Nine of the 43 levers explain nearly all the variance between scenarios, and one of them, price per token, swings the answer further than any other input in the chain. I put the uncertainty on the table here, quantified, because a projection you cannot take apart is a projection you cannot check.
The model has 43 levers; nine are shown here. They explain nearly all the variance between scenarios. Each one carries what to verify and where to verify it — that's the work of every earnings season.
One assumption decides this model, and these two charts are built to prove it. Read each bar as a controlled experiment. The full chain is re-run with a single lever swung from its bear value to its bull value, while every other lever stays at base. The length of the bar is exactly what that one assumption is worth. The tick marks the base case. The left chart ranks the levers by how far they move 2030 revenue; the right, by how far they move coverage, the verdict on whether the infrastructure ever pays for itself. The ranking is the finding: price per token (B5) dominates both, and no physical lever comes close — not the chip count, not the fleet, not efficiency. What the world spends on AI turns less on how many tokens it runs than on what each one sells for.
B6 (adoption) moves neither bar, and that is a finding: this model derives revenue from the supply side — chips, fleet, tokens. Adoption is the plausibility check (see the adoption bridge above), not a driver.
Annual drift in the blended realized price
next year's price = prior price × (1 + B5)
bear −30%: a16z cost curve and the flight to cheap models on OpenRouter · bull 0%: Anthropic raised Sonnet 5 +50% and its tokenizer emits ~30% more tokens
Effect on 2030 revenue: US$700 billion ↔ US$2.92 trillion
The single highest-leverage variable in the model. A token getting 30% cheaper per year versus one that doesn't changes 2030 revenue by a factor of six. API list prices have fallen hard, but the mix has shifted toward pricier reasoning models — the two forces fight.
US enterprise AI adoption in 2030
2030 adoption = today's measured 19.8% + linear ramp to B6
official Census BTOS series: 19.8% in May-2026, flat for six months
Determines how many workers are on the other side to absorb the tokens. Base assumes 35% of firms use AI; bull assumes 55%. The official series (Census BTOS) is running in the high single digits — the widest gap between what the market discounts and what gets measured.
Annual unit growth — NVIDIA (2027-2030)
year's units = prior units × (1 + A1)
bull +24% is NVIDIA's own guidance · base +15% the houses' midpoint
Effect on 2030 revenue: US$1.36 trillion ↔ US$2.13 trillion
Drives all four layers at once: more units means more revenue, more memory, more gigawatts and more tokens. It's also the most verifiable number — it comes straight from quarterly filings.
Monetization factor (% of tokens billable via API)
billable = tokens produced × B3
measured at Google: only 26% of its tokens go through the API; blended with OpenAI/Anthropic (~100%) gives ~55%
Effect on 2030 revenue: US$1.59 trillion ↔ US$1.91 trillion
Nearly half the tokens that run are never billed: internal platform consumption, free tiers and test traffic. That cushion is why cost per billable token sits so far above price.
Each quarter: API vs. consumer mix in lab disclosures; OpenRouter and hyperscaler commentary on internal inference.
Blended realized price — 2026 anchor
2026 price = accounting revenue ÷ billed tokens
US$55bn accounting ÷ 39.6 quadrillion = US$1.40 · run-rate would give US$2.00 — 43% too high
Effect on 2030 revenue: US$1.44 trillion ↔ US$2.50 trillion
The most uncertain input in the entire model, and the one the industry most often calibrates wrong. It anchors against ACCOUNTING revenue, not run-rate: using run-rate inflates the price by over 40% and drags the error through to 2030.
Each quarter: recognized revenue (not ARR or run-rate) from OpenAI, Anthropic, Google Cloud AI and Azure AI, divided by estimated tokens.
Fleet utilization factor
TWh = GW × 8,760 × PUE × D2
Epoch AI measures 71% real utilization; assuming 100% inflates power demand by 41%
Effect on 2030 coverage: 1.33× ↔ 1.36×
Converts installed silicon into electricity consumed. Raise it 10 points and the energy gap widens proportionally — without buying a single extra chip. It's the bridge between layer 1 and layer 4.
Each quarter: hyperscaler commentary on utilization and capacity contracts; datacenter operator filings; utility interconnection data.
Total AI revenue multiplier over the model layer
total AI revenue = model-layer revenue × B8
rule of three on 2026: US$166bn total AI ÷ US$55bn model layer = 3.0
Effect on 2030 revenue: US$1.28 trillion ↔ US$2.33 trillion
For every dollar labs charge for tokens, how many dollars does the full ecosystem charge — cloud, software, applications? Base assumes 3x. It's what separates 'models don't make money' from 'the industry does'.
Silicon service life (years)
annual depreciation = what was bought over A7 years ÷ A7
Epoch AI uses 5 years; hyperscalers stretched stated life from 3 to 5-6 years
Effect on 2030 revenue: US$1.57 trillion ↔ US$1.77 trillion
Decides how much fleet stays alive and how much depreciates each year. Cutting it from 5 to 3 years raises annual cost and sinks the coverage ratio, without changing a single dollar of capex.
Each quarter: changes to stated useful life in hyperscaler depreciation notes (Microsoft, Alphabet, Amazon, Meta).
Chip ASP growth
year's price = prior price × (1 + A5)
half of 2026 growth was price, not volume: +31% price, +46% units
Effect on 2030 coverage: 1.29× ↔ 1.43×
The lever nobody models. Consensus projects units; silicon revenue is price × units, and the average price rose with every generation (H100→B200→GB300). If ASP stops climbing, the silicon ceiling drops even with unchanged units.
Implied average price in NVIDIA/AMD SEC filings: datacenter revenue ÷ estimated units.
What it was validated against
A model that can explain any outcome explains nothing. The only test that means anything is whether it can be proven wrong. So the chain is held against numbers it did not use to build itself: four figures published by other people, on their own methods, that it has to reproduce or fail. Each row below sets the model's own output beside that independent reference and shows the gap between them. Three of the four land within a few percent, which is the evidence that the derivation is sound rather than fitted to the answer. The fourth is a deliberate difference in scope, flagged as such, not a miss. Agreement here is what earns the right to believe the numbers above it that no one else has published yet.
Sum of SEC filings: NVIDIA Data Center + Broadcom AI semis + AMD Data Center + Marvell Data Center
China is excluded because those makers don't file with the SEC. A low single-digit deviation confirms the unit-and-price reconstruction is right.
Observed ACCOUNTING revenue of the labs (OpenAI, Anthropic, Google, Azure)
The comparison uses recognized revenue, not annualized run-rate. Using run-rate (~$80bn) inflates implied price per token by over 40% and drags the error through to 2030.
TrendForce: total $889.3bn (DRAM $618.7bn + NAND $270.6bn)
The model reaches memory from silicon (units × GB per package × price per Gb), never using the market figure. Matching a house that measures it from the supply side is an independent check.
Four-hyperscaler capex consensus (Goldman Sachs)
These are NOT expected to match, and the gap is informative: the model covers every buyer of AI silicon (neoclouds, sovereigns, enterprises, China), while the consensus covers only Microsoft, Alphabet, Amazon and Meta. The excess measures how much spending happens outside the big four.
How it gets updated
This dashboard does not update itself, and saying so matters. At the close of every earnings season a fixed ritual runs: read the quarterly filings, confront them against what the model assumed, adjust the levers that moved, and regenerate the data file. Every change is logged below.
Read what was reported
NVIDIA, Broadcom, AMD and Marvell publish datacenter revenue. That gives the quarter's actual unit growth, against what the model assumed.
Listen to the suppliers
SK hynix, Micron and Samsung for memory; TSMC for packaging capacity; utilities for interconnection. Suppliers announce the constraint before customers do.
Measure adoption
The Census Bureau publishes the BTOS survey every two weeks. It's the only public, official, frequent series on enterprise AI use. The base scenario needs it to reach 35% by 2030.
Check the price
Published list prices from the labs, changes to discounts and caching, and recognized revenue divided by estimated tokens. It's the most uncertain variable and the one that moves the outcome most.
Regenerate and log
Levers get adjusted, the engine runs, and the data file is republished under a new version. What changed and why goes in the changelog.
Methodology, limits and sources
- —It is not an investment recommendation or a price target on any instrument. It is an estimate of an industry's physical and monetary demand.
- —It does not model Chinese demand at the same quality, because those makers don't file with the SEC. Whenever it's validated against public data, China is excluded.
- —Realized price per token is the most uncertain input in the entire exercise. It's anchored to accounting revenue, not run-rate, and even so the bear-bull range is nearly two to one.
- —The 2030 projections assume the physical relationship between silicon, memory and energy holds. An architecture jump would break it in either direction.
Every model input references one of these identifiers. Without an identifier, the input doesn't enter the model. Every source pill goes somewhere: public ones open the original document, ones without a public link open their entry below.