Financial Forensic Analysis

AI Supercycle Forensic Monitor

Grounded in Jim Chanos' asset-liability and accounting-arbitrage framework. Monitoring the structural imbalances between front-loaded capital outlays and deferred downstream revenue across the 2025–2026 AI infrastructure build-out.

Every figure here is sourced and dated in the reference log →

What changed — August 2026

10–12 Aug 2026

Three developments in one week. Two of them did not add data to this monitor — they invalidated instruments.

  1. The neo-clouds grew fast and lost more money. CoreWeave revenue +112% with net loss widening to $626M on financing costR2; Nebius spent 9.7× its revenue on capexR6. Utilisation is fine. The deterioration arrived through the interest line, which is the wrong end of the income statement from where this monitor was lookingR9.
  2. A $500B plan to move chip purchases off balance sheet. Nvidia and six asset managers signed memoranda of understanding to finance compute through special-purpose vehicles, collateralised on the GPUs and funded by insurance and retirement capitalR11R13. Nvidia retains residual-value support of up to 25%R14. Compute bought this way would never enter reported capex or CIP — nothing has been drawn yetR72, but the structure puts this monitor's central balance-sheet instrument on notice: it reads a floor, not a measurementR18.
  3. The efficiency shock came from inside. Nvidia paid $20B to license Groq and shipped an inference part claiming ~10× energy efficiencyR46R47. The monitor assumed such a shock would arrive from outside and damage Nvidia. It arrived captured: Nvidia's revenue is insulated and the obsolescence risk sits with its customers and their creditorsR49. Nvidia is now guaranteeing the residual value of the chips its own new chip devaluesR50.
Net read. Demand is real and contracted. The deterioration is in financing. The measurement problem is now the central problem. Evidence log →

Hyperscaler Prints — Q2 2026

Capex, guidance, and whether operations still cover it.

As of 30 Jul 2026

← swipe to see all columns →

Company Q2 CapEx FY26 Guide Free Cash Flow The tell Tape
Alphabet GOOGL $44.9B $195–205B −$5.9B First negative FCF as a public company. Coverage 0.87×.R45 −7%
Microsoft MSFT $41.0B ~$175B ↓ headline only +$19.6B Two accounting levers at once: building lives 15→25yr, and ~$15B of leases reclassified out of headline capex with no change to spend.R31R37 +8%
Meta META $31.1B ×1.8 YoY $130–145B floor ↑ $784M Ad impressions decelerated to +14% (from 19%, 18%). Revenue carried by price, not volume.R41R42 −7%
Amazon AMZN $53.1B net PP&E $220B ↑ on memory −$7.6B TTM Buildout now partly debt-funded. Offsetting: AWS +36.7%, margin +650bps, backlog $496B.R31aR33R34 +8%

The tell this quarter is not the size of the spend — it is that operating cash flow stopped covering it. Two of the big four are free-cash-flow negative in the same quarter, and Microsoft's relief rally came from holding the envelope flat while extending useful livesR39. Aggregate 2026 guidance now sits near $725–800B, a 77% step-up on 2025R60.

Meta is the exception that qualifies the rest. At Amazon, Microsoft and Alphabet the financial leg deteriorated while demand accelerated. Meta is the first name where both moved together — and it has the least contracted revenue to fall back onR43.

An accounting story becomes a credit story at the point where the marginal dollar has to be borrowed. That point has passed: Amazon confirmed debt issuanceR33, the FOMC held on a 9–3 vote with three dissents for a hike, and the 30-year pushed past 5.19%R38.

Amazon's $53.1B is net property & equipment purchases; Microsoft's $41.0B includes finance leases. Indicative, not like-for-like. Sources & full figures →

Related-Party Disclosure

Circular Revenue — What The Filings Actually Say

The claim is that chip vendors and hyperscalers fund their own customers, so some reported revenue is the seller's own money returning. A register of 29 announced arrangements was built and every quotable figure traced to an SEC document or withdrawn. The exercise refuted more of the thesis than it confirmed — eight refuting findings against twelve confirming onesR30b.

Disclosed circular revenue, Microsoft→OpenAI
$24.1B
FY2026, compelled by ASC 850 R20
Peers disclosing the same
0 of 3
Amazon, Alphabet, Nvidia all outside ASC 850 R21
Measurable ceiling, after verification
$5.62B /qtr
Was $25.1B before unsourced figures withdrawn R24
Amazon mark-up on Anthropic, one quarter
$50.5B
Runs through earnings; no operating cause R30

The disclosure that was supposed to be impossible

The structures are built to stay non-voting and below the significant-influence threshold, so that ASC 850 related-party disclosure never fires. At Microsoft it fired. The October 2025 recapitalisation pushed its holding to an ~25% as-converted equity-method interest, which compelled the exact figure the structure was meant to keep private: $24.1B of FY2026 revenue from OpenAI, $6.0B receivableR20.

The three peers holding comparable positions disclose none of itR21. So the rule is neither that ASC 850 never fires nor that it does: the avoidance structure holds at three of four filers and fails only when a recapitalisation pushes a stake across the line. That makes Microsoft's $24.1B a natural experiment for what the other three are not required to reportR22.

The headline commitments mostly have no issuer source

Oracle's $300B OpenAI contract and AMD's ~$90B could not be traced to any filing by either party. A sweep of 212 documents — every Oracle and AMD accession since mid-2025, exhibits and XBRL included — returns zero hits for $300 billion and zero for Stargate. Both figures were carried here and have been withdrawnR23.

Rule

In this market, the larger the headline number, the less likely any issuer has ever said it.

The penny warrant, the margin test, and what this evidence cannot see

The $0.01 warrant is an industry instrument, not a one-off: AMD→OpenAI, AMD→Meta, Google→TeraWulf, Google→Cipher, CoreWeave→Core Scientific. A warrant struck at one cent is stock delivered on a conditionR25. Alphabet discloses $43.8B of credit derivatives, nearly tripled in six months from $16.9BR26. But Alphabet names no counterparty, so the linkage to those specific deals is inference, not disclosureR27.

The margin test came back clean — AWS margin expanded ~645bps while growth accelerated, Google Cloud margin roughly doubledR28. That is narrower evidence than it appears. The test sees concessions delivered as price and is blind to concessions delivered as equity — for the Nvidia, Amazon and Microsoft structures it has no power by construction, not merely no signal yetR29.

Method. Every figure is quoted verbatim or marked as inference. Disproven findings are retained with their retraction rather than deleted. The conclusions were put to an adversarial panel instructed to attack them; it returned kill, which is what prompted the 212-document sweep. The claims survived, one Oracle overstatement was corrected, and the margin-test scope was narrowedR30b.

Open: Microsoft's 15→25yr extension is not yet in a filingR44 · AMD's first warrant vestingR63 · watch for any other holder crossing into equity-method treatmentR64 Full register evidence →

Coverage Through Time

The forensic question is not how much the hyperscalers spend — it is whether the business still generates the cash to pay for it. This is trailing-twelve-month operating cash flow divided by trailing-twelve-month capital expenditure, straight from SEC filings, quarter by quarter. Below 1.0x, the buildout is being funded from the balance sheet rather than from operations.

Figure 4 — Operating cash flow as a multiple of capital expenditure, trailing twelve months. Every tracked builder has compressed toward the 1.0x line since 2022; Oracle and Amazon have crossed it.R75

Table view — latest filed quarter
Company Coverage, 2022 Coverage, latest TTM Op. Cash Flow TTM CapEx TTM Free Cash Flow Period end
Microsoft MSFT3.71x1.58x$182.9B$115.9B$67.0B2026-06-30
Alphabet GOOGL3.42x1.40x$185.7B$132.4B$53.3B2026-06-30
Amazon AMZN0.61x0.98x$148.5B$151.0B−$2.5B2026-03-31
Meta META3.00x1.64x$124.0B$75.7B$48.3B2026-03-31
Oracle ORCL2.12x0.57x$32.0B$55.7B−$23.7B2026-05-31

Source: SEC EDGAR XBRL company filings. Most quarters are derived by differencing year-to-date cumulative filings, since few issuers tag discrete quarterly cash-flow figures. Fiscal years are not aligned — Microsoft's ends in June, Oracle's in May, the rest in December — so read trajectory rather than like-for-like quarters. Trailing-twelve-month basis is used deliberately: single quarters are dominated by working-capital seasonality (Amazon's Q1 2022 operating cash flow was genuinely negative). Filings lag earnings by 22–30 days, so the final point may predate the most recent press release.

Early Warning Forensics Matrix

Current values against bubble-top thresholds. Hover any metric for what it means. Six instruments were amended in August 2026why.

Updated:

← swipe to see all columns →

Monitoring Metric Underlying Vulnerability Current Live Metric Thresholds Status
CIP to Net PP&E Ratio Deferred Depreciation Shield

Construction-in-Progress represents capital spent on data centers not yet active. Assets don't depreciate until "placed in service", letting companies park capital here to defer massive depreciation expenses.

Accounting shield via deferred depreciation. Reads a floor since Aug 2026 — SPE-financed compute never enters this ratio.R18 28.4%R70 floor · computed < 15% (Historical mean, below is acceptable) Impaired
Long-End Financing Cost Accounting Story → Credit Story

The build-out has crossed from internally-funded to externally-funded. Rising long rates lift the hurdle on every incremental capex dollar just as incremental returns compress, and directly squeeze the debt-service coverage of leveraged neo-clouds borrowing above 10%.

External funding meets a rising cost of capital. Promoted to primary read — price now leads issuance volume.R57 30Y 5.19% · 5Y CDS ~75bp 30Y sustained > 5.5%, or IG tech spreads +75bp from trough Elevated
Ad-Engine Volume Growth Spending More to Sell Less

For ad-funded capex (Meta, Alphabet), impression volume is the demand signal underneath the revenue line. Revenue held up by price per ad while volume decelerates is the classic late-cycle tell — the engine funding the build-out is weakening beneath a flattering headline.

Demand-side crack beneath a price-carried revenue print Meta +14% (from 19%) Volume growth −5pp over two quarters while capex guidance rises Triggered
Tech Buyback Volumes Capex Cannibalization Risk

When massive capital expenditure requirements eat into operating cash flow, companies must slow discretionary share buybacks to protect liquid reserves.

Free cash flow exhaustion under capex strain GOOGL $0
big tech −17% YoYR66
> 0% (Organic growth, above is expected) Breached
Operating Cash Flow CapEx Coverage The Self-Funding Test

Operating cash flow divided by capital expenditure. Above 1.0x the build-out is paid for out of the business; below 1.0x it must be funded from cash reserves, debt, or leases.

Build-out no longer self-funded — shift to debt & lease financing 0.87× (GOOGL) > 1.0× (Self-funded, below requires external capital) Breached
GPU Spot Leasing Prices Organic Demand Reality Check

While primary contracts mask real demand, real-time hourly secondary rental prices reflect true industry utilization.

Secondary capacity oversupply. Blended index invalid — prior gen collapsing while current gen is rationed; the read is the spread.R55 H100 −64–75%
Blackwell/Rubin rationed
Per generation. Widening current-vs-prior spread = obsolescence Spread widening
Grid Utility Lead Times The Power Lead-Time Paradox

High lead times strand capital: Completed facilities can't get power. Equipment sits idle in non-depreciating CIP, deteriorating ROIC.

Rapid decreases trigger crash: Easing bottlenecks allows massive backlogged compute online instantly, flooding the market and collapsing leasing margins.

Physical transmission bottlenecks & stranded asset risk 24–72 mo
5–7 yr where constrainedR69
< 30 mo = pricing pressure · > 30 mo = stranded capital Strained
Neo-Cloud Interest Burden Replaces the utilisation read

This monitor originally watched utilisation and lease rates for signs the leveraged intermediaries could not service their debt. Q2 2026 showed both healthy while losses widened anyway — the failure is arriving through interest expense, not demand.

Debt service outrunning rental incomeR9 Loss widening on 2× revenue Interest/revenue rising across quarters, or widening adj-EBITDA-to-GAAP gap, while revenue grows Triggered
Inference Efficiency Transfer Re-pointed, not re-scaled

The original instrument assumed an efficiency breakthrough would arrive from outside and damage Nvidia. It arrived from inside and was bought: Nvidia licensed Groq for $20B. Revenue is insulated; obsolescence risk transfers to owners of the prior generation and their creditors.

Obsolescence risk assigned to cohorts 2–3 and SPE creditorsR49 Realised $/token vs $45/M claim Falling realised price = deflation thesis · rising = Nvidia thesis Unresolved
Guided 2026 Big-Four AI CapEx
$725B – $800B
+77% on 2025's ~$410B · vs. ~$100B total Dot-Com telecom vendor financingR60
Hyperscalers Now FCF-Negative
2 of 4
Alphabet −$5.9B (first ever)R45 · Amazon trailing −$7.6BR31a
Neo-Cloud GPU Life Assumed
6yr CRWV / 4yr NBIS
Against a ~3-yr architecture cadenceR68 · a 3-yr schedule cuts EPS 6–15%R53
S&P 500 EPS, Last CapEx Digestion
−50%+
2000–2002 trailing operating EPS more than halvedR67

Macroeconomic Precedents

Historical Parallels

The current expansion mirrors the speculative patterns of the 1998–2000 telecommunications build-out and the 2005–2007 subprime credit expansion. While the Dot-Com era was driven by a demand myth regarding data traffic, the current cycle faces a unique reversal: hyperscalers are bearing the costs while their downstream customers remain heavily unprofitable.

Figure 1 — Comparative infrastructure spending. A single hyperscaler's 2026 guide is already 2× the entire 1998–2002 Dot-Com telecom build-out; the big four combined are roughly 8×.

Industry Cohort Risk Matrix

Value-Chain Exposure

Five distinct layers present unique balance sheet exposures — from CIP deferrals in hyperscalers to heavy leverage in neo-cloud intermediaries.

Hyperscalers

MSFT · GOOGL · AMZN · META

Primary RiskCIP Manipulation
VulnerabilityiROIC Decay

Neo-Cloud Intermediaries

CRWV · NBIS · Fluidstack

Primary RiskGPU-Collateral Debt
VulnerabilityInterest BurdenR9

Colocation Operators

EQIX · DLR · CORZ

Primary RiskAsset-Light Squeeze
VulnerabilityGrid Bottlenecks

Hardware & Power Enablers

GEV · Siemens · Bloom Energy

Primary RiskBacklog Concentration
VulnerabilityValuation Compression

Upstream Silicon & Suppliers

NVDA · TSMC · AMD · MU

Primary RiskCustomer Concentration
VulnerabilityDouble-Ordering Reversal

Forensic Accounting: The CIP Backlog

By early 2026, Construction in Progress (CIP) balances have reached unprecedented levels. Under GAAP rules, hardware classified as CIP is exempt from depreciation — creating a temporary shield for operating margins as obsolescence risk accumulates off the P&L.

  • Alphabet: $78.6B in capitalized assets not yet in service (55% YoY increase)R51 — and Q2 2026 capex of $44.9B, double the year-ago quarter, feeding the same backlog.
  • Meta: Extended server useful life to 5.5 years, deferring $2.9B in annual depreciation.R52 Q2 2026 capex $31.1B against $784M of free cash flow.
  • Microsoft: 6-year depreciation schedule added $3.7B to pre-tax incomeR52; on the Q2 2026 call the CFO confirmed further lengthening of building useful lives while Q2 capex ran $41.0B (+69% YoY).
  • Amazon: Q2 2026 net PP&E purchases of $53.1B — the largest of the four — against trailing free cash flow of −$7.6B. Jassy defended the asset lives directly on the call: data centers absorb capital ~2 years before earning and last 30+ years; servers break even in under three years against five-to-six-year lives, with AI capacity contracted at least five years. No updated CIP balance was disclosed this quarter, leaving the ~$29B figure stale.
  • The Q2 2026 pattern: spend accelerated, guidance moved up, and the depreciation recognised against it was pushed further out — CIP and useful-life extension are now doing the same work at the same time.

CIP balance by entity — the deferred depreciation wave.

Capital Efficiency

The iROIC Decay Monitor

Incremental Return on Invested Capital evaluates the profitability generated by each additional dollar of capital. A sustained decline toward 10% indicates that returns on new hardware may no longer cover the weighted average cost of capital.

Microsoft's cumulative incremental ROIC across the AI capex era (FY2022–FY2026) is 24.5%, against 51.2% for the pre-AI era (FY2019–FY2022)R61 — roughly a halving. This is a monitor computation on a consolidated basis from SEC XBRL filings, not a reported figure and not a segment figure: Microsoft discloses segment operating income but not segment invested capital, so no segment-level iROIC is checkable from disclosure. The pre-AI window starts at FY2019 to avoid an invested-capital tagging discontinuity in FY2017–FY2018.R70

Return Compression

The Air Gap Is Real — And It Is Not The Explanation

All four hyperscalers earn a lower return on invested capital than they did at the start of 2024. They are not failing businesses: operating income grew 36% to 79% over the same period. Their capital bases grew 67% to 160%. Computed quarterly from SEC filings; latest data 30 June 2026.

Pretax ROIC, annualised from the quarter.R74 Q4 is absent by construction — 10-Ks report full-year durations. Microsoft’s fiscal year ends in June.

Excluding capital that is not yet earning

Assets under construction sit in the denominator earning nothing, so a falling return could be timing rather than decline. Dashed lines remove that capital entirely. The gap between solid and dashed is the air gap; the slope is the answer.

Excluding it lifts the level by roughly 10–12 points — but does not flatten the decline. Alphabet’s −6.9pp becomes −7.5pp; Meta’s −14.6pp becomes −15.7pp. Both steeper without it.R74 Alphabet’s operating income rose 60% over this window and its return on capital already in service still fell. Meta reports “construction in progress”; Alphabet reports “assets not yet in service”. Microsoft discloses neither and Amazon annually only, so the test covers two of four.

The same quarter, from the other side of the balance sheet

Alphabet’s purchase commitments and other contractual obligations went from $332.4bn to $811.0bn in one quarter. The disclosure heading, categories and wording are identical across both filings, so this is growth rather than a change in what is disclosed. Short-term commitments grew 1.45×; long-dated commitments grew 3.14×, and 87% of the $478.6bn increase is long-dated. Over the same two quarter-ends, invested capital rose $166bn. Two independent measures, one company, one quarter, the same direction.R73

Read from the 10-Q text. This figure is not a tagged XBRL value — the concept API returns an unrelated $7.7bn.R73

What this does not show

  • It does not attribute the capital growth to AI. Invested capital includes everything.
  • It does not say returns are inadequate. Microsoft at 40% and Alphabet at 23% pretax remain high. The finding is the trajectory, not the level.
  • The air-gap test covers two of the four companies.
  • The relationship between capital restraint and return preservation rests on four data points.

What Frontier Capability Actually Costs

The buildout is justified by the premium that the most capable models command. So it is worth asking what that premium actually is. Each point below is the cheapest model available at or above its capability level — the efficient frontier of the market as it is priced today. Cost rises with capability throughout, and then accelerates sharply near the top: the steepest single step on the curve buys the last stretch of measured capability.

Closed — API access only Open weights

The capability/cost frontier, August 2026. Capability is the Epoch Capabilities Index; cost is list price per million tokens, blended 3:1 input:output, on a logarithmic axis. Open-weight models are marked separately.R76

Percentile of capability range Index Cheapest available Step
0% (floor)112.4$0.015
25%124.9$0.0281.9x
50%137.5$0.1003.5x
75%150.0$0.4504.5x
90%157.5$4.50010.0x
100% (ceiling)162.5$20.0004.4x

Read the right-hand column. Each quarter of the capability range costs more than the one before it — 1.9x, then 3.5x, then 4.5x — and then the curve turns: the 75th-to-90th percentile band alone costs 10x. In absolute terms, index 156.2 costs $0.45 per million tokens and index 161.5 costs $10.00: a 3.4% gain in measured capability costs twenty-two times as much.

The composition of the frontier splits just as sharply, and along the same line. Five of the six frontier models below index 156 are open-weight or Chinese. Above it, none of the five are. The cheap frontier is largely open; the expensive frontier is entirely closed and American. That is consistent with the separate finding that seven of the ten most-used models on OpenRouter are open-weightR56.

One of these prices has an expiry date printed on it. Gemini 3.7 Flash sits on the frontier at $1.50 blended, but Google's own pricing page states that rate holds only through 31 December 2026, rising to double on 1 January 2027R77. At the new rate it leaves the frontier entirely. A frontier partly composed of introductory pricing is not a stable frontier — and an aggregated price feed reports the current number with no expiry attached. This one was visible only by reading the vendor's page.

What this measures, and what it does not
  • The cost axis is price per token, not price per task. A verbose reasoning model emits more tokens to answer the same question, so per-token price understates its true cost per task — and understates it most at the top of the range, which is exactly where this chart is most interesting. The direction of that bias is known; its size is not.
  • List prices only. No batch discount, no cache discount, no negotiated enterprise rate. Large buyers do not pay these numbers.
  • Capability indices are constructions, not measurements. The index used here is a composite over one organisation's benchmark suite. A different suite would move the points; the question is whether it would move the shape.
  • 200 of 283 indexed models carry a published price. Every current-generation unpriced model in the band where it could displace a frontier point was priced by hand against its vendor's page, because a missing price silently drops a model out of the frontier — and dropping a cheap capable model would exaggerate the premium in the direction this page already argues. The remainder are superseded previews and duplicate host listings.
  • Two different kinds of price sit on this chart. Where a model's creator sells it, the creator's published list price is used. Where the creator only released weights and runs no API of its own — gpt-oss, Gemma, Mistral NeMo, Qwen2.5-Coder — there is no such price, so the cheapest third-party hosted rate is used instead. Hosted rates for a single model span up to 19x across providers; the minimum is taken deliberately, because the frontier asks what is cheapest available.
  • The capability index does not cover every lab. Tencent, among others, has no entry in it at all, so its models cannot appear here however cheap or capable they are. A benchmark suite that omits a vendor removes that vendor's points from the cheap end of the frontier, which flatters the premium rather than understating it.

Capability: Epoch AI, "AI Benchmarking Hub", published online at epoch.ai, licensed CC-BY. Cost: each vendor's own published pricing page, read 20 August 2026, with an MIT-licensed aggregated price map used as a cross-check. Where the two disagreed by more than 5% the vendor page was used — this happened once, on GPT-5 nano, where the aggregator was 9.1% high.

What A Task Actually Costs

Both charts above price capability per token, which is the number vendors publish and the wrong number for anyone buying work. This one uses measured spend: Cursor runs a benchmark of real, ambiguous, multi-file coding tasks and reports what each model actually cost to finish one. No blended rate, no assumptions — just the bill.

The cheapest configuration reaching each score, against measured cost per task. Each model appears several times — once per reasoning-effort setting.R80

The shape of the first chart survives on measured money. Going from 70.8% to 72.9% — just over two points of score — costs 6.4x, from $2.81 to $18.02 a task. That is the same conclusion the capability/cost frontier reached from list prices and an abstract index, arrived at independently with no pricing assumptions at all. Two constructions agreeing is worth more than either alone.

It also shows something list prices cannot. The effort dial is the cost dial. Claude Fable 5 costs $6.80 a task at medium effort and $18.02 at maximum — a 2.6x range for one model at one published price, which is more than switching vendor buys across most of the range. The panels above price a model. This one prices a decision about how hard to run it.

And it corrects something this page had been assuming. Work back from the measured bill and these tasks turn out to be enormous: a single task consumes at least a million and a half tokens, most of it context re-read across dozens of agent steps — 76 steps for the most expensive configuration.R81

Which means the per-token charts above are not understating the cost of this work — if anything they overstate the rate, because agentic tasks are dominated by input tokens and input is a fifth the price of output. The cost comes from volume, not from the rate. So there is no correction factor: a per-token price simply cannot tell you what a task costs, in either direction. That is why the charts above say what they measure rather than adjusting for it.

What is measured, what is inferred, and the limits
  • Measured: the score, the cost per task, the steps per task, and the token count Cursor reports. Cursor states cost is computed by applying each model's published pricing for input, cache read, cache write and output to the tokens it used.
  • Inferred: the million-token figure. Cursor does not define what its token column comprises; the arithmetic implies it is output only. Because cache reads bill at about a tenth of input, the reconstruction charges those tokens too much, so a million and a half is a floor, not an estimate — the true volume is higher. Both corrections push the same way, which is why the conclusion holds even though the field is undocumented.
  • Tested against a second domain, and it held. An earlier version of this note said nothing here generalises. That was too strong. ARC-AGI-2 is abstract visual reasoning — no repository, no tool loop, no resent context — and 32 model configurations appear on both benchmarks with the identical effort setting. For those, the price per token is the same number on both sides and cancels exactly, so the ratio of costs is the ratio of token volume, with no pricing assumption at all. An agentic coding task costs a single-digit multiple of an abstract-reasoning task — median 1.72x against ARC-AGI-2 and 4.20x against the easier ARC-AGI-1 — and on the harder benchmark 7 of the 32 are cheaper on the coding side. Million-token tasks are not peculiar to agent loops. Quote the range, not one figure: the estimate moves with the difficulty of whichever reasoning benchmark sits in the denominator, because an easier one burns fewer tokens. The order of magnitude replicates; the point estimate does not.R82
  • Two domains is not all domains. Both benchmarks tested here are long-generation work. Chat, summarisation, classification and retrieval remain untested, and a single-turn task should be smaller by orders of magnitude. The caveat has narrowed, not disappeared.
  • These are the benchmark's costs, not a buyer's. No enterprise discount, no prompt engineering to cut steps. A competent operator pays less.
  • Coverage is 67 configurations across four model families. A cheaper model outside those families would not appear, and the 6.4x step at the top is a jump between two specific models rather than a smooth market curve.

Source: CursorBench, cursor.com/cursorbench, and ARC-AGI-2, both via Epoch AI, "AI Benchmarking Hub", epoch.ai (licensed CC-BY). Read 21 August 2026. List prices used for the reconstruction are the same vendor pages as the panels above; the cross-domain test uses no price data.

What An Hour Of Autonomous Work Costs

The panel above prices capability against an index. This one prices it against something a buyer actually decides about: how long a task a model can carry out on its own. METR measures that directly — the task duration, in human time, at which a model succeeds half the time. Plotted against the same list prices, it moves the story. The expensive step is not at the top. It is the step into hour-long work.

The cheapest model able to sustain a task of each length, against blended list price per million tokens. Both axes logarithmic. Bars show METR's fitted interval, which is wide at the top.R78

Task lengthCheapest model that can sustain itStep
15 minutes$0.100
30 minutes$0.1001.0x
1 hour$1.92519.2x
2 hours$3.4381.8x
4 hours$4.5001.3x
8 hours$10.0002.2x

Crossing from half-hour to hour-long tasks costs 19.2x. Going from one hour to eight costs about five times in total. Once a model can sustain an hour of autonomous work, extending that horizon is comparatively cheap. Acquiring the first hour is what costs. That is the opposite shape to the index panel above, and both are true — they are different axes. But they answer different questions, and this is the one a buyer faces.

Now the part that matters more than the price curve. Everything above is measured at a 50% success rate — a coin flip. METR also publishes the horizon at 80%, which is nearer a bar anyone would actually deploy against. The horizons do not shrink gently. They collapse.R79

Model50% horizon80% horizon$/1M
Claude Opus 4.612.0h1.17h$10.000
Gemini 3.1 Pro6.4h1.50h$4.500
GPT-53.4h0.64h$3.438

The twelve-hour model is a seventy-minute model at a threshold you would rely on. And the ranking inverts: at 80%, the most expensive model on the chart drops off the frontier entirely, beaten on horizon by one costing less than half as much. The buildout is underwritten by an expectation of long autonomous capability. Measured at a bar anyone would deploy against, the best horizon in this data is about ninety minutes — and paying twice as much does not buy more of it.

What this measures, and three ways it is limited
  • The uncertainty is large, and largest where the chart is most interesting. Horizons are fitted curves, not stopwatch readings. Claude Opus 4.6's twelve hours carries a fitted range of 5.3 to 65.8 hours, which overlaps Gemini's substantially. The shape — a steep step at the bottom, a flat stretch above it — survives that. The top-end level does not, and nothing here should rest on it.
  • This is price per token, not price per task, and the mismatch is sharper here than on the panel above: an eight-hour agentic job emits vastly more tokens than a thirty-minute answer. The real spread across task lengths is therefore wider than shown. This is the price of the tokens, arranged by task length.
  • The newest models are absent. Claude Opus 5, Fable 5 and the GPT-5.6 family have no measured horizon at all — and they are the expensive ones, so this understates what the current top tier costs.
  • Three of the six models on this frontier are no longer sold. OpenAI has withdrawn o4-mini and GPT-5 from its price list; Moonshot has withdrawn K2 Thinking. Those figures are historical. This is structural, not sloppy: measuring long-horizon capability takes long enough that the models measured have since been retired. Read this as a snapshot of a market that has already moved.

Capability: METR, "Measuring AI Ability to Complete Long Tasks" (arXiv:2503.14499) and "Task-Completion Time Horizons of Frontier AI Models" (Time Horizon 1.1), metr.org/time-horizons. Horizons are the 50% success threshold except where the 80% figure is named. Cost: vendor list prices, read 20 August 2026, on the same basis as the panel above. The two models on this frontier that are still sold were checked against their vendors' pages and both matched exactly. METR measured the horizons. The frontier construction, the pricing and every conclusion drawn on this page are this monitor's own, and METR does not underwrite any of them. METR confirmed on 21 August 2026 that its public work may be cited on that basis.

The Short Seller's Comparative Playbook

The Fracking Trap: Like shale oil wells, high-end compute clusters suffer rapid technological decline curves. A GPU cluster purchased today might lose 80% of its competitive economic utility inside 36 months, yet it is capitalized with slow, optimistic depreciation schedules.

Subprime Securitization: Off-balance-sheet SPVs and neo-cloud leasing transfer hardware risk to private credit. As of August 2026 this is no longer an analogy — $500B of GPU-collateralised SPE financing was announced, funded by insurance and retirement capital.R19 Long-term construction leases are paired against volatile, cancelable user compute demands — creating a severe asset-liability mismatch.

Dot-Com Echo: Circular vendor financing is back. Semi suppliers invest heavily in customers who immediately recycle that cash into purchases of the supplier's advanced silicon chips, driving optical revenue growth.

Chanos Rule of Thumb
"Bull markets price dreams. Bear markets analyze structural balance-sheet decay."

Physical Decline Curves vs. Accounting Lifespans

Comparing the actual productivity drop of fracking wells to the economic obsolescence rate of GPUs vs. aggregate straight-line depreciation assumptions.

Red: True GPU Obsolescence Rapid loss of hardware competitiveness as faster silicon launches.
Yellow: Fracking Decline Curve Structural baseline for natural resource exhaustion profiles.
Green: Straight-Line Accounting Linear book-value depletion currently reported on corporate filings.

Depreciation Arbitrage Simulator

Explore how manipulating useful life assumptions distorts EBITDA and conceals rapid hardware obsolescence.

GPU Acquisition CapEx$10.0 Billion
Assumed Accounting Life6 Years
True Economic Useful Life3 Years
GAAP Reported Depreciation (Annual):$1.67B
True Economic Depreciation Required:$3.33B
Artificial Net Income Uplift:+$1.67 Billion

*By extending server useful lives beyond true technology utility, companies artificially pad gross operating margins.

The Circular Vendor Financing Loop

How semiconductor builders, neo-cloud operators, and credit syndicates recycle capital to inflate earnings and offload systemic equipment risks.

Phase 1: Equity Recycle

The Silicon Monopoly

Leading chipmakers purchase strategic equity stakes in Neo-Cloud structures or back key venture funds.

Phase 2: Collateral Debt

The GPU Borrowing Base

Neo-Clouds leverage this backing to raise massive asset-backed credit facilities from private lenders using GPUs as primary collateral.

Phase 3: The Order Rush

Buying Back Silicon

Borrowing capacity is immediately sent back to the supplier to lock in orders of advanced chips, showing record sales metrics.

Chanos Parallel
"This creates a recursive economic machine. Suppliers validate their own market growth projection via financial equity injections. If actual customer utilization turns down, this loop reverse-leverages violently as high depreciating inventory becomes stranded."

Interactive Fragility Diagnostic

Toggle real-time observed triggers to calculate the structural risk threat index.

Calculated Fragility Level: MODERATE

The Commitment Gap: Off-Balance-Sheet Obligations vs. Trailing Capex

Contracted forward obligations that sit outside the balance sheet — purchase and unconditional obligations, plus leases signed but not yet commenced — set against each company's trailing-twelve-month capex. The multiple is how many years of current spending is already contracted. Alphabet 6.8×, Meta 7.0×, Microsoft 4.8×, Amazon 1.5×; $2.35tn against $511bn of TTM capex across the four, or 4.6× in aggregate.R73 Read from filing text, not XBRL: these figures are not tagged and structured-data queries return unrelated, far smaller values. On-balance-sheet lease liabilities are excluded for every company so the four are on one basis. Microsoft's $329.1bn is disclosed directly in its FY2026 lease note as leases “primarily for datacenters” that have not yet commenced. As of Q2 2026 for Alphabet, Meta and Amazon; FY2026 (30 June) for Microsoft.

Strategic Implications

The transition from speculative expansion to normalization is inevitable. Institutional portfolios should pivot toward industrial "bottleneck beneficiaries" like the turbine oligopoly — GE Vernova now filling 2028–29 slots with some customers pulled into 2030R62 — while reducing exposure to leveraged intermediaries dependent on continuous venture injections.

Final Conclusion

"When the supply of new issuance exceeds institutional demand, the market becomes highly vulnerable to a sharp valuation adjustment."R56 As of August 2026 the harder problem is that CIP and reported capex have both been engineered downward — the monitoring has to follow the risk into the credit markets.R18