The Three-Layer War: Who Captures the AI Stack

Every player in AI is integrating vertically toward the application layer. But if everyone is doing it, it can't be differentiating. A structural analysis of who captures the economics of the AI stack — and what would prove the thesis wrong.

The Three-Layer War: Who Captures the AI Stack

It started with chips. Then it was models. Now it is applications.

Over the past eighteen months nearly every serious player in AI has made the same move: reach into the layer above or below their home turf. Hyperscalers pushed custom AI silicon into production at a scale that made Nvidia's margins impossible to defend. Nvidia, almost on cue, began shipping open models. Model providers — Anthropic, OpenAI — launched coding agents, design tools, and browser products that compete directly with their own largest customers. The application companies that built on those APIs suddenly discovered they were feeding the companies that would later try to replace them.

This is not chaos. It is a predictable restructuring of a maturing value chain. The strategic question that matters for investors is not whose model tops a benchmark this month. It is: who captures the economics of the AI stack, and why?

The answer requires mapping the stack into three layers with very different economics, and then tracking why every player is integrating vertically in the same direction — toward the application layer, where durable value capture may live. One caveat up front: "the application layer" here means the application relationship and the compute and distribution behind it. Pure-play application companies, as we will see, are among the most exposed. Whether the integrated stack is actually the winning pattern — or whether, as in prior technology stack transitions, modular specialists eventually tear it apart — is the question this analysis engages but does not fully resolve.

The Map: Three Layers, Three Economics

The AI stack, as it exists in mid-2026, has three layers. They differ in capital intensity, in defensibility, and — most importantly — in where the margin ultimately accrues.

LayerWhat it isCapital intensityMoatMargin today
ComputeChips (Nvidia, custom silicon), foundries (TSMC), data center infrastructureExtremeScale + tooling + supply contractsHigh (Nvidia, TSMC)
ModelsFoundation models — frontier (GPT-5.6, Claude Opus 4.7, Gemini) and open (Llama, DeepSeek, Qwen, Kimi)Very high, recurringScale, data access, talent — but erodingNegative to thin
ApplicationsCoding agents, enterprise workflows, consumer superapps, vertical toolsModerateData flywheels, workflow lock-in, distributionEmerging, uneven

Two assets cut across all three layers and are not located in any single one: data (proprietary training data, user interaction data, domain-specific workflow data — the moat the application layer generates and the model layer needs) and distribution (owned consumer surfaces, enterprise sales channels, regulated procurement relationships — the factor that determines whether a layer's moat is actually monetizable). The post treats these as cross-cutting rather than as a fourth and fifth layer, because they are inputs to every layer rather than stages in a chain. There is also a binding physical constraint underneath all of this — power. The ~$725B of 2026 hyperscaler capex is increasingly gated not by chip availability but by electricity: data center power demand is projected to grow faster than grid capacity can expand in many regions, making energy access a quiet but decisive competitive factor. Power is not a layer in the AI stack, but it may be the constraint that determines how fast the compute layer can actually grow.

The most important fact in this table is the middle layer's margin profile. Foundation models, as a standalone business, are not yet profitable and may never be highly profitable on their own. OpenAI did roughly $24B in annualized run-rate revenue by Q1 2026 against a cost base — training, inference, talent, and a hyperscaler tax to Microsoft — widely reported to exceed revenue. Anthropic, valued at $965B in its May 2026 Series H (having surpassed OpenAI's $852B — the market pricing in, rather than proving, the application-revenue thesis), was at $9B ARR at the end of 2025; bull-case analyst estimates project $93B by May 2027, with Claude Code contributing ~$21B of that. Those are extraordinary growth assumptions. They are also assumptions that the model layer does not commoditize faster than the application layer can be built out.

The reason everyone is integrating vertically is that the table above is unstable. Compute has strong moats — but they are contestable, and they are under attack from the compute layer's own largest customers. Models have weak moats and are commoditizing in real time. Applications have the moats that are most defensible — proprietary data, embedded workflows, customer relationships — but standalone application companies are vulnerable to the very model providers that power them. Each layer is trying to capture the layer where value will ultimately pool.

Layer One: Compute and the Silicon Rebellion

The compute layer is, for now, where the money is. Nvidia's position through 2025 was the closest thing to a monopoly rent the technology industry has seen since Intel in the 1990s. The hyperscalers paid it, grudgingly, because they had no choice.

They now have a choice, and they are exercising it. The custom silicon push of 2025–2026 is the single most important structural event in the compute layer:

  • Google has the most mature program. Trillium (TPU v6) is in production, Ironwood (TPU v7) followed, and a five-year, $46B Broadcom contract to build TPUs through 2031 was disclosed in an April 2026 SEC filing. Google has been designing its own AI silicon for a decade; this is no longer experimental.
  • Amazon went from late to aggressive. Trainium2 is in production, Trainium3 has been announced, and Project Rainier — a cluster of 500,000+ Trainium2 chips dedicated to training Anthropic's models — is among the largest non-Nvidia AI compute deployments in the world. AWS is reportedly in active discussions to sell Trainium externally.
  • Microsoft announced Maia 200 on January 26, 2026, an inference-focused part not sold externally.
  • Meta confirmed in July 2026 that its first purpose-built training chip, "Iris" (designed with Broadcom), enters production in September 2026, with a planned cadence of a new chip every six months through 2027.

The strategic logic is not that these chips outperform Nvidia's. They generally trail on peak performance per chip, though the gap narrows at the system level and depends heavily on workload. The logic is that the Nvidia tax — the gross margin Nvidia extracts on every accelerated-compute dollar — is large enough to justify tens of billions in custom silicon capex. At hyperscaler scale, even a chip materially less efficient than Nvidia's — say, in the 70–80% range of peak performance — but at a fraction of the cost can be a decisive economic win, because the alternative is sending margin to a single supplier who is also, increasingly, a strategic competitor. (These ratios are illustrative, not measured; the point is directional.)

Nvidia's response is the most revealing move in the entire stack. At GTC in March 2026 it announced Vera Rubin, the platform widely understood to succeed Blackwell. But the more interesting announcement came at Computex in June: Nemotron 3, the latest in its open model family, with a 550B-parameter MoE Ultra variant and a 30B Nano variant with one-million-token context. Nvidia, a hardware company, is now a notable open-weights publisher.

The strategic logic here is not, as a first reading might suggest, about locking inference to CUDA — open weights are portable and will run on AMD, TPU, or Trainium within weeks of release. The coherent logic is commoditize-the-complement. Nvidia's long-term threat is not that inference moves to rival chips; it is that a model provider accumulates enough scale and pricing power to become a monopsony buyer of compute, dictating terms back up the stack. By flooding the model layer with free, good-enough open weights, Nvidia ensures no single model provider can corner the layer above and turn that corner into buyer power over chips. This actually strengthens the model-squeeze thesis developed below: Nvidia wants the model layer fragmented, because a fragmented layer cannot compress Nvidia's margins. The model layer, for Nvidia, is not a business. It is a defensive move to keep its customers many and small.

The takeaway for investors: the compute layer will remain profitable, but the supernormal margins of 2024–2025 are not structural. They are being competed away by the very customers who created them. Nvidia stays highly profitable; it does not stay a monopoly.

Layer Two: Models, Commoditization, and the Open Source Problem

The model layer is where the economics are hardest to defend. Consider what a frontier model requires in 2026: multi-billion-dollar training runs, clustered compute measured in tens of thousands of accelerators, a talent war for senior researchers, and an inference cost base that must be subsidized to win developers. Now consider what protects that investment: almost nothing durable. Architectures leak. Researchers move. Training recipes diffuse. TechCrunch argued in July 2026 that "the real AI race may no longer be at the frontier" — meaning the frontier itself may no longer be where the competition is decided.

The result is visible in pricing. Through 2025 and into 2026, token prices across the frontier have fallen repeatedly, and Milkroad and others have described an active "pricing collapse." But it is worth pausing on what falling prices actually prove, because three different stories produce the same price chart and they have opposite implications for who captures value:

  1. Genuine commoditization — undifferentiated substitutes drive price toward marginal cost. This is the reading that supports the "model layer gets squeezed" thesis.
  2. Cost-curve pricing with stable or expanding margins — AWS cut prices over a hundred times between 2006 and 2020 while its margins grew, because each price cut reflected efficiency gains the incumbents captured. Falling prices can be evidence of a widening scale moat, not its absence.
  3. Deliberate loss-leader share-buying — subsidized pricing to capture developers, which implies players expect future pricing power once share is locked in. This is the opposite of commoditization; it is a bet against it.

The honest position is that the current price chart cannot distinguish between these three, and the model layer's long-term economics depend on which one is actually operating. If it is story 1, the thesis below holds and model providers become utilities. If it is story 2, the frontier oligopoly earns AWS-style rents. If it is story 3, today's losses are a land grab that converts to pricing power. The rest of this analysis proceeds on the assumption that stories 2 and 3 are partially operating but that story 1 is the dominant force — open-weights pressure from below (DeepSeek, Qwen, Kimi) makes it structurally hard for any single frontier provider to hold pricing power even if they wanted to. But this is a judgment call, and the reader should hold the alternative outcomes in mind.

The open source question — who pays, and is there a future — has now been answered, directionally, by Meta. On April 8, 2026 Meta released Llama 5 as open weights (~600B parameters, five-million-token context). On the same day it launched Muse Spark as a closed, proprietary model. Press coverage widely read the same-day split as Meta effectively abandoning open weights as its frontier strategy. (CNBC reported on July 9 that an open Muse Spark variant is "in development." Meta's pivot is also over-determined as evidence: going closed is consistent with wanting to monetize a hit model directly, competitive-leakage concerns, and regulatory positioning — not only with "open can't win the frontier." But the strategic emphasis has plainly shifted.)

The investor-relevant question about open source is not whether open weights are a standalone business — nobody claims they are. The question is whether open weights destroy the pricing power of the proprietary frontier, and on that question the answer is yes, for the same reason Android destroyed the OS-licensing business model despite never being a standalone business itself. Android was funded by Google's advertising motives and given away free; it nonetheless captured 70%+ of global mobile OS share and eliminated the ability of any rival to charge for an operating system. The sustainability of the funder is irrelevant to the competitive effect — what matters is whether the funder's strategic motive persists. China's geostrategic motive (~80% of US AI startups reportedly use Chinese open-source models, per the US-China Economic and Security Review Commission), Nvidia's complement-commoditization motive (see the Nemotron analysis above), and DeepSeek's hedge-fund AGI bet are all persistent. As long as they are, they ship frontier-adjacent capability for free within months of the frontier, and no proprietary provider can hold pricing power on workloads where the open tier is "good enough."

This is the key distinction: open source may never produce the single best model, but it does not need to. It needs only to be close enough, soon enough, to cap what the frontier can charge. DeepSeek V4 (April 2026, ~1T-parameter MoE, fully open, funded by High-Flyer's 57% hedge-fund returns) and Kimi K3 (July 2026, 2.8T params, open, funded by Moonshot AI's venture capital) are both frontier-adjacent. They are not sustainable standalone businesses — they are strategic bets funded by non-model revenue — but that distinction matters to the funders' accountants, not to the frontier's pricing power. An MIT Sloan analysis estimated optimal closed-to-open reallocation could save the global ecosystem ~$25B/year; that is a welfare argument about efficiency, not a business model, but it confirms the direction of travel.

The conclusion is not that "open source loses" — it is that open source wins nothing for itself but denies everyone else the win. The open-weights tier will own the "good enough" market: the long tail of use cases that do not require the absolute frontier, enterprise self-hosting for sensitive data, and academic research. The frontier itself, where the capital intensity is highest and the talent is most concentrated, will remain proprietary — but it will remain proprietary at compressed margins, because the open tier sets the ceiling on what the frontier can charge. Meta's pivot is the clearest evidence that open weights are not a frontier strategy (Meta could not win the frontier that way); but open weights are nonetheless a frontier pricing strategy's worst enemy, whether or not Meta participates.

This has a second-order implication that matters for the application layer. If the frontier is proprietary and the open tier is a commodity, then the only way model providers earn a durable return is by moving up the stack into applications — where data, workflow, and distribution create real moats. Which is exactly what they are doing.

Layer Three: Applications and the Trojan Horse

A note on naming before we proceed: in February 2026 SpaceX acquired xAI at a $250B valuation and folded it in as SpaceXAI. So when this section refers to "xAI's parent SpaceX" or "SpaceX/SpaceXAI," it is the same post-acquisition entity — distribution via X, models from the former xAI, compute from its own clusters.

If you want to understand why Anthropic launched Claude Design, why OpenAI is reportedly merging ChatGPT, Codex, and its Atlas browser into a single desktop superapp ahead of IPO, and why SpaceX acquired Cursor for $60B in all-stock in June 2026, the answer is the same: the application layer is where value capture consolidates, and model providers cannot afford to let someone else own it.

The logic is simple. A model accessed via API is a commodity input. The company that builds the application on top of that API owns the customer relationship, the usage data, the workflow embedding, and ultimately the pricing power. Over time, the application company captures the margin and the model provider captures the inference cost. Every model provider in 2026 has looked at this dynamic and drawn the same conclusion: we must own the application.

This is the "trojan horse" problem, and it is the most discussed strategic risk in the application layer today. The mechanism is intuitive: when a startup builds on the OpenAI or Anthropic API, the model provider sees the use cases that work, sees the prompts that convert, sees the data that flows through the workflow — and then, if the market is large enough, enters it. Vanderbilt's "AI Neutrality" paper in January 2026 formalized this vertically-integrated bundling problem; Business Insider reported in June on the broad pattern of AI companies "rapidly expanding into each other's markets." For enterprise buyers it is now a procurement concern: the question a CIO asks is not "which model is best" but "if I standardize on this provider, how long until they compete with me?"

It is worth being precise, though, about what has actually been demonstrated versus what is still a feared mechanism. The trojan-horse story — model provider watches your usage, learns your market, then enters and undercuts you — has not yet played out cleanly in a major 2026 case. The closest thing to a canonical application-layer event, Cursor, actually demonstrates a different mechanism:

In December 2025, Cursor's CEO publicly argued that competition would not crush his company. Six months later — four days after SpaceX's own IPO at $135/share on June 12 gave it fresh acquisition currency — Cursor was acquired by SpaceX for $60B in all-stock, the largest venture-backed startup acquisition on record. The framing — xAI losing the AI-coding race and buying the winner — is one plausible reading of the timeline. But note what did not happen: Cursor was not captured by the model providers it built on (it routed across Anthropic and OpenAI — it was, by the post's own definition, model-neutral). It was bought by a third party with a fresh balance sheet. That is evidence for application-layer consolidation at enormous premiums, and evidence that application-layer value is real and monetizable — Cursor turned roughly $400M of raised capital into $60B in four years while renting its models. It is not clean evidence for the replicate-and-undercut trojan horse. The honest read: the application layer is where so much value is being created that every entity with a balance sheet is forced to buy its way in. That is closer to a bull case for application builders than a warning.

So the trojan-horse risk remains a risk — widely feared, procurement-tested, not yet fully executed in a headline case. What is already demonstrably true is the broader pattern: model providers are integrating upward into applications because they have to. Anthropic's moves are the clearest expression of this. Claude Code — widely characterized in early-2026 coverage as having its "ChatGPT moment" — is projected (in the same bull-case estimates cited above, for May 2027) to contribute ~$21B to Anthropic's ARR. Against a $9B end-2025 actual, that would make Claude Code alone larger than most standalone software companies — if the projection holds. Claude Design launched in April 2026, received a major overhaul on June 17, and added two-way integration with Claude Code on June 18. Anthropic filed a confidential S-1 on June 1, 2026 — suggesting the application-revenue thesis will be central to its public-market story. Anthropic is not building a model and hoping developers use it. Anthropic is building the applications that capture the workflow, because that is where the revenue and the data live.

The implication for the application layer is a structural shift in who is positioned to stay independent. Application companies that depend on a single proprietary model provider are the most visibly exposed. The companies that are better positioned to remain independent have one or more of three things:

  1. Model neutrality — the ability to route across multiple providers or self-host, so that no single model provider can hold them hostage or learn from their data. (Cursor had this and still sold — but it sold at a record premium, which is a different outcome than being undercut.)
  2. Proprietary data that does not flow through the model provider — domain-specific data, customer data that cannot be exposed, regulatory moats (healthcare, finance, defense).
  3. Distribution that the model provider cannot replicate — embedded workflows, enterprise sales relationships, regulated procurement channels.

The Vertical Integration Logic, by Origin Layer

Every player in the stack is integrating vertically, but the direction and the motivation depend on where they start. The pattern is clarifying:

Origin layerMoving towardWhy
Compute (Nvidia)Up into models (Nemotron)To commoditize the complement and prevent model-provider monopsony
Hyperscalers (Google, Amazon, Microsoft, Meta)Down into silicon, up into appsTo escape the Nvidia tax and own the customer relationship
Model providers (OpenAI, Anthropic)Up into applicationsBecause raw models face pricing pressure; apps are where margin lives
Integrated distributors (xAI/SpaceX)Into apps via acquisition (Cursor)Because distribution (X) + models + application = full stack

The direction that is common to all of them is movement toward the application layer. This is the predictable consequence of the economics: applications are where value capture consolidates, because that is where the proprietary data, the workflow lock-in, and the customer relationship all live.

But here is the problem with declaring vertical integration "the winning pattern": if everyone is doing it, it cannot be differentiating. The table above shows four categories of players all integrating toward the same endpoint. Universal strategies produce competed-away returns, not moats. Five fully-integrated AI stacks fighting for the same enterprise and consumer customers looks less like Apple's vertical domination of mobile and more like the airline industry — capital-intensive, undifferentiated at the margin, and chronically struggling to earn its cost of capital. The post's own Google analysis (below) makes this concrete: Google already holds all three layers and is not yet dominant. Integration is necessary to compete; it does not appear sufficient to win.

There is also a deeper historical argument against integration as the endgame, and the post needs to engage it rather than ignore it. Christensen's interdependence/modularity cycle describes how technology stacks evolve: integration wins while the product is not good enough (cross-layer optimization produces compounding advantages), but the moment the product becomes good enough, modular specialists tear the integrated stack apart (Wintel over IBM's mainframe stack; best-of-breed SaaS over Oracle's integrated suite; AWS's modular layers over integrated hosting). The post's own open-source section argues the "good enough" tier is large and growing — and by modularity logic, that growing tier should be won by modular specialists (routing layers, neutral inference, best-of-breed apps), not by integrated stacks. The condition under which integration wins is narrow: the frontier must stay performance-constrained, so that cross-layer optimization — model-on-own-silicon, application-tuned-to-model — keeps producing advantages that modular specialists cannot match. The condition under which it loses is the open-weights tier converging close enough to the frontier that "good enough" becomes the dominant segment and modular specialists take it. Which condition we are in depends on whether the frontier pulls away from or converges with the open tier over the next 18–24 months.

The capital being deployed is staggering regardless. Combined hyperscaler capex (Google, Amazon, Microsoft, Meta) is on pace for roughly $725B in 2026, and CNBC projects it will top $1 trillion in 2027. This is the financial signature of a land grab: each player is spending to secure a position in the layer where they do not yet have a moat, on the assumption that the window to establish one is closing. Whether that land grab produces Apple-style winners or airline-style capital traps is the question the rest of this analysis tries to answer — and, as the counter-arguments section will acknowledge, does not fully resolve.

Who Captures the Pie

With the map laid out, the question of long-term value capture has a reasonably clear answer, though it is uncomfortable for anyone hoping a single layer will dominate.

Short term (through 2026–2027): compute wins — unless the squeeze is priced in early. The capex flowing into AI has to be spent on something, and that something is chips, data centers, power, and cooling. Nvidia, TSMC, Broadcom, and the equipment names capture this spend directly. The hyperscalers' custom silicon programs take years to erode Nvidia's position, and in the interim Nvidia's revenue and margins remain extraordinary. The pick-and-shovel thesis holds here — with one caveat developed in the counter-arguments section: compute revenue is model-layer capex, and if capital markets reprice the model layer's economics before the capex cycle completes, the compute trade can unwind faster than the timeline suggests.

Medium term (2027–2029): the model layer gets squeezed — or consolidates into a choke point. This is the layer with the weakest moat and the most aggressive competition. Token prices will continue to fall. The frontier will remain a small oligopoly — OpenAI, Anthropic, Google, plus Meta's Muse Spark if it reaches the frontier, and whatever the SpaceX/SpaceXAI entity becomes — but the economic returns of being in that oligopoly depend on which of the pricing stories above is operating. If the open-weights tier caps frontier pricing (the base case), model providers that do not build successful application businesses will find their economics resemble a utility — high volume, thin margin, capital-intensive. If instead the frontier pulls away and training costs compound toward a TSMC-style supply concentration, the oligopoly earns rents instead. (Nvidia's Nemotron, which this post treats as a compute-layer moat rather than a model-layer business, sits outside this oligopoly by design.) Either way, model providers that do build application businesses (OpenAI via its superapp, Anthropic via Claude Code and Design) will look more like software companies than model companies — which is why every model provider is integrating upward.

Long term (2029+): value pools at the application relationship, gated by compute access and distribution. Note that this conclusion is framed in three ingredients (application relationship, compute, distribution) rather than the three layers of the opening map. Distribution is the cross-cutting factor that determines whether a layer's moat is actually monetizable: a great application with no distribution pays a toll to whoever owns the channel. With that framing, the durable winners will be the entities that hold all three ingredients simultaneously: a customer relationship and proprietary data flywheel (applications), the compute capacity to train and serve at scale (compute), and the distribution to reach users without paying a toll to a rival (distribution).

The strongest test of this thesis is Google. Google already holds all three: TPUs and the $46B Broadcom contract (compute), Gemini (models), and Search, Workspace, and Cloud (distribution and application relationships). Under the framework here, Google should be the presumptive long-term winner. That it is not yet the clear frontrunner — Gemini is competitive but not dominant, and Google's application-layer AI products have not pulled away — is itself informative. Full vertical integration is necessary but not sufficient; execution, organizational speed, and product quality still decide who converts an integrated position into market leadership. The framework predicts who is positioned to win, not who will.

Run against the same test: the hyperscalers have compute and distribution by default and are buying application capability. OpenAI has distribution (ChatGPT) and is building applications on top of its own models, but does not own its compute — a structural dependency on Microsoft that cuts both ways. Anthropic has the strongest application positioning among pure model providers but is compute-dependent on Amazon and Google — a structural vulnerability. SpaceX/SpaceXAI has distribution (X), its own compute footprint, and just bought a category-defining application (Cursor). Meta has distribution (its apps) and now compute (Iris) but is starting late on applications and just shifted away from the open-weights strategy that might have made it the default model layer.

The most exposed category is pure-play application companies built on a single proprietary model — exposed as independent entities, that is, not necessarily as investments (see the Cursor distinction above). They are, structurally, R&D outsourced to the model layer.

Implications

Three conclusions that follow from the analysis, each of which an investor or operator in this space should internalize:

1. The model layer's economics are the decisive uncertainty. The investment case for OpenAI ($852B) and Anthropic ($965B) depends entirely on which of the pricing-decomposition stories above is actually operating. If the model layer commoditizes (story 1), frontier model providers become utilities and must escape into applications to justify their valuations — Claude Code and the ChatGPT superapp become the entire bet. If the model layer consolidates into a 2–3 player choke point with compounding entry costs (the TSMC outcome), the model layer itself earns oligopoly rents and the valuations are defensible on model revenue alone. The current evidence — falling token prices but enterprises paying premiums for the best coding models — does not cleanly distinguish these outcomes. This is the single most important question for anyone underwriting AI at current valuations, and it is not yet settled.

2. Open source denies the frontier its pricing power, even if it never wins the frontier. Meta's pivot is the clearest signal that open weights are not a frontier strategy, but the investor-relevant point is different: open weights are a frontier pricing strategy's worst enemy. DeepSeek V4 and Kimi K3 ship frontier-adjacent capability for free within months of the frontier, funded by non-commercial motives (Chinese geostrategy, Nvidia's complement-commoditization, venture subsidization). Whether those funders are commercially sustainable is beside the point — as long as their strategic motives persist, proprietary frontier providers cannot hold pricing power on workloads where the open tier is "good enough." This is why the model-layer squeeze thesis holds even if no open lab ever produces the single best model. Open source wins nothing for itself, but it denies everyone else the win.

3. The application layer is where the next wave of value is being created — and where independence is hardest to keep. Every model provider is now an application competitor, and every entity with a balance sheet is buying its way in (Cursor at $60B is the proof of application value, not of application fragility). Companies that solve model neutrality (multi-model routing, open-weights enterprise deployments, self-hosting) are solving a problem the market has only just begun to price. The honest framing: application-layer builders are creating enormous value and will often be acquired at premiums for it; whether they can remain independent category leaders depends on whether they hold model neutrality, proprietary data, or distribution the model providers cannot replicate.

The throughline across all three: in a maturing stack, value migrates toward the layer with the strongest moat and away from the layer with the weakest — if the stack stays integrated. Compute has moats but is under attack from its own customers. Models have weak moats and face open-weights pricing pressure that will persist regardless of any single player's strategy. Applications have the moats that are most defensible — proprietary data, workflow lock-in, distribution — but whether those moats accrue to integrated stacks or to modular specialists depends on whether the frontier stays performance-constrained (integration wins) or "good enough" becomes the dominant tier (modularity wins). The entire industry is reorganizing itself around the assumption that integration is the answer. Whether that assumption holds is the bet.

What would prove this wrong

A thesis worth publishing is worth stress-testing. Three counter-arguments a sophisticated reader will raise, and the conditions under which each would invalidate the framework above:

The modularity counter-thesis. As argued in the vertical-integration section above, every prior technology stack eventually went horizontal rather than vertical: Wintel over IBM, best-of-breed SaaS over Oracle, AWS over integrated hosting. The framework would be wrong if frontier capability stops being the binding constraint on most real workloads — at which point the "good enough" tier expands, modular specialists (routing, neutral inference, best-of-breed apps) take it, and vertical integration confines the integrated players to a shrinking premium segment. The condition to watch: whether the frontier pulls away from or converges with the open-weights tier over the next 18–24 months.

Reflexivity between the phases. The short/medium/long timing thesis treats the layers as sequencing on a schedule. They do not. Compute revenue is model-layer capex — the $725B of 2026 hyperscaler spend is justified by expected returns in the model and application layers. Capital markets are forward-looking. The moment the medium-term model-squeeze becomes consensus, the short-term compute trade reprices immediately, not on a polite three-phase timeline. This is how the 2000 telecom equipment bust unfolded: Cisco and Nortel's revenue was a derivative of the capital-raising capacity of the layer above them, and when that layer's economics were revealed as unsustainable, equipment demand collapsed inside two quarters rather than unwinding over years. The timing thesis would be wrong if model-layer commoditization is priced in before the compute capex cycle completes — in which case the compute "win" and the model "squeeze" arrive simultaneously, not sequentially. The decoupling condition: inference demand becomes driven by application-layer revenue (end customers paying for delivered AI work) rather than by model-layer fundraising. We are not there yet.

The TSMC outcome. The thesis assumes the model layer earns utility returns because it commoditizes. But capital-intensive layers with 2–3 players usually earn rents, not utility returns — TSMC being the canonical case. In the 1990s, semiconductor manufacturing was the consensus "commoditizing layer" with value supposedly moving to design; instead, capital intensity concentrated supply until one player had pricing power over the entire industry. If frontier training costs keep compounding — $500M for GPT-4, reported billions for Grok 4, and rising — the model layer could consolidate to 2–3 players who become the choke point, and value migrates toward them rather than away. The thesis would be wrong if enterprises treat frontier models as interchangeable on price alone, with no willingness to pay a premium for the best model on high-value workloads. Current evidence (coding agents, agentic workflows) suggests the opposite — enterprises pay premiums for capability — which is the empirical question that determines whether the model layer goes the way of utilities or the way of TSMC.

The model is not the business. The model is the cost of being in the business — provided the model layer commoditizes (story 1 above). If it does not, the model layer is the business, and everyone else pays rent to it. Which of those worlds we are entering is the decisive uncertainty this analysis cannot resolve — and the one that matters most.

Subscribe to Agentic Capital

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe