Essay · Artificial Intelligence
The Economic Structure of the Generative AI Ecosystem
Capital Intensity, Capability Diffusion, and Competitive Differentiation Across the Value Chain
Keywords

Abstract
Digital platform ecosystems built on generative AI are not single markets but six-layer value chains—spanning semiconductors, cloud infrastructure, foundation model development, inference services, platform orchestration, and applications—in which capital requirements, entry barriers, and value capture mechanisms differ sharply across layers. This paper analyzes the economic structure of this networked business ecosystem through a systematic literature review of 36 studies, evaluates six structural hypotheses, and advances three formal propositions extending platform theory and complementary-asset frameworks to multi-layer AI ecosystems. The findings reveal structural asymmetry: capital concentration coexists with rapid capability diffusion at the model layer and materially lower entry barriers at the application layer. Platform orchestration layers show the strongest theoretical basis for value capture as model capabilities commoditize. Firms spanning multiple layers internalize inter-layer value flows, explaining why vertical integration confers durable competitive advantages that technical competence alone cannot replicate.
Keywords: Generative AI, digital platform ecosystems, value chain, platform theory, AI infrastructure, vertical integration, capability diffusion.
1. Introduction
Generative artificial intelligence has evolved into a multi-layered commercial ecosystem in which the logic of competition differs sharply across layers. Firms competing in semiconductor hardware face a different cost structure, concentration dynamic, and value capture mechanism than firms competing in foundation model development, inference services, or downstream applications. Yet existing research treats this ecosystem as an undifferentiated category. Studies of AI scaling laws, platform dynamics, productivity effects, and governance risks address individual components of the stack in isolation; none integrates them into a systematic account of economic structure. As a result, firms lack an analytical framework for deciding where to compete, how to position across layers, or why vertical integration strategies produce different outcomes than single-layer focus.
The structural properties that make this ecosystem analytically distinctive can be traced to the emergence of foundation models—large neural systems trained on broad datasets and adapted to multiple downstream tasks (Bommasani et al., 2021). Three technological developments shaped these properties. First, the introduction of the transformer architecture (Vaswani et al., 2017) and the subsequent demonstration of emergent few-shot capabilities (Brown et al., 2020) established a new paradigm of general-purpose language systems. Second, scaling law research showed that frontier model performance is a predictable function of sustained compute investment (Kaplan et al., 2020; Hoffmann et al., 2022), making capital intensity a structural feature of the ecosystem rather than a transitional condition. Third, the deployment of these models through standardized APIs (OpenAI, 2023; Anthropic, 2023) created a layered supply chain in which downstream developers can access general-purpose intelligence without replicating the underlying infrastructure (Cusumano et al., 2025). Together, these developments produced an ecosystem characterized by capital concentration at the infrastructure layer, rapid capability diffusion at the model layer, and asymmetric entry conditions across layers—precisely the structure this paper analyzes.
This ecosystem is now embedded within broader economic and governance dynamics. Researchers have documented risks related to bias, misinformation, and misuse (Bender et al., 2021), proposed taxonomies of language model risks (Weidinger et al., 2022), and assessed transparency practices among frontier developers (Bommasani et al., 2023). Economists have examined how AI reduces prediction costs (Agrawal et al., 2018), shapes labor markets (Acemoglu and Restrepo, 2019), and generates productivity effects in knowledge work (Noy and Zhang, 2023; Brynjolfsson et al., 2023). Platform scholars have shown that technological infrastructures tend to evolve into multi-sided coordination mechanisms (Gawer, 2014; Hagiu and Wright, 2015; Evans and Schmalensee, 2016; Cusumano et al., 2019). Empirical work on LLM markets further shows that competition in these ecosystems is already shaped by pricing strategies, infrastructure constraints, and routing intermediaries (Demirer et al., 2025). This study integrates these strands into a unified economic analysis of the generative AI ecosystem.
Despite rapid progress across these areas, existing literature remains fragmented across disciplinary boundaries. Machine learning research primarily focuses on algorithmic capabilities and model scaling, while economic studies emphasize productivity and labor market effects. Platform research analyzes ecosystem dynamics but rarely integrates the technological constraints associated with AI development. From an information systems perspective, the generative AI value chain represents a new class of networked business infrastructure whose economic structure determines how digital platform competition unfolds across layers—yet no integrated IS-grounded account of this structure exists. From the perspective of industrial organization, the growing role of AI may also reshape market structures, particularly in industries characterized by strong complementarities between software platforms and cloud infrastructure (Varian, 2018).
Crucially, this fragmentation extends to the literature on technological ecosystems and innovation systems that is most directly relevant to the paper's argument. Research on sectoral patterns of technical change has long shown that industries differ systematically in their regimes of appropriability, entry barriers, and the relationship between capital intensity and value capture (Pavitt, 1984). The generative AI ecosystem exhibits an extreme version of this pattern—with one of the sharpest capital-intensity gradients ever documented in a nascent technology sector—yet no study has applied this analytical tradition to its structure. Similarly, the theory of complementary assets (Teece, 1986) predicts that the ability to profit from innovation depends not on the innovation itself but on who controls the complementary resources required to deliver value; this prediction has direct implications for which layers of the AI value chain will capture rents as model capabilities diffuse, but has not been applied to this context. Most recently, Jacobides, Cennamo, and Gawer's theory of ecosystems (Jacobides et al., 2018) develops a framework for understanding how firms with different architectural positions capture value through control of bottlenecks and interfaces—a framework that maps onto the generative AI value chain in ways this paper makes explicit. As a result, the broader economic structure of the generative AI ecosystem has not been systematically analyzed from within the innovation studies tradition that best equips scholars to address it.
This study addresses this gap by integrating insights from machine learning, digital platform economics, and innovation studies. It makes three contributions to the literature. First, it proposes a six-layer conceptual model of the generative AI value chain that organizes the ecosystem according to technological function, capital intensity, competitive dynamics, and value capture mechanisms—a unified structure that prior work has not systematically developed. Second, it provides a theory-grounded empirical evaluation of six structural hypotheses about the ecosystem, showing which claims are supported, which are partially supported, and which remain to be verified as the market matures. Third, it extends platform theory and the innovation-studies tradition on complementary assets and ecosystem architecture to a multi-layer generative AI context through three formal propositions: Proposition 1 (diffusion-driven platform shift) establishes that platform power migrates from the capability layer to the orchestration layer as model capabilities commoditize—an extension of the Teece (1986) complementary-assets framework to learning-system ecosystems; Proposition 2 (multi-layer platform simultaneity) identifies a structural configuration requiring layer-specific market definitions absent from existing ecosystem models (Jacobides et al., 2018); and Proposition 3 (integration premium) reframes the rationale for vertical integration from cost reduction to rent internalization across inter-layer bottlenecks, advancing the architectural-control logic of Jacobides et al. (2018) to the specific configuration of AI value chains.
Building upon a systematic literature review presented in Section 2, the paper develops the conceptual model in Section 3 and evaluates the following six hypotheses, each derived from the theoretical and empirical literature reviewed:
- H1 Capital intensity is concentrated in the compute infrastructure and frontier model training layers.
- H2 Compute supply is highly concentrated, though the concentration is not absolute or irreversible.
- H3 Capability improvements in foundation models diffuse rapidly across the ecosystem, compressing prices and eroding technical exclusivity.
- H4 Entry barriers in the application layer are materially lower than in compute and model layers, though platform dependence remains high.
- H5 Platform layers are positioned to capture a disproportionate share of ecosystem value through coordination and routing mechanisms.
- H6 Generative AI produces measurable and heterogeneous productivity gains in knowledge work, and the distribution of these gains across firms and sectors is conditioned by the supply-side structure of the value chain—particularly by competitive conditions in the inference and application layers.
Section 4 evaluates these hypotheses against the corpus. Section 5 discusses implications for platform theory, innovation policy, and competitive structure. Section 6 concludes with a summary of findings, theoretical contributions, and implications for research and practice.
2. Systematic Literature Review
This study adopts a systematic literature review (SLR) to identify and organize the literature relevant to the technological and economic structure of the generative AI ecosystem. The objective of the review is to integrate insights from machine learning research, digital platform economics, and innovation studies to support the conceptual framework developed in this paper.
2.1. Review Design
The review was designed to capture interdisciplinary research addressing generative AI from both technological and economic perspectives. Because the phenomenon investigated in this study spans multiple research domains—including computer science, information systems, industrial organization, and innovation economics—the review intentionally integrates literature from these fields.
The final corpus includes 36 publications released between 2014 and 2025. Of these, 20 are peer-reviewed works (11 journal articles, 4 conference papers, and 5 academic books); the remaining 16 are grey literature comprising working papers from recognized research institutions (NBER, Harvard Business School), arXiv preprints, and technical reports from AI research organizations and industry. This distinction is maintained throughout the analysis: empirical claims derived from grey literature sources are treated as indicative and qualified accordingly, while conclusions regarding structural patterns draw on the convergence of evidence across source types. The inclusion of grey literature is methodologically justified by the pace of development in this field, where peer-reviewed publication lags the phenomena under study by 18–36 months; it is disclosed transparently to enable readers to assess the evidential basis of each finding independently.
This temporal window was selected because it captures both the consolidation of research on platform ecosystems and the emergence of large-scale neural architectures and foundation models that underpin modern generative AI systems. Database searches were conducted between January and March 2025. All screening and selection decisions were made by the author; limitations arising from single-reviewer assessment are acknowledged in Section 5.4.
2.2. Data Sources
The search process relied on multiple academic sources to capture the interdisciplinary nature of the field.
The primary academic databases used for the search were:
- ACM Digital Library
- IEEE Xplore
- ScienceDirect
These databases were selected because they contain a substantial portion of peer-reviewed research in computer science, artificial intelligence, information systems, and software engineering.
To complement these sources, additional publications were identified through cross-database indexing and targeted searches in major scientific journals, conference proceedings, institutional research repositories, and technical reports from established AI research centers and industry organizations. Complementary venues included *Science*, *Nature*, *Journal of Economic Perspectives*, *Research Policy*, *Strategic Management Journal*, the NBER Working Paper Series, Harvard Business School Working Papers, and technical reports from recognized AI research institutions. The inclusion of *Research Policy* and *Strategic Management Journal* as complementary venues was deliberate: the study's theoretical framing draws on innovation-studies traditions (sectoral patterns of technical change, complementary assets, ecosystem architecture) that are primarily published in these journals, and excluding them would systematically under-represent the theory base most relevant to the paper's contributions.
2.3. Search Strategy
The search strategy combined keywords associated with five analytical dimensions relevant to this study:
- generative AI and large language models
- AI infrastructure and computational scaling
- platform ecosystems and digital markets
- economic and productivity implications of artificial intelligence
- innovation economics, sectoral patterns of technical change, and complementary assets
Search strings combined terms such as “large language models”, “foundation models”, “AI scaling laws”, “AI infrastructure”, “platform ecosystems”, “multi-sided platforms”, “AI productivity”, and “AI economics”. An example composite string applied across title, abstract, and keyword fields was: (*“generative AI” OR “large language models” OR “foundation models”*) AND (*“platform ecosystems” OR “AI infrastructure” OR “AI economics” OR “AI productivity”*). Searches were applied to titles, abstracts, and keywords. Only English-language publications were considered.
2.4. Inclusion Criteria
Publications were included in the review if they satisfied the following criteria:
- The study addressed at least one of the analytical dimensions relevant to this research: AI foundations and architectures, scaling and infrastructure, digital platform ecosystems, or economic impacts of AI adoption.
- The publication presented conceptual, empirical, or technical contributions directly relevant to understanding the structure or dynamics of the generative AI ecosystem.
- The publication appeared in a recognized academic or institutional venue, including peer-reviewed journals, conference proceedings, scholarly books, working paper series, institutional research reports, or technical reports with identifiable authorship and bibliographic metadata. This criterion explicitly covers working papers from recognized research institutions, such as Harvard Business School and the National Bureau of Economic Research, as well as technical reports from established industry organizations whose work directly informs the empirical analysis of AI infrastructure and market dynamics.
- The publication fell within the temporal window of 2014–2025.
The inclusion criteria were intentionally broad to capture the interdisciplinary nature of generative AI research, which spans computer science, economics, and technology management.
2.5. Exclusion Criteria
The following exclusion criteria were applied during the screening process:
- Publications focusing exclusively on narrow algorithmic improvements without broader implications for infrastructure, ecosystems, or market dynamics.
- Non-academic sources such as blog posts, press articles, promotional materials, or opinion pieces lacking methodological transparency.
- Publications lacking identifiable authorship, publication venue, or bibliographic consistency.
- Duplicated records retrieved across multiple databases.
Grey literature sources—working papers, arXiv preprints, institutional technical reports, and industry annual reports—were not excluded on grounds of publication type, but were subject to a higher threshold of relevance and were flagged separately in the corpus (see Section 2.1). Empirical claims derived exclusively from grey literature are treated as indicative throughout the analysis and are explicitly qualified in the text. This approach follows established practice in fast-moving technology domains where peer-reviewed publication typically lags the phenomena under study by 18–36 months.
2.6. Selection Process
The search initially returned 146 records across the selected databases and complementary sources. After removing 32 duplicated entries retrieved from multiple databases, the remaining 114 records were screened based on titles and abstracts to assess their relevance to the analytical scope of the review.
Of those 114 records, 50 were excluded based on title and abstract screening, leaving 64 publications retained for full-text evaluation. Full-text screening was subsequently conducted to assess alignment with the inclusion and exclusion criteria defined in the review protocol. Publications focusing exclusively on narrow algorithmic improvements, domain-specific applications without ecosystem implications, or lacking methodological transparency were excluded at this stage.
This process resulted in a final corpus of 36 studies (31 identified through the primary PRISMA screening process, supplemented by 5 additional studies identified through targeted citation searching of key innovation-studies sources), which form the analytical basis of the systematic literature review presented in this paper. Figure 1 summarizes the study selection process.

*Figure 1. PRISMA diagram of the study selection process.*
2.7. Corpus Organization
The selected studies were organized into seven thematic categories reflecting the main dimensions of the generative AI ecosystem: AI foundations, large language models and architectures, infrastructure and scaling, AI risks and governance, platform ecosystems and digital markets, innovation studies and ecosystem theory, and economic and productivity impacts.
Table 1 summarizes the 36 studies included in the review.
*Table 1. Studies included in the systematic literature review (36 studies: 20 peer-reviewed, 16 grey/technical literature).*
| Category | Studies |
|---|---|
| AI Foundations | (LeCun et al., 2015), (Goodfellow et al., 2016), (Jordan and Mitchell, 2015) |
| Large Language Models and Architecture | (Vaswani et al., 2017), (Brown et al., 2020), (Kaplan et al., 2020), (Hoffmann et al., 2022), (Bommasani et al., 2021), (OpenAI, 2023), (Anthropic, 2023) |
| AI Infrastructure and Scaling | (Patterson et al., 2022), (NVIDIA Corporation, 2024) |
| AI Risks and Governance | (Bender et al., 2021), (Weidinger et al., 2022), (Bommasani et al., 2023) |
| Platform Ecosystems and Digital Markets | (Gawer, 2014), (Hagiu and Wright, 2015), (Evans and Schmalensee, 2016), (Cusumano et al., 2019), (Gawer and Cusumano, 2002), (Cusumano et al., 2025), (Cusumano, 2025), (Demirer et al., 2025) |
| Innovation Studies and Ecosystem Theory | (Dosi, 1982), (Pavitt, 1984), (Teece, 1986), (Jacobides et al., 2018), (Cockburn et al., 2018) |
| AI Economics and Productivity | (Agrawal et al., 2018), (Acemoglu and Restrepo, 2019), (Brynjolfsson et al., 2017), (Noy and Zhang, 2023), (Peng et al., 2023), (Brynjolfsson et al., 2023), (Dell'Acqua et al., 2023), (Varian, 2018) |
3. Conceptual Model of the Generative AI Value Chain
The systematic literature review presented in Section 2 indicates that generative AI is best understood not as a single product category, but as a layered technological and economic system. Existing studies address individual parts of this system, such as large-scale training (Kaplan et al., 2020; Hoffmann et al., 2022; Patterson et al., 2022), foundation models (Bommasani et al., 2021; OpenAI, 2023; Anthropic, 2023), platform dynamics (Gawer, 2014; Hagiu and Wright, 2015; Evans and Schmalensee, 2016; Cusumano et al., 2019), and the economic effects of AI adoption (Agrawal et al., 2018; Cockburn et al., 2018; Brynjolfsson et al., 2023). Empirical evidence on AI infrastructure deployment and market dynamics further illustrates the capital intensity and competitive structure of the industry (NVIDIA Corporation, 2024; Cusumano, 2025; Demirer et al., 2025). However, these dimensions are rarely integrated into a single model of the generative AI value chain.
To address this gap, this paper proposes a conceptual model of the generative AI value chain. The model organizes the ecosystem into a set of interdependent layers that differ in technological function, capital intensity, competitive dynamics, and mechanisms of value capture. This layered perspective provides the analytical basis for evaluating the six hypotheses introduced in Section 1.
3.1. Rationale for a Layered Model
The value-chain perspective is appropriate because generative AI systems depend on a sequence of complementary technological and organizational activities. Progress at one layer often depends on capabilities and constraints located in another. For example, improvements in model performance depend on computational scale (Kaplan et al., 2020; Hoffmann et al., 2022); the availability of models through APIs affects application entry barriers (Bommasani et al., 2021); and value capture in digital ecosystems often depends on platform control rather than on the ownership of a single technical component (Gawer, 2014; Hagiu and Wright, 2015).
The generative AI ecosystem is therefore best understood as a structured stack rather than an undifferentiated market. The lower layers provide essential computational and infrastructural inputs. Intermediate layers transform those inputs into general-purpose model capabilities. Upper layers coordinate access, integration, and downstream applications.
3.2. Layers of the Generative AI Value Chain
The proposed model is composed of six layers:
- Semiconductors and specialized hardware
- Cloud and data-center infrastructure
- Foundation model development
- Inference and model serving
- Platforms and developer ecosystems
- Applications and domain-specific solutions
Although these layers are analytically distinct, they are economically interdependent.
3.3. Semiconductors and Specialized Hardware
The first layer consists of the semiconductor technologies required to execute AI workloads at scale. This includes GPUs and other specialized accelerators used to train and serve large neural networks. This layer is analytically central: research demonstrates that model performance improves with access to compute (Kaplan et al., 2020; Hoffmann et al., 2022), and empirical evidence documents the scale of hardware infrastructure required to support frontier AI workloads (NVIDIA Corporation, 2024). Without specialized hardware, the training and deployment of frontier models would be infeasible.
This layer is characterized by high fixed costs, specialized manufacturing requirements, and strong technological concentration. These properties make it one of the most capital-intensive segments of the generative AI ecosystem.
3.4. Cloud and Data-Center Infrastructure
Above the hardware layer lies the cloud and data-center infrastructure layer. This segment includes hyperscale computing clusters, storage systems, network interconnects, orchestration tools, and energy-intensive data-center operations. Research on the computational and energy costs of large neural network training demonstrates that frontier workloads require substantial resources (Patterson et al., 2022; NVIDIA Corporation, 2024).
This layer transforms hardware resources into usable computing environments for model training and inference. It is also where scale economies become especially pronounced, since the efficiency of large distributed clusters depends on integrated management of compute, storage, and networking resources (Varian, 2018; NVIDIA Corporation, 2024).
3.5. Foundation Model Development
The third layer consists of organizations that develop foundation models. These models are trained on large and heterogeneous datasets and are designed to support many downstream tasks (Bommasani et al., 2021). Recent systems such as GPT-4 and Claude illustrate the role of this layer as the producer of general-purpose generative capabilities (OpenAI, 2023; Anthropic, 2023).
This layer is strategically central because it defines the core capabilities made available to downstream users. At the same time, it remains closely tied to the lower layers, since model development depends on computational scale, data availability, and training efficiency.
3.6. Inference and Model Serving
The fourth layer concerns the operational delivery of model capabilities through inference systems. This includes the deployment, optimization, and pricing of model access through APIs and serving frameworks. The economics of this layer differ from those of training. Training requires large up-front investment, whereas inference converts model capabilities into ongoing operational services.
Recent market evidence shows that competition in this layer is shaped by pricing, model availability, and service differentiation (Demirer et al., 2025). As more providers offer comparable capabilities, inference is becoming a principal locus of price competition and commercial standardization.
3.7. Platforms and Developer Ecosystems
The fifth layer consists of platforms that connect model capabilities to developers, organizations, and users. These platforms include API gateways, orchestration layers, developer tools, integration frameworks, and distribution interfaces. Research on digital platforms shows that such coordination layers often become central to value creation and value capture in technological ecosystems (Gawer, 2014; Hagiu and Wright, 2015; Evans and Schmalensee, 2016; Cusumano et al., 2019).
Recent work suggests that generative AI systems are evolving toward this kind of platform structure, where model capabilities support broader ecosystems of applications and services (Cusumano et al., 2025). This layer gains strategic significance because it can reduce switching frictions, standardize access, and shape the terms under which complementors participate in the ecosystem.
3.8. Applications and Domain-Specific Solutions
The sixth layer consists of applications that embed generative AI capabilities into specific workflows, products, and use cases. Examples include coding assistants, enterprise copilots, customer service tools, productivity applications, and domain-specific systems.
Compared with lower layers, this segment generally exhibits lower fixed-capital barriers to entry because developers can access AI capabilities through platforms and APIs rather than training models independently. However, the application layer is also likely to be highly competitive, since many firms can build on the same underlying foundation models and infrastructure.
The long-term sustainability of firms at this layer may therefore depend on workflow integration, domain expertise, distribution channels, and user retention rather than exclusive technical capability.
3.9. Structural Characteristics Across Layers
The layered model implies that different segments of the generative AI value chain are likely to exhibit different economic properties.
First, capital intensity is concentrated in the lower layers. Hardware and cloud infrastructure require large fixed investments and specialized capabilities. Foundation model development also remains expensive because it depends on sustained access to large-scale compute.
Second, capability diffusion is most visible in the model and inference layers (Layers 3–4). Once model architectures, training methods, and alignment approaches become more widely known, competing organizations may replicate or approximate similar capabilities (Cusumano, 2025; Demirer et al., 2025).
Third, entry barriers in the upper layers are lower, but coordination and value capture may shift toward platform intermediaries. Platform research establishes that the layers controlling ecosystem participation, interfaces, and interactions are positioned to capture a disproportionate share of economic value (Gawer, 2014; Hagiu and Wright, 2015; Cusumano et al., 2025). At the same time, once general-purpose capabilities become accessible through APIs, application developers can enter at lower capital cost, making coordination and distribution the primary basis of competitive differentiation.
These three structural properties motivate the empirical evaluation conducted in the next section.
3.10. Implications for Hypothesis Evaluation
The conceptual model developed here provides the basis for interpreting the six hypotheses introduced in Section 1.
The first two hypotheses concern the lower layers of the stack, where compute infrastructure and specialized hardware shape capital intensity and concentration. The third hypothesis concerns the model and inference layers (Layers 3–4), where capability diffusion has been empirically documented. The fourth and fifth hypotheses concern the upper layers, where application entry barriers and platform coordination influence value capture. The sixth hypothesis connects the entire value chain to downstream productivity outcomes.
The model therefore clarifies that generative AI is an ecosystem in which technological layers are economically differentiated. This is the central analytical claim that the next section evaluates using empirical evidence from the literature.
4. Empirical Evaluation of Hypotheses
4.1. Analytical Strategy
This section evaluates the six hypotheses introduced in Section 1 using the systematic corpus defined in Section 2 and the value-chain framework presented in Section 3. The evidence base is heterogeneous. Some studies quantify scaling behavior and compute requirements (Kaplan et al., 2020; Hoffmann et al., 2022; NVIDIA Corporation, 2024). Others document pricing, supply, and demand in LLM API markets (Demirer et al., 2025). A further set measures productivity effects of generative AI in applied settings (Noy and Zhang, 2023; Peng et al., 2023; Brynjolfsson et al., 2023; Dell'Acqua et al., 2023). The relevant empirical question is not whether a single study proves each hypothesis, but whether the combined evidence points in a consistent direction.
Two cautions are necessary. First, the generative AI market remains at an early stage of development. Many observed relationships are structural tendencies rather than settled equilibria. Second, direct profit and loss data by layer of the generative AI value chain are scarce. The analysis therefore relies on observable proxies: compute scale, pricing trajectories, market concentration, entry patterns, usage concentration, and productivity outcomes. Where evidence permits, the discussion integrates both quantitative findings and qualitative structural attributes to assess each hypothesis.
4.2. H1 — Capital intensity is concentrated in compute infrastructure and frontier model training
H1 states that capital intensity is concentrated in the infrastructure and frontier model layers. The theoretical basis for this hypothesis derives from scaling law research, which shows that frontier model performance is a function of sustained compute investment (Kaplan et al., 2020; Hoffmann et al., 2022), and from industrial organization analysis, which identifies high fixed costs and specialized inputs as structural sources of entry barriers (Varian, 2018).
The convergent evidence from scaling research and infrastructure data supports H1. Kaplan et al. show that performance improves with additional compute, model size, and data, while Hoffmann et al. show that competitive training is compute-optimal only under a careful balance between parameters and tokens (Kaplan et al., 2020; Hoffmann et al., 2022). These are not abstract technical results: economically, they mean that frontier competition requires sustained investment in training capacity, not isolated experimentation.
NVIDIA's infrastructure data give the scale of that requirement. According to the company's own reporting, the data-center segment accounted for approximately 78% of NVIDIA's total revenue in fiscal year 2024, and third-party estimates suggest that training a frontier model such as GPT-4 required on the order of tens of thousands of GPUs and hundreds of millions of dollars in compute cost—though these figures are not officially disclosed and should be interpreted as indicative rather than precise (NVIDIA Corporation, 2024). The energy and carbon footprint of large-scale training provide a further proxy for resource intensity: Patterson et al. document the computational costs of earlier-generation large neural networks, establishing a methodological baseline for estimating the order of magnitude of frontier training requirements (Patterson et al., 2022). Upstream, the concentration of advanced semiconductor manufacturing further constrains access to frontier compute capacity (NVIDIA Corporation, 2024).
Cusumano's analysis of DeepSeek refines rather than weakens this result. It reports DeepSeek's own claim of about US$5.6 million using 2,000 H800 GPUs, but also cites outside estimates of access to as many as 50,000 GPUs, more than US$500 million in GPU hardware, and about US$1.3 billion in broader development cost (Cusumano, 2025). The methodological implication is straightforward: estimates vary according to whether one prices a single training run or the capital base required to make repeated frontier runs possible.
H1 is therefore supported.
4.3. H2 — Compute supply is concentrated, but the concentration is not absolute
H2 states that compute supply is highly concentrated. The hypothesis is grounded in industrial organization theory, which predicts that markets with extremely high fixed costs and specialized inputs tend toward concentration (Varian, 2018), while also recognizing that contestability and countervailing buyer power can moderate that tendency.
The evidence supports a concentration story, but not an irreversible one. According to industry analyst estimates cited in NVIDIA's own investor communications, the company held approximately 98% of the data-center GPU market in 2023 (NVIDIA Corporation, 2024). This figure is self-reported and has not been independently verified in peer-reviewed literature; it should be interpreted as an indicative order of magnitude rather than a precisely audited market-share statistic. Independent corroboration is available in a qualitative sense: NVIDIA's data-center segment accounted for approximately 78% of total company revenue in fiscal year 2024, and the absence of named alternative suppliers at comparable volume is structurally consistent with extreme concentration, even if the precise share figure is uncertain (NVIDIA Corporation, 2024; Patterson et al., 2022).
The same data show concentration on the demand side: a few hyperscale cloud providers account for the majority of GPU procurement, and a substantial portion of NVIDIA's data-center revenues derives from a handful of large customers (NVIDIA Corporation, 2024). The resulting market structure is bilateral rather than broad-based, concentrated among a few large suppliers and a correspondingly small number of hyperscale buyers.
At the same time, the evidence shows active attempts to reduce that dependence. Hyperscalers have invested in developing proprietary AI accelerators—such as AWS Trainium and Inferentia and Google TPUs—as alternatives to third-party GPUs, with reported gains in price-performance for selected workloads (NVIDIA Corporation, 2024). These numbers should be interpreted cautiously, as some figures derive from manufacturers' own performance benchmarks. Even so, they show that hyperscalers are not behaving as if the compute layer were permanently closed.
The pricing behavior of NVIDIA warrants specific analytical attention. NVIDIA's data-center GPU gross margins have been reported above 70%, and the company has maintained premium pricing even as demand from hyperscalers has accelerated (NVIDIA Corporation, 2024). This pricing posture is not straightforwardly explained by competitive pressure alone. A substantial part of NVIDIA's durability derives from CUDA, its proprietary parallel computing platform and programming interface. Because CUDA is deeply embedded in the machine learning software stack—training frameworks, inference libraries, developer toolchains, and research workflows have been built around it for over a decade—switching to alternative hardware entails not only hardware replacement costs but software migration costs that are substantially harder to quantify and bear. CUDA therefore functions as a complementary asset that extends NVIDIA's pricing power beyond what hardware concentration alone would support (NVIDIA Corporation, 2024; Cusumano et al., 2019).
This structure suggests a strategic logic consistent with the economics of durable monopoly under anticipated competition. An incumbent that expects its technical lead to erode over time has a rational incentive to extract rents aggressively while its advantage holds, rather than pricing to deter entry. The available evidence is consistent with this interpretation, though it does not definitively establish strategic intent: direct evidence of NVIDIA's internal strategic reasoning—such as executive statements specifically linking pricing strategy to anticipated hardware commoditization—is not available in the academic literature reviewed here. The interpretation should therefore be treated as an analytically motivated hypothesis about NVIDIA's competitive position rather than a demonstrated conclusion. What the evidence does show is the following: hyperscalers are investing in proprietary alternatives precisely because GPU costs are prohibitively high (NVIDIA Corporation, 2024); open-source alternatives such as DeepSeek's architecture demonstrate that frontier-quality capabilities can be reproduced at substantially lower compute cost (Cusumano, 2025); and inference market data show that model serving is becoming commoditized (Demirer et al., 2025). NVIDIA's observable expansion into software and services—CUDA ecosystem tools, inference frameworks, and cloud-GPU partnerships—is structurally consistent with a firm migrating value capture up the stack before hardware margins compress, regardless of whether that outcome is the result of deliberate strategy or rational response to market signals.
This observation carries a broader implication for H2. The concentration documented in the compute layer is real, but it is not a static equilibrium. It reflects a temporally bounded window in which a single firm controls a necessary input while downstream buyers have not yet scaled viable alternatives. The incumbency advantage is durable in the short run due to CUDA dependency, but structurally contested in the medium run as large buyers develop in-house capabilities and open-source compute frameworks mature.
Demirer et al. reveal a second notable distinction: the number of inference providers grew from 27 to 90 during 2025 and some open-source models were available from more than 20 providers (Demirer et al., 2025). Concentration is strongest where physical capital and manufacturing constraints bind most tightly; it is weaker where standardized model serving can be layered on top of shared infrastructure.
*Table 2. Evidence relevant to H2: concentration in compute supply.*
| Dimension | Quantitative finding | Source | Interpretation |
|---|---|---|---|
| Data-center GPU market share | ~98% Nvidia share | (NVIDIA Corporation, 2024) | Strong supplier concentration in accelerators |
| Leading-edge logic manufacturing | >90% share in advanced fabs | (NVIDIA Corporation, 2024) | Strong upstream fabrication concentration |
| Frontier cluster ownership | Small number of hyperscalers account for majority of GPU procurement | (NVIDIA Corporation, 2024) | Concentrated demand among hyperscalers |
| Revenue concentration | Substantial portion of Nvidia revenues from a few data-center customers | (NVIDIA Corporation, 2024) | Bilateral concentration in supply and demand |
| Inference provider growth | 27 → 90 providers in 2025 | (Demirer et al., 2025) | Service layer is less concentrated than training layer |
| Multi-hosted open models | Some models served by >20 providers | (Demirer et al., 2025) | Openness expands competition at inference layer |
The available evidence therefore does not support the view that the compute layer is monopolized, but rather that it is highly concentrated at the hardware and frontier training levels, while becoming more plural at the inference-service level.
Table 2 summarizes the quantitative evidence reviewed for H2.
H2 is partially supported.
4.4. H3 — Capability diffusion is rapid and commercially consequential
H3 states that improvements in foundation model capabilities diffuse rapidly across the ecosystem. The hypothesis draws on the economics of knowledge diffusion (Cockburn et al., 2018) and on platform theory's observation that once capabilities become broadly accessible, competition shifts from technical exclusivity toward price, integration, and complementary services (Gawer, 2014; Cusumano et al., 2019).
The strongest evidence comes from the market evolution documented by Demirer et al. During 2025, the number of models increased from about 253 to over 651, the number of model creators rose from 43 to 85, and the number of inference providers rose from 27 to 90 (Demirer et al., 2025). These figures indicate an ecosystem undergoing rapid technical and commercial turnover, inconsistent with a slow-diffusion interpretation.
The same study documents dramatic price compression: models that were frontier systems in 2023 experienced price declines on the order of 1000× by 2025 (Demirer et al., 2025). This figure warrants methodological qualification: Demirer et al. measure price per token at a given measured intelligence level, meaning the comparison controls for capability rather than comparing nominal prices of specific models over time. The magnitude is therefore partly a function of how capability is measured and how the frontier is defined at each point in time. Even with these caveats, the directional finding—that scarcity premiums attached to frontier capability erode rapidly—is corroborated by the independent pricing data from Cusumano's DeepSeek analysis and is consistent with the general economics of knowledge-intensive industries (Cockburn et al., 2018). It should also be noted that Demirer et al. is a working paper at time of writing and its findings are subject to peer review revision; the analysis here treats its quantitative results as indicative rather than definitive.
Cusumano's analysis of DeepSeek provides a concrete illustration of this process. V3 and R1 achieved benchmark-level performance close to frontier proprietary models while charging radically lower API prices: US$0.07/million input tokens and US$0.11/million output tokens for V3, and US$0.55/million input tokens and US$2.19/million output tokens for R1, compared with US$15 and US$60 per million tokens for GPT-4 at the time of the comparison (Cusumano, 2025). Even accounting for benchmark limitations, the pricing differential is commercially meaningful.
At the same time, diffusion is not uniform across use cases. Demirer et al. show that no single model dominates across all tasks, that programming tends to use models closer to the frontier, and that categories such as roleplay and translation rely more heavily on lower-intelligence models (Demirer et al., 2025). Capability diffusion therefore does not eliminate segmentation; it narrows technical exclusivity while preserving differentiated demand.
*Table 3. Evidence relevant to H3: capability diffusion and price compression.*
| Metric | Quantitative finding | Source | Interpretation |
|---|---|---|---|
| Number of models | 253 → 651 in 2025 | (Demirer et al., 2025) | Rapid expansion of supply |
| Number of creators | 43 → 85 in 2025 | (Demirer et al., 2025) | Frontier capability is no longer confined to a few labs |
| Number of providers | 27 → 90 in 2025 | (Demirer et al., 2025) | Commercial diffusion through serving markets |
| Price decline of former frontier models | ~1000× | (Demirer et al., 2025) | Frontier capability loses scarcity quickly |
| Open vs. closed model price gap | Open models ~90% cheaper at similar intelligence | (Demirer et al., 2025) | Capability diffusion reshapes market pricing |
| DeepSeek API pricing | V3: $0.07/$0.11; R1: $0.55/$2.19; GPT-4: $15/$60 | (Cusumano, 2025) | Frontier-like capability can diffuse through low-price challengers |
Table 3 consolidates the quantitative metrics supporting H3.
H3 is supported.
4.5. H4 — Application-layer entry barriers are materially lower, but dependence on platforms remains high
H4 states that entry barriers are lower in the application layer than in the compute or model layers. The theoretical basis derives from the layered structure of the generative AI value chain: once foundation model capabilities are accessible through standardized interfaces, downstream actors can integrate them without independently replicating the underlying infrastructure (Bommasani et al., 2021; Cusumano et al., 2025). This logic parallels the general principle in platform economics that upper-layer participants rely on lower-layer capabilities without bearing their capital costs (Gawer, 2014; Cusumano et al., 2019).
The available empirical evidence supports H4. The core economic argument for lower application-layer entry barriers is comparative: building a frontier AI capability requires sustained investment in the hundreds of millions of dollars (NVIDIA Corporation, 2024; Cusumano, 2025), while accessing that same capability through an API requires no capital expenditure beyond the marginal cost of tokens consumed. This cost asymmetry means that organizations without access to frontier compute—the vast majority of firms—can nonetheless build AI-enabled applications, a condition that was structurally impossible before the emergence of model APIs (Bommasani et al., 2021; Cusumano et al., 2025). The MIT platform analysis confirms this: closed models accessed by API are usable by organizations that lack compute infrastructure or machine-learning teams, allowing them to rent intelligence on demand rather than produce it internally (Cusumano et al., 2025).
The rapid growth of application-layer activity is consistent with this structural condition, though it measures experimentation rather than commercial entry in the economic sense. After the introduction of custom GPTs, approximately 3 million custom versions were created within two months (Cusumano et al., 2025)—a figure that reflects the low barrier to prototyping rather than the number of commercially viable entrants. Demirer et al. provide complementary evidence on demand-side specialization: programming accounts for roughly 50% of token usage, while roleplay and technology each account for approximately 15% (Demirer et al., 2025), indicating rapid category formation around high-usage applications.
Nevertheless, the decline in entry barriers should not be overstated. Lower barriers to building do not mean lower barriers to distribution, trust, workflow integration, or monetization. The same application ecosystem is heavily mediated by platforms that standardize APIs, routing, pricing, and discoverability (Cusumano et al., 2025; Demirer et al., 2025). The application layer is easier to enter technically but remains dependent on platform-controlled interfaces.
*Table 4. Evidence relevant to H4: entry conditions in the application layer.*
| Metric / feature | Quantitative finding | Source | Interpretation |
|---|---|---|---|
| Custom GPT creation | ~3 million in 2 months | (Cusumano et al., 2025) | Rapid application-layer experimentation |
| Programming share of token usage | ~50% | (Demirer et al., 2025) | Application development responds quickly to accessible model APIs |
| Roleplay share of token usage | ~15% | (Demirer et al., 2025) | Non-enterprise application experimentation is also substantial |
| Technology share of token usage | ~15% | (Demirer et al., 2025) | Usage is concentrated in a few categories |
Table 4 summarizes the evidence for H4.
H4 is supported.
4.6. H5 — The platform layer has the strongest basis for value capture, but direct value allocation evidence remains limited
H5 states that the platform layer is likely to capture a disproportionate share of value in the generative AI ecosystem. The theoretical foundation is well established: multi-sided platforms create value by coordinating interactions across user groups and can extract rents by controlling access, visibility, switching, and participation (Gawer, 2014; Hagiu and Wright, 2015; Evans and Schmalensee, 2016; Cusumano et al., 2019). As model capabilities diffuse and become more substitutable, the theory predicts that coordination layers should become more, not less, strategically central.
The strongest empirical evidence is indirect but meaningful. OpenRouter provides a unified API across multiple models and providers, standardizes access, exposes price and performance information, and charges about 5–5.5% in fees depending on the usage mode (Demirer et al., 2025). This is a coordination rent: the platform does not need to own the underlying model to monetize the market interface.
The MIT platform analysis treats foundation models and their toolchains as a new application-development platform and argues that enterprises are likely to use a mix of closed large models and smaller open models, which raises the strategic value of orchestration and integration (Cusumano et al., 2025). As model supply expands and models become more substitutable, the platform that coordinates access, routing, governance, and workflow integration becomes more strategically central.
This does not yet establish that the platform layer captures the largest share of profit. Transparent cross-layer financial data showing value allocation among chip vendors, cloud firms, model creators, inference intermediaries, and application platforms are not yet available. But the observable market mechanisms point in the same direction as platform theory.
*Table 5. Evidence relevant to H5: platform coordination and monetization.*
| Platform mechanism | Quantitative / observed evidence | Source | Economic implication |
|---|---|---|---|
| Unified access across models | Multi-model standardized interface | (Demirer et al., 2025) | Reduces search and switching costs |
| Platform monetization | ~5–5.5% fee on usage | (Demirer et al., 2025) | Coordination can be monetized directly |
| Multi-model enterprise use | Mix of closed and open models expected | (Cusumano et al., 2025) | Raises value of orchestration layer |
| Custom GPT ecosystem | ~3 million custom versions in 2 months | (Cusumano et al., 2025) | Ecosystem growth increases platform relevance |
Table 5 collects the observable evidence for H5.
H5 is theoretically grounded, but direct empirical verification remains incomplete.
4.7. H6 — Productivity gains are measurable and heterogeneous
H6 states that generative AI produces measurable and heterogeneous productivity gains in knowledge work, and that the distribution of these gains across firms and sectors is conditioned by the supply-side structure of the value chain. The first part of the hypothesis is grounded in the economics of general-purpose technologies, which predicts that once a new capability becomes widely accessible, productivity effects should emerge in tasks for which the technology provides a good fit (Brynjolfsson et al., 2017; Agrawal et al., 2018). The second part—the distribution claim—follows from the value-chain model: if capability access is mediated by platform layers that can extract coordination rents (H5), and if inference costs are shaped by compute concentration (H2), then the terms on which productivity gains reach end users are determined by the competitive structure of layers 4–6, not merely by the existence of capable models.
Among the six hypotheses, H6 is supported by the most direct empirical evidence. Noy and Zhang show that participants using generative AI completed professional writing tasks 40% faster and produced output rated 18% higher in quality (Noy and Zhang, 2023). Peng et al. find that developers using GitHub Copilot completed coding tasks 55.8% faster than the control group (Peng et al., 2023). Brynjolfsson, Li, and Raymond, using data on 5,172 customer-support agents, report average productivity gains of approximately 14%, with larger gains for less experienced workers (Brynjolfsson et al., 2023). Dell'Acqua et al. show that consultants using AI completed 12.2% more tasks and 25.1% faster, while also achieving better quality ratings in relevant cases (Dell'Acqua et al., 2023).
These effect sizes are economically meaningful and consistent across diverse task settings, spanning laboratory experiments, field experiments, and quasi-experimental designs. But the literature also shows that the effects are heterogeneous. Gains depend on task structure, baseline worker skill, and workflow compatibility with LLM assistance. The “jagged frontier” framing in Dell'Acqua et al. carries particular analytical weight: generative AI does not improve all knowledge work equally (Dell'Acqua et al., 2023).
*Table 6. Evidence relevant to H6: productivity effects of generative AI.*
| Study | Context | Quantitative effect | Key qualitative result |
|---|---|---|---|
| Noy & Zhang (2023) | Writing tasks | 40% faster; 18% higher quality | Strong gains in structured language production |
| Peng et al. (2023) | Coding with Copilot | 55.8% faster | Large gains in bounded programming tasks |
| Brynjolfsson et al. (2023) | 5,172 customer-support agents | ~14% productivity on average | Less experienced workers benefit more |
| Dell'Acqua et al. (2023) | Consulting / knowledge work | +12.2% tasks; 25.1% faster | Gains vary sharply by task type |
Generative AI does not uniformly transform all knowledge work; the evidence supports the conclusion that measurable, often large productivity gains exist in a meaningful subset of knowledge-intensive tasks, and these gains are sufficiently large to be economically consequential.
Table 6 summarizes the productivity studies supporting H6.
Connecting H6 to the value-chain model reveals a structural implication that the productivity literature alone does not surface. The application layer (layer 6) and the inference layer (layer 4) are the primary delivery mechanisms through which productivity gains reach end users. This means that the productivity effects documented in H6 are contingent on the competitive conditions in layers 4–6: as inference prices fall (H3) and application-layer entry barriers decline (H4), productivity-enhancing AI tools become accessible to a wider range of firms and workers. Conversely, if compute concentration (H2) persists and platform dependency (H5) intensifies, access to productivity-enhancing AI may be mediated by platform gatekeepers who can extract rents from the demand side. H6 therefore cannot be read in isolation from the supply-side structure of the ecosystem: the distribution of productivity gains across firms and sectors depends on who controls the layers through which those gains are delivered.
H6 is supported.
4.8. Consolidated Assessment
*Table 7. Overall hypothesis assessment.*
| Hypothesis | Assessment | Main empirical basis |
|---|---|---|
| H1 — Capital intensity is concentrated in infrastructure and model training | Supported | (Kaplan et al., 2020; Hoffmann et al., 2022; Patterson et al., 2022; NVIDIA Corporation, 2024; Cusumano, 2025) |
| H2 — Compute supply is highly concentrated | Partially supported | (NVIDIA Corporation, 2024; Demirer et al., 2025) |
| H3 — Capability improvements diffuse rapidly across the ecosystem | Supported | (Demirer et al., 2025; Cusumano, 2025) |
| H4 — Application development has lower barriers to entry | Supported | (Cusumano et al., 2025; Demirer et al., 2025) |
| H5 — Platform layers capture disproportionate value | Theoretically grounded; empirically incomplete | (Gawer, 2014; Hagiu and Wright, 2015; Evans and Schmalensee, 2016; Cusumano et al., 2019; Demirer et al., 2025) |
| H6 — Generative AI improves productivity in knowledge work | Supported | (Noy and Zhang, 2023; Peng et al., 2023; Brynjolfsson et al., 2023; Dell'Acqua et al., 2023) |
The strongest empirical conclusions concern H1, H3, H4, and H6. The most uncertain are H2 and H5, not because the theory is weak, but because the market has not yet matured enough to expose stable long-run outcomes; preserving this distinction strengthens the analytical credibility of the assessment. Table 7 presents the consolidated assessment.
5. Discussion and Implications
5.1. Interpreting the Structure of the Generative AI Ecosystem
The empirical evaluation in Section 4 supports the six-layer model but also reveals dynamics that the static model does not fully capture. Three interpretive observations extend the model beyond its structural description.
First, the value chain is structurally asymmetric in a specific way: the layers with the highest capital intensity (semiconductors, cloud infrastructure) are not the same layers with the highest potential for sustained value capture (orchestration and platform layers). This asymmetry has a direct implication for competitive analysis: the firms that absorb the largest capital investment are not necessarily those that will capture the largest economic returns as the market matures. The evidence from H3 suggests that model-layer margins will compress as capabilities diffuse, while orchestration layers may maintain coordination rents even as underlying model capabilities commoditize. This decoupling of capital intensity from value capture is structurally distinctive relative to prior technology ecosystems, where infrastructure providers (e.g., telecommunications carriers, cloud computing providers) typically captured rent commensurate with their capital base.
Second, the generative AI ecosystem is best understood as a system under transition rather than a mature equilibrium. The compute concentration documented in H2 reflects conditions of the 2023–2025 period and is already being contested by hyperscaler in-house chip programs and by the architectural efficiency demonstrated by DeepSeek's models (NVIDIA Corporation, 2024; Cusumano, 2025). The productivity effects documented in H6 reflect early-stage adoption studies whose long-run implications for organizational structure and labor markets remain unresolved (Brynjolfsson et al., 2017; Acemoglu and Restrepo, 2019). Analytical conclusions drawn from this period should be treated as structural tendencies rather than permanent equilibria.
Third, the stack exhibits a logic of competitive layering in which advantages at lower layers create options—but not guarantees—at upper layers. A firm controlling semiconductor supply has structural advantage but does not automatically convert that position into application-layer returns; the conversion requires deliberate investment in intermediate layers. Conversely, a firm with strong application-layer presence but no lower-layer position is exposed to dependency risks that can be exploited by integrated competitors. This logic of layered optionality is absent from both standard industrial organization models (which treat market structure as layer-independent) and from single-layer platform models (which treat the platform as the unit of analysis).
From an information systems perspective, the six-layer model offers a contribution that extends beyond describing a single industry. The generative AI value chain is the first instance of a digital platform ecosystem in which the core platform capability—the foundation model—is itself a continuously improving learning system subject to rapid performance diffusion. Prior IS research on digital platform ecosystems (Gawer, 2014; Hagiu and Wright, 2015; Evans and Schmalensee, 2016; Cusumano et al., 2019) developed its frameworks in contexts where the core platform capability (an operating system, a marketplace protocol, a social graph) was relatively stable between major versions, and where the primary dynamic was the accumulation of complementors and users around a fixed interface. The generative AI context violates this assumption: the core capability improves by orders of magnitude within a single year, diffuses to competing implementations almost immediately, and thereby continuously redefines what counts as a bottleneck and what counts as a commodity. This has direct implications for IS research on platform governance, standardization, and competitive strategy: models calibrated on stable-platform contexts may systematically underestimate the speed at which value migrates across layers in AI ecosystems.
The findings also carry implications for how organizations design their digital market strategies in AI-intensive environments. The six-layer value-chain model provides a practical diagnostic tool for identifying which layers of the AI ecosystem are accessible to a given organization given its capital base, technical capabilities, and existing relationships. Organizations that lack the capital to compete at layers 1–3 can nonetheless build differentiated positions at layers 5–6, provided they invest in the integration depth, domain expertise, and user relationships that constitute the sustainable source of competitive advantage in the application layer. The critical strategic error to avoid is conflating the low capital barrier to entry at layer 6 (any developer can access a foundation model via API) with a low barrier to sustainable competitive position (few application-layer firms achieve durable differentiation). The value-chain model makes this distinction analytically precise: low entry barriers and low moat are independent dimensions, and conflating them has led to overinvestment in undifferentiated AI applications and underinvestment in the platform and orchestration capabilities that create durable value.
Finally, the observation that value migrates toward coordination and orchestration layers as model capabilities commoditize has direct relevance for the design of electronic markets and networked business platforms. Electronic markets and B2B platforms that integrate generative AI capabilities as a feature enhancement may find themselves structurally repositioned as AI orchestration becomes the primary coordination mechanism for enterprise digital workflows. The implication for digital market design is that platform owners should invest in orchestration intelligence—routing, optimization, context management, workflow integration—rather than in the underlying AI capabilities themselves, which are subject to commoditization dynamics already well underway. This is a design principle grounded in the structural analysis of the value chain, with direct implications for platform architecture decisions in the electronic markets and networked business domain.
5.2. Implications for Platform Theory and Innovation Economics
The findings carry specific implications for research on digital platforms and motivate three formal propositions that extend existing theory to the multi-layer AI context. Each proposition is explicitly grounded in a prior theoretical tradition, and the contribution is defined in terms of what the paper adds to that tradition.
Traditional platform theory focuses on systems that coordinate interactions among multiple user groups through a shared technological interface (Gawer, 2014; Hagiu and Wright, 2015; Evans and Schmalensee, 2016; Gawer and Cusumano, 2002). Classic examples—operating systems, digital marketplaces, social networks—share a defining feature: the core platform is a stable coordinating infrastructure whose capabilities change slowly relative to the ecosystem it supports. Value capture derives from controlling access, participation, and switching costs at the interface layer. The theory of complementary assets (Teece, 1986) extends this logic by showing that profit from innovation accrues not to the innovator per se but to whoever controls the complementary assets required to bring capabilities to market. The theory of ecosystem architecture (Jacobides et al., 2018) further specifies that value capture depends on controlling architectural bottlenecks—the interfaces, standards, and coordination layers that other participants cannot bypass.
Generative AI departs from these configurations in two ways that existing theory does not explicitly model. First, the core platform capability is itself a learning system whose capabilities diffuse to competing systems relatively quickly, violating the static-capability assumption of prior platform models. Second, platform-relevant properties—multi-sidedness, network effects, complementor coordination—are present simultaneously at multiple layers of the stack, not only at a single coordination interface. These structural differences motivate the following propositions.
Proposition 1 (Diffusion-driven platform shift): In ecosystems where the core platform capability is a learning system subject to rapid diffusion, the locus of platform power is predicted to migrate over time from the layer controlling the capability (the model layer) to the layer controlling access, routing, and integration (the orchestration layer). This mechanism extends Teece's (1986) complementary-assets logic in a specific direction: when the innovation itself diffuses rapidly, the relevant complementary asset is not manufacturing or distribution (as in Teece's original examples) but coordination and integration infrastructure. The orchestration layer is the complementary asset that remains scarce as model capabilities commoditize. Available evidence is directionally consistent with this prediction but does not yet establish the dynamic trajectory: model-layer prices declined by approximately 1000× between 2023 and 2025, while orchestration-layer monetization mechanisms (e.g., routing fees on the order of 5–5.5%) were documented at a similar level during the same period (Demirer et al., 2025). The proposition should be understood as prospective—predicting where competitive advantage will concentrate as the ecosystem matures. Longitudinal data tracking margin trajectories across layers will be required to test it definitively.
Proposition 2 (Multi-layer platform simultaneity): In generative AI ecosystems, multiple layers exhibit platform characteristics simultaneously—the semiconductor layer exhibits supply-side concentration analogous to platform bottlenecks; the model layer mediates between compute suppliers and application developers in a two-sided configuration; the orchestration layer coordinates across models and users in a classic multi-sided market. This configuration extends Jacobides et al.'s (2018) ecosystem architecture framework, which identifies architectural bottlenecks as the primary locus of value capture, by showing that in AI ecosystems the bottleneck is not fixed at a single layer but distributed across layers simultaneously—and that different layers exhibit bottleneck-like concentration through different mechanisms (hardware lock-in at layer 1, CUDA dependency at layer 2, capability leadership at layer 3, routing control at layer 5). This implies that competitive analysis requires layer-specific market definitions rather than ecosystem-level characterizations, and that regulatory frameworks developed for single-interface platforms may be systematically mis-specified for multi-layer AI ecosystems.
Proposition 3 (Integration premium): Firms that span multiple platform-relevant layers can capture coordination rents that would otherwise be distributed across independent market transactions. This proposition extends both Teece (1986) and Jacobides et al. (2018): from Teece, it takes the insight that value capture requires control of complementary assets; from Jacobides et al., it takes the insight that architectural control determines who appropriates ecosystem surplus. The specific contribution is the mechanism of rent internalization: a firm controlling a sufficient span of the stack does not simply reduce transaction costs (the standard vertical integration argument) but converts inter-layer dependency—which single-layer firms experience as a constraint on their pricing power—into intra-firm value flow. The integration premium grows with the number of value-transferring inter-layer relationships internalized. The Google case illustrates this concretely: TPUs and Cloud at layers 1–2, Gemini at layer 3, inference via Google Cloud at layer 4, and Search/Android/Workspace at layers 5–6. This full-stack position means that Google internalizes every major inter-layer value transfer in the ecosystem. A firm occupying only layers 1–2 and 5–6, without controlling layers 3–4, would remain dependent on external model providers at the most strategically contested segment of the stack (Cusumano et al., 2019).
These three propositions represent theoretical advances over existing work. Proposition 1 introduces a dynamic complementary-asset mechanism absent from Teece's original static formulation. Proposition 2 identifies a multi-bottleneck configuration not captured by Jacobides et al.'s single-bottleneck model. Proposition 3 specifies a rent-internalization rationale for vertical integration that is distinct from both transaction cost economics and the cost-reduction rationale. Together, they provide a basis for future theoretical and empirical work on AI ecosystem governance, innovation policy, and competitive strategy.
5.3. Implications for Innovation Policy and Competitive Structure
The six-layer model and the three propositions advanced in the preceding section carry direct implications for innovation policy, market regulation, and the competitive structure of AI-intensive industries.
The first implication concerns the regulation of compute markets. The analysis of H2 shows that extreme concentration in GPU supply—sustained in part by the CUDA software stack as a complementary lock-in mechanism—creates a structural bottleneck that conditions innovation dynamics throughout the entire value chain. This is not a standard monopoly problem (a single price-setting seller in a stable market) but a dynamic bottleneck problem in which the incumbent extracts rents during a transitional window of technological leadership while downstream innovation is effectively taxed. The appropriate policy response is not conventional antitrust remedies, which address static market power, but interoperability requirements and open-interface standards for AI compute infrastructure—instruments more analogous to the network unbundling policies applied to telecommunications infrastructure (Dosi, 1982; Pavitt, 1984).
The second implication concerns the distributional effects of AI productivity gains. The evidence from H6 shows that productivity gains are real and substantial, but the analysis connecting H6 to the value-chain structure reveals that these gains flow to end users through specific bottleneck layers (inference and application). If compute concentration (H2) persists and platform dependency (H5) intensifies, the terms on which productivity-enhancing AI tools are accessible will be set by intermediaries who can extract rents from the demand side. This implies that innovation policy aimed at ensuring broad access to AI productivity gains must address the supply-side structure of the value chain, not merely the existence of AI capabilities.
The third implication concerns the industrial dynamics of AI ecosystems. Following Pavitt's (1984) framework for sectoral patterns of technical change, the generative AI value-chain exhibits characteristics of at least two distinct Pavittian regimes simultaneously: the semiconductor and infrastructure layers resemble the *scale-intensive* regime (high fixed costs, concentration, incremental innovation driven by capital investment), while the model and application layers resemble the *information-intensive* regime (rapid diffusion, low entry barriers, innovation driven by organizational knowledge). This co-existence of regimes within a single value chain is a structural novelty that creates distinctive innovation dynamics—including the decoupling of capital intensity from value capture documented in this paper—and warrants dedicated empirical investigation as the market matures.
Framework for layer positioning. The six layers differ along three strategically relevant dimensions: capital intensity of entry, durability of competitive advantage, and degree of platform dependency. Table 8 organizes these dimensions across the value chain as a reference for both researchers designing layer-level empirical studies and practitioners navigating positioning decisions.
*Table 8. Layer positioning framework: strategic dimensions across the generative AI value chain.*
| Layer | Capital intensity of entry | Durability of advantage | Primary basis of competition |
|---|---|---|---|
| Semiconductors / Hardware | Extremely high | High (manufacturing lock-in, CUDA ecosystem) | Proprietary architecture, fabrication access, software lock-in |
| Cloud infrastructure | Very high | Moderate-high (scale economies) | Infrastructure scale, energy cost, geographic coverage |
| Foundation model development | High | Low-moderate (rapid diffusion) | Research velocity, data access, safety reputation |
| Inference / model serving | Moderate | Low (price competition) | Cost efficiency, latency, model breadth |
| Platforms / orchestration | Low-moderate | High if network effects established | Routing control, API standardization, developer lock-in |
| Applications | Low | Low (competitive entry) | Domain expertise, workflow integration, user retention |
This framework yields three strategic archetypes. *Infrastructure incumbents* (e.g., NVIDIA, major hyperscalers) compete on capital depth and switching cost engineering; their moat is structural rather than technological. *Model competitors* (e.g., OpenAI, Anthropic, Google DeepMind) compete on research velocity but face rapid capability diffusion; sustainable differentiation requires either continuous frontier investment or migration toward orchestration and platform control. *Application specialists* face low entry barriers but also low moat; their sustainable advantage lies in domain integration depth and user retention rather than in the AI capability itself, which is increasingly available as a commodity input.
The table assigns high durability of advantage to the platforms and orchestration layer contingent on network effects being established—a condition that requires further elaboration. Network effects in AI orchestration platforms arise through two distinct mechanisms. The first is *data-side accumulation*: a routing platform that processes a large volume of requests across many models and use cases accumulates information about model performance, pricing, and user preferences that lower-volume rivals cannot replicate, creating an informational asymmetry that compounds over time. The second is *complementor lock-in*: as developers build workflows, integrations, and toolchains on top of a specific orchestration platform's APIs and abstractions, the switching cost of migrating to an alternative platform grows with the depth of integration. These mechanisms are structurally analogous to the network effects documented in prior platform ecosystems (Gawer, 2014; Hagiu and Wright, 2015), but they operate at a different layer and with different timing: they accumulate gradually as the ecosystem grows, rather than emerging at launch. This implies that early orchestration platforms that achieve sufficient developer adoption may establish durable positions, while late entrants face the classic platform chicken-and-egg problem even in a technically mature market.
Organizations competing in the infrastructure layer must invest heavily in compute capacity, data centers, and specialized hardware. These investments require large capital commitments but may benefit from economies of scale (NVIDIA Corporation, 2024). Firms operating at this layer therefore tend to be large technology companies or specialized semiconductor firms.
Organizations competing in the model layer must invest in machine learning research, training infrastructure, and data resources. The empirical evidence demonstrates that capabilities in this layer diffuse relatively quickly once techniques become widely known (Demirer et al., 2025; Cusumano, 2025). As a result, sustaining long-term differentiation may require continuous investment in research and infrastructure, or strategic migration toward platform and orchestration positions.
The application layer, by contrast, allows entry with far lower capital investment. Developers can access model capabilities through APIs or open-weight models (Cusumano et al., 2025). This enables rapid experimentation and the development of specialized applications across industries. However, application developers may depend heavily on upstream model providers and platform interfaces; the sustainability of their competitive position depends on factors orthogonal to AI capability—distribution, trust, and workflow integration.
The six-layer model also illuminates the strategic logic of vertical integration across the stack. A firm that competes across multiple layers simultaneously does not need to maximize margin at each layer independently. Instead, it can use ownership of lower layers to subsidize or strengthen its position in upper layers, and use upper-layer relationships to capture data, feedback, and usage signals that reinforce advantages in lower layers. This cross-layer consolidation of position can produce durable competitive advantages that single-layer firms cannot replicate, even if those single-layer firms are technically competitive within their own layer.
Google's vertical position in the generative AI ecosystem illustrates this logic concretely. Google designs and operates its own Tensor Processing Units (TPUs) at the semiconductor and infrastructure layer, reducing its dependence on NVIDIA's GPU supply chain and providing cost and latency advantages in model training and inference (NVIDIA Corporation, 2024). At the model layer, Google develops Gemini and related frontier systems, competing directly with OpenAI and Anthropic. At the platform and application layer, Google controls Search, Android, Chrome, and Workspace—distribution surfaces that allow it to deploy AI capabilities to billions of users without relying on third-party intermediaries. This vertical span means that Google can accept below-market margin in any individual layer, because the value of its position derives from the integrated system rather than from any single component. A firm that competes only in model development or only in applications faces Google not just as a rival at their own layer but as a system integrator whose cost structure and data flywheel—the self-reinforcing cycle in which usage generates data that improves the product and attracts further usage—extend across the entire stack (Cusumano et al., 2019; Bommasani et al., 2021).
This structural observation extends H5. The hypothesis states that platform layers are positioned to capture a disproportionate share of ecosystem value through coordination and routing mechanisms. Vertical integration across layers suggests a related but distinct mechanism: firms that span multiple layers can internalize value flows that would otherwise be distributed across independent market transactions. In this configuration, the question of which layer is most valuable becomes secondary to the question of which firm controls the most layers simultaneously. The strategic advantage of full-stack integration is precisely that it converts inter-layer dependency—which single-layer firms experience as a constraint—into an intra-firm asset.
These structural differences imply that firms must align their strategy with their capabilities and resources. Organizations with strong capital bases and infrastructure expertise may seek to expand across multiple layers of the value chain, as illustrated by the infrastructure investments of major cloud providers. Smaller firms may focus on specialized applications or domain-specific services built on foundation model APIs, competing on integration depth rather than on AI capability per se.
5.4. Limitations of the Current Evidence
Four limitations should be considered when interpreting the findings of this study.
First, the generative AI ecosystem is evolving rapidly. Many of the market patterns described in this paper are based on recent observations that may change as the industry matures.
Second, the available empirical literature provides limited visibility into profit allocation across the generative AI value chain. While studies document infrastructure costs, pricing trends, and productivity effects, few studies provide direct evidence regarding long-term margins or value capture across layers of the value chain.
Third, the six-layer model excludes several phenomena that may matter for competitive outcomes. It does not capture intra-layer asymmetries: within the application layer, the strategic position of a consumer application serving millions of users differs fundamentally from an enterprise application serving regulated industries, yet both appear at the same layer in the model. It does not capture layer convergence: a foundation model that performs inference, orchestration, and application functions simultaneously—as increasingly capable models are beginning to do—collapses the distinction between layers 3, 4, and 5, rendering the layer boundaries analytically fluid. And it does not capture deliberate non-integration: firms that refuse vertical integration for reasons of ecosystem positioning, regulatory risk, or partnership strategy may achieve outcomes that the integration premium logic would not predict. Empirical work refining the model should treat these boundary conditions as active research questions.
Fourth, all screening and selection decisions in the systematic review were made by a single reviewer. The absence of a second independent reviewer represents a limitation of the review protocol and may introduce unintentional selection bias. Future replications of this review should employ multi-reviewer procedures to enhance methodological rigor.
Each limitation points to a specific direction for further empirical work as the industry matures.
5.5. Directions for Future Research
The results of this study point to four directions for further research.
First, subsequent studies should construct layer-level financial accounts of the generative AI value chain. Firm-level revenue and margin data, disaggregated by infrastructure, model, platform, and application segments, would allow direct measurement of where profits accumulate and how that distribution shifts as the market matures.
Second, researchers should analyze the competitive dynamics of foundation model markets. As new architectures and training techniques emerge, it remains unclear whether the market will converge toward a small number of dominant models or remain fragmented.
Third, additional research is needed on the long-term productivity effects of generative AI adoption. Existing studies measure short-term task performance, but the long-term effects on organizational productivity and labor markets remain uncertain.
Finally, generative AI ecosystems offer a productive empirical setting for extending existing theories of digital platforms and technological ecosystems. Understanding how coordination, innovation, and value capture evolve in these multi-layer ecosystems remains an unresolved theoretical challenge.
6. Conclusion
6.1. Summary of Findings
This paper examined the economic structure of the generative AI ecosystem through a systematic analysis of 36 studies—20 peer-reviewed works (journal articles, conference papers, and academic books) and 16 grey literature sources (NBER working papers, arXiv preprints, and institutional technical reports)—organized around a six-layer value-chain model and evaluated against six structural hypotheses.
Four hypotheses were supported by the available empirical evidence. Capital intensity is concentrated in the infrastructure and frontier training layers (H1), consistent with scaling law research and infrastructure cost data. Capability improvements diffuse rapidly across the model and inference layers, compressing prices by orders of magnitude and eroding technical exclusivity (H3). Entry barriers are materially lower at the application layer, enabling rapid ecosystem growth built on API-accessible model capabilities (H4). Generative AI produces measurable and heterogeneous productivity gains in knowledge work, with effect sizes ranging from 14% to 55% across multiple experimental settings (H6) (Noy and Zhang, 2023; Peng et al., 2023; Brynjolfsson et al., 2023; Dell'Acqua et al., 2023).
One hypothesis was partially supported. Compute supply is highly concentrated—particularly at the hardware and frontier training levels—but the concentration is not absolute or irreversible; inference-layer competition is already more plural, and hyperscalers are investing in proprietary alternatives (H2) (NVIDIA Corporation, 2024; Demirer et al., 2025).
One hypothesis is theoretically grounded but empirically incomplete. Platform orchestration layers have the strongest theoretical basis for value capture as model capabilities commoditize, and observable coordination mechanisms are consistent with this prediction; however, direct cross-layer profit data sufficient for definitive verification are not yet available (H5) (Gawer, 2014; Demirer et al., 2025).
Two additional findings extend the core results. First, the compute layer concentration reflects a temporally bounded window shaped by CUDA-based switching costs and pricing behavior consistent with rent extraction under anticipated competition, not a stable equilibrium. Second, vertical integration across multiple stack layers converts inter-layer dependency into intra-firm value flow, producing competitive advantages that no single-layer competitor can replicate through technical excellence alone. Both findings should be understood as theory-grounded structural observations, not empirically verified conclusions: cross-layer profit data and long-run competitive outcomes are not yet available.
6.2. Theoretical Contributions
This paper makes three contributions to the literature.
First, it proposes a six-layer conceptual model of the generative AI value chain that integrates insights from scaling law research, platform economics, and the innovation-studies tradition on sectoral patterns of technical change (Pavitt, 1984; Dosi, 1982). The model identifies a structural novelty—the co-existence of scale-intensive and information-intensive innovation regimes within a single value chain—that prior work on AI has not systematically characterized. While previous studies have examined individual components of the ecosystem, none has analyzed how these layers interact within a unified structure that accounts for their distinct capital intensity, competitive dynamics, and value capture mechanisms.
Second, it provides an empirical evaluation of six structural hypotheses using a systematic corpus of 36 studies, distinguishing between claims supported by peer-reviewed evidence, claims supported by grey literature and working papers (treated as indicative), and claims that remain theoretically grounded but empirically incomplete. This disciplined treatment of evidential uncertainty is itself a methodological contribution to the emerging literature on AI market structure, where the pace of development has outrun peer-reviewed publication.
Third, it extends platform theory and the complementary-assets framework to the multi-layer generative AI context through three formal propositions. Proposition 1 advances Teece's (1986) complementary-assets logic by identifying coordination infrastructure as the relevant complementary asset in rapidly-diffusing learning-system ecosystems. Proposition 2 extends Jacobides et al.'s (2018) single-bottleneck ecosystem model to a multi-bottleneck configuration in which different layers exhibit concentration through different mechanisms. Proposition 3 specifies a rent-internalization mechanism for vertical integration that is distinct from both transaction cost economics and cost-reduction rationales, advancing the architectural-control logic of ecosystem theory to AI value chains. Together, these propositions provide a theoretically grounded basis for future empirical work on AI ecosystem governance, innovation policy, and competitive strategy.
6.3. Implications for Research and Industry
The findings of this study have direct implications for both research and practice.
For researchers, the generative AI ecosystem offers a productive empirical setting for extending existing theories of digital platforms and technological ecosystems. The evidence presented here demonstrates that generative AI introduces a multi-layer architecture in which different layers exhibit distinct competitive dynamics. Understanding how these layers interact and co-evolve remains an unresolved theoretical challenge.
For industry practitioners, the results clarify the strategic logic of layer positioning within the generative AI value chain. Firms competing in infrastructure or model development must sustain large capital investments, while firms operating in downstream application layers may benefit from lower entry barriers but face greater competitive pressure. Platform orchestration layers are likely to become more strategically central as the industry grows in complexity.
Acknowledgment
This research was conducted as part of the author's extended research and residency activities at the Sloan School of Management within the Visiting Fellows Program at the Massachusetts Institute of Technology (MIT). The author, a professor and researcher at the CESAR School (Recife Center for Advanced Studies and Systems), gratefully acknowledges the academic environment and interdisciplinary dialogue that supported the intellectual development of this study.
References
- Acemoglu, Daron, and Restrepo, Pascual (2019). Automation and New Tasks: How Technology Displaces and Reinstates Labor. *Journal of Economic Perspectives*, 33(2):3–30. https://doi.org/10.1257/jep.33.2.3
- Agrawal, Ajay, Gans, Joshua, and Goldfarb, Avi (2018). *Prediction Machines: The Simple Economics of Artificial Intelligence*. Harvard Business Review Press, Boston, MA.
- Anthropic (2023). Claude's Model Card. Technical Report, Anthropic. Available: https://www.anthropic.com/claude.
- Bender, Emily M., Gebru, Timnit, McMillan-Major, Angelina, and Shmitchell, Shmargaret (2021). On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? In *Proceedings of the ACM Conference on Fairness, Accountability, and Transparency*, pages 610–623.
- Bommasani, Rishi, Hudson, Drew A., Adeli, Ehsan, et al. (2021). On the Opportunities and Risks of Foundation Models. arXiv preprint arXiv:2108.07258.
- Bommasani, Rishi, et al. (2023). The Foundation Model Transparency Index. Technical Report, Stanford Center for Research on Foundation Models (CRFM). Available: https://crfm.stanford.edu/fmti.
- Brown, Tom B., Mann, Benjamin, Ryder, Nick, Subbiah, Melanie, Kaplan, Jared, et al. (2020). Language Models are Few-Shot Learners. In *Advances in Neural Information Processing Systems*, volume 33, pages 1877–1901.
- Brynjolfsson, Erik, Rock, Daniel, and Syverson, Chad (2017). Artificial Intelligence and the Modern Productivity Paradox: A Clash of Expectations and Statistics. Working Paper 24001, National Bureau of Economic Research.
- Brynjolfsson, Erik, Li, Danielle, and Raymond, Lindsey (2023). Generative AI at Work. Working Paper 31161, National Bureau of Economic Research.
- Cockburn, Iain M., Henderson, Rebecca, and Stern, Scott (2018). The Impact of Artificial Intelligence on Innovation. Working Paper 24449, National Bureau of Economic Research.
- Cusumano, Michael A., Gawer, Annabelle, and Yoffie, David B. (2019). *The Business of Platforms*. Harper Business, New York, NY.
- Cusumano, Michael A., Farias, Vivek, and Ramakrishnan, Rama (2025). Generative AI as a Platform for Applications Development. MIT CISR Research Briefing; forthcoming in *Communications of the ACM*.
- Cusumano, Michael A. (2025). The DeepSeek Moment: Strategic Implications for AI Platforms. Forthcoming in *Communications of the ACM*.
- Dell'Acqua, Fabrizio, McFowland III, Edward, Mollick, Ethan R., Lifshitz-Assaf, Hila, Kellogg, Katherine, Rajendran, Saran, Krayer, Lisa, Candelon, François, and Lakhani, Karim R. (2023). Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality. Working Paper 24-013, Harvard Business School.
- Demirer, Mert, Fradkin, Andrey, and Peng, Sida (2025). The Emerging Market for Intelligence: Pricing, Supply, and Demand for LLMs. Working Paper 34608, National Bureau of Economic Research.
- Dosi, Giovanni (1982). Technological Paradigms and Technological Trajectories. *Research Policy*, 11(3):147–162. https://doi.org/10.1016/0048-7333(82)90016-6
- Evans, David S., and Schmalensee, Richard (2016). *Matchmakers: The New Economics of Multisided Platforms*. Harvard Business Review Press, Boston, MA.
- Gawer, Annabelle, and Cusumano, Michael A. (2002). *Platform Leadership: How Intel, Microsoft, and Cisco Drive Industry Innovation*. Harvard Business School Press, Boston, MA.
- Gawer, Annabelle (2014). Bridging Differing Perspectives on Technological Platforms: Toward an Integrative Framework. *Research Policy*, 43(7):1239–1249. https://doi.org/10.1016/j.respol.2014.03.003
- Goodfellow, Ian, Bengio, Yoshua, and Courville, Aaron (2016). *Deep Learning*. MIT Press, Cambridge, MA.
- Hagiu, Andrei, and Wright, Julian (2015). Multi-Sided Platforms. *International Journal of Industrial Organization*, 43:162–174. https://doi.org/10.1016/j.ijindorg.2015.09.005
- Hoffmann, Jordan, Borgeaud, Sebastian, Mensch, Arthur, Buchatskaya, Elena, Cai, Trevor, Rutherford, Eliza, et al. (2022). Training Compute-Optimal Large Language Models. arXiv preprint arXiv:2203.15556.
- Jacobides, Michael G., Cennamo, Carmelo, and Gawer, Annabelle (2018). Towards a Theory of Ecosystems. *Strategic Management Journal*, 39(8):2255–2276. https://doi.org/10.1002/smj.2904
- Jordan, Michael I., and Mitchell, Tom M. (2015). Machine Learning: Trends, Perspectives, and Prospects. *Science*, 349(6245):255–260. https://doi.org/10.1126/science.aaa8415
- Kaplan, Jared, McCandlish, Sam, Henighan, Tom, Brown, Tom, Chess, Benjamin, Child, Rewon, Gray, Scott, Radford, Alec, Wu, Jeffrey, and Amodei, Dario (2020). Scaling Laws for Neural Language Models. arXiv preprint arXiv:2001.08361.
- LeCun, Yann, Bengio, Yoshua, and Hinton, Geoffrey (2015). Deep Learning. *Nature*, 521(7553):436–444. https://doi.org/10.1038/nature14539
- Noy, Shakked, and Zhang, Whitney (2023). Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence. *Science*, 381(6654):187–192. https://doi.org/10.1126/science.adh2586
- NVIDIA Corporation (2024). Annual Report, Fiscal Year 2024. Technical Report, NVIDIA Corporation. Available: https://investor.nvidia.com/financial-information/annual-reports/.
- OpenAI (2023). GPT-4 Technical Report. arXiv preprint arXiv:2303.08774.
- Patterson, David, Gonzalez, Joseph, Le, Quoc, Liang, Chen, Munguia, Lluis, Rothchild, Daniel, So, David, Texier, Maud, and Dean, Jeff (2022). Carbon Emissions and Large Neural Network Training. *IEEE Micro*, 42(4):59–68. https://doi.org/10.1109/MM.2022.3163226
- Pavitt, Keith (1984). Sectoral Patterns of Technical Change: Towards a Taxonomy and a Theory. *Research Policy*, 13(6):343–373. https://doi.org/10.1016/0048-7333(84)90018-0
- Peng, Sida, Kalliamvakou, Eirini, Cihon, Peter, and Demirer, Mert (2023). The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. arXiv preprint arXiv:2302.06590.
- Teece, David J. (1986). Profiting from Technological Innovation: Implications for Integration, Collaboration, Licensing and Public Policy. *Research Policy*, 15(6):285–305. https://doi.org/10.1016/0048-7333(86)90027-2
- Varian, Hal (2018). Artificial Intelligence, Economics, and Industrial Organization. Working Paper 24839, National Bureau of Economic Research.
- Vaswani, Ashish, Shazeer, Noam, Parmar, Niki, Uszkoreit, Jakob, Jones, Llion, Gomez, Aidan N., Kaiser, Lukasz, and Polosukhin, Illia (2017). Attention Is All You Need. In *Advances in Neural Information Processing Systems*, volume 30, pages 5998–6008.
- Weidinger, Laura, Mellor, John, Rauh, Maribeth, Griffin, Conor, et al. (2022). Taxonomy of Risks Posed by Language Models. In *Proceedings of the ACM Conference on Fairness, Accountability, and Transparency*, pages 214–229. https://doi.org/10.1145/3531146.3533088
Comments
Every comment is moderated before it appears here. Nothing is published automatically.
Loading…