Building the Intelligence Age: AI’s Seventy-Five-Year Journey and the Race to 2031
How seventy-five years of research became a global industrial system—and how infrastructure, agents and policy will determine who captures the next wave of value.
Modern AI is the product of decades of research and engineering across universities, public laboratories and technology companies. Advances in computation, semiconductors, learning algorithms, data systems and cloud infrastructure have transformed machine intelligence into a global industry. Through 2031, the opportunity is to convert this expanding foundation into reliable action, sustainable economic growth and broadly shared human and public benefit.
Do not add the headline numbers. Infrastructure capex becomes supplier revenue. Enabled value is an opportunity pool. Captured value is the realized share. Net benefit subtracts enterprise implementation, run, governance and transition cost. Vendor revenue, profit, consumer surplus and GDP are separate accounting objects.
This expanded edition integrates the economic model in Agentic AI: $16Tn in Global Economic Value Enabled and $8Tn Captured, the workflow architecture in Agentic AI: From Better Answers to Entirely New Industry Economics, and the industry/startup theses in Agentic AI: Transforming Industries.
Explore the expanded edition
- Executive briefing
- Prologue. Before AI had a name: 1936–1955
- 0. 1956–2009: research foundations and the hidden computing platform
- 1. 2010–November 2022: analytics, deep learning and foundation models become industrial
- 2. December 2022–2026: mass distribution, multimodality and agents
- 3. The AI factory: chips, data centers, power and water
- 4. The model and cloud economy: laboratories, training and inference
- 5. What current AI cannot reliably do—and the research attempting to fix it
- 6. Adoption becomes application: consumers, enterprises, industries and government
- 7. The money flow: capex, venture capital, revenues, valuations and returns
- 8. National and human consequences: sovereignty, work, security and trust
- 9. 2027–2031: four futures and one operating agenda
- 10. Beyond 2031: six questions that will define the 2035 intelligence economy
- Conclusion: the race is to turn machine intelligence into shared capability
- Methods, definitions and publication package
Executive briefing
The question this article answers
This article explains AI as one industrial system across four broad periods: the intellectual and computing foundations laid before 2010; the enterprise analytics, deep-learning and foundation-model buildout of 2010–2022; the distribution and capital discontinuity after ChatGPT; and the governed intelligence architecture that may emerge by 2031. It follows the chain from algorithms, data and programmable compute through warehouses, semiconductor tools, chips, data centers, electricity, cloud, frontier models, applications, governments and users. Its central question is which inventions become usable capacity, which investments become measurable outcomes, who retains the benefit and who bears the risk.
The article keeps four ledgers separate. Capital formation pays for assets. Supplier revenue records payments to providers. Value enabled estimates the feasible improvement. Net benefit is what remains after implementation, operation, control and transition. Combining them would count the same dollar several times.
Central thesis
Modern AI emerged from a long continuum of research, enterprise investment and infrastructure development. Symbolic reasoning, expert systems, backpropagation, recurrent memory, statistical learning, distributed computing, programmable GPUs and large labeled datasets supplied the ingredients before a conversational product made them visible. From 2010 onward, enterprises spent heavily on data warehouses, Hadoop and Spark clusters, business-intelligence tools, data integration, statistical software and specialist teams. They trained regression, classification, recommendation, anomaly-detection and time-series models to price risk, forecast demand, detect fraud, predict churn and optimize operations. Deep learning then replaced manual feature engineering in important perception tasks; transformers and pretraining made capabilities reusable across tasks. After November 30, 2022, conversational distribution made general-purpose capability immediately accessible at global scale. Experimentation accelerated, users arrived at mass-market speed and developers gained a reusable platform.
That product event became an industrial buildout. Amazon, Alphabet, Meta and Microsoft moved from roughly $95 billion of combined gross capital expenditure or property-and-equipment additions in 2020 to $151.1 billion in 2022 and about $360 billion actually reported for 2025. Their latest calendar-2026 plans and discussion total approximately $735–760 billion. Those figures include logistics, offices and conventional cloud assets as well as AI. They measure the direction and financing pressure of the buildout, not a clean AI-only total.
The same transition is visible in realized revenue. NVIDIA’s data-center revenue rose from $6.7 billion in FY2021 and $15.0 billion in FY2023 to $194 billion in FY2026. Cloud-infrastructure revenue rose from $129 billion in 2020 to $227 billion in 2022 and $500 billion for the trailing twelve months through Q2 2026. The first returns accrued upstream, where scarcity was easiest to price. The next test is whether installed capacity produces customer revenue, productivity and public benefit before compute depreciates and model prices compress. The race is shifting from impressive output to reliable action.
Coding agents provide the clearest early bridge from generated content to completed work. Claude Code, Codex, Gemini Code Assist and CLI, Grok Build, Qwen Code and open models from DeepSeek, Qwen, GLM and Llama can inspect repositories, edit several files, run commands and tests, recover from failures and prepare reviewable changes. Claude Code reached a company-reported $1 billion revenue run rate within six months. Across three randomized company experiments involving 4,867 developers, AI code completion increased completed tasks by about 26%; a smaller METR study of experienced open-source developers working in familiar repositories found early-2025 tools made tasks 19% slower. The contrast is central to this article: access and code volume are not value. Task selection, context, verification, expertise and workflow design determine whether agent speed becomes financial return.
Ten eras in the construction of machine intelligence
These are analytical boundaries, not claims that one person or company invented an era alone. Research, hardware and commercial adoption overlap; an algorithm can precede its largest economic effect by decades.
| Era | What became technically possible | Representative contributors and industrial enablers | Economic inheritance and unresolved constraint |
|---|---|---|---|
| 1936–1955: computation, information, feedback and electronics | General computation, mathematical neurons, information theory, feedback control and electronic switching created the prerequisites for machine intelligence. | Alan Turing; Warren McCulloch and Walter Pitts; Claude Shannon; Norbert Wiener; wartime operations-research groups; Bell Labs transistor teams | Computation became programmable and scalable in principle; machines, memory and data remained extremely scarce. |
| 1956–1969: symbolic foundations and early learning | Logic, search, planning and perceptrons framed intelligence as an engineering problem. | John McCarthy, Marvin Minsky, Allen Newell, Herbert Simon and Frank Rosenblatt; Dartmouth, MIT, Carnegie Mellon and IBM | Established planning and formal representation; broad claims exceeded compute and robust knowledge. |
| 1970–1987: knowledge engineering and AI winters | Expert knowledge could drive narrow inference, diagnosis and configuration systems. | Edward Feigenbaum, Bruce Buchanan, Joshua Lederberg and Edward Shortliffe; Stanford, DEC, IBM and Japan’s Fifth Generation project | Created specialist decision support; knowledge acquisition, exceptions and maintenance prevented general scale. |
| 1988–2005: statistical and probabilistic learning | Models learned scoring, recognition and prediction from examples instead of receiving every rule. | Vladimir Vapnik, Corinna Cortes, Judea Pearl, Leo Breiman, Yann LeCun, Sepp Hochreiter and Jürgen Schmidhuber; databases, CPUs and internet platforms | Powered risk, fraud, forecasting, search and recommendation; labeled data and feature engineering remained expensive. |
| 2006–2011: scalable data and representation learning | Distributed data systems, shared benchmarks and programmable GPUs made larger learned representations practical. | Geoffrey Hinton, Yoshua Bengio, Yann LeCun and Fei-Fei Li; Google distributed systems, Hadoop, AWS and NVIDIA CUDA | Created the data–compute flywheel; training stability, accelerator access and specialist talent constrained adoption. |
| 2012–2016: deep learning becomes industrial | CNNs, sequence models, GANs and reinforcement learning transformed perception, translation and generation. | Alex Krizhevsky, Ilya Sutskever, Geoffrey Hinton, Ian Goodfellow, Demis Hassabis and David Silver; GPUs, TPUs and hyperscale clouds | Displaced many hand-engineered pipelines; robustness, labels and accelerated infrastructure remained bottlenecks. |
| 2017–2020: transformers and reusable pretraining | Attention, self-supervision and scaling made one pretrained model useful across many tasks. | The transformer, BERT, GPT and scaling-law teams; Google TPUs, NVIDIA datacenter GPUs, OpenAI, Google, Meta and cloud platforms | Created foundation-model economics; high-quality data, compute and serving cost concentrated frontier production. |
| 2021–November 2022: alignment and multimodality | Compute-optimal training, instruction tuning, preference learning and diffusion made foundation models more useful and accessible. | Chinchilla, InstructGPT and diffusion researchers; HBM, advanced packaging, AI supercomputers and frontier laboratories | Prepared a general model for mass interaction; truthfulness, alignment, distribution and inference economics remained unresolved. |
| December 2022–2026: mass distribution and the AI factory | Conversation, multimodality, reasoning, open-weight competition and bounded tool use made AI a consumer and board-level market. | OpenAI, Anthropic, Google DeepMind, Meta, xAI, DeepSeek, Qwen, Kimi and GLM; NVIDIA, TSMC, memory suppliers, clouds and utilities | Created mass use and unprecedented capital formation; utilization, power, reliable action and institutional absorption became decisive. |
| 2027–2031: governed agentic systems—scenario | Routed model portfolios, memory, tools, specialist agents and independent verification may perform bounded work across institutions. | No settled winner: laboratories, open communities, enterprises, standards bodies and governments all contribute. | Continuous inference and workflow redesign enlarge value; trust, accountability, diffusion and benefit sharing determine legitimacy. |
Quantified signals that define the transition
The dates and denominators matter. Observed results, estimates, external forecasts and article scenarios are labeled rather than blended.
| System signal | Past: 2020–2022 | Current: 2025–2026 | Future: 2030–2031 | Evidence label and interpretation |
|---|---|---|---|---|
| Enterprise data and AI investment | big-data technology/services rose from less than $5B in 2010 to roughly $15–20B in 2015; four-company broad capex reached ~$95B in 2020 and $151.1B in 2022 | ~$360B actual 2025; ~$735–760B 2026 plans for the same four-company basket | Goldman: $7.6T cumulative AI ecosystem buildout, 2026–2031 | Observed estimates/guidance/scenario. Historical big-data revenue and company capex have different boundaries; do not add them. |
| Data-center investment and buildout | global investment subsequently nearly doubled from 2022 | ~$500B in 2024 | Goldman: $7.6T cumulative AI ecosystem buildout, 2026–2031 | IEA estimate / bank scenario. Scope overlaps company capex; do not add. |
| Data-center electricity | 240–340 TWh in 2022 | 415 TWh in 2024; 485 TWh in 2025 | ~950 TWh IEA / >1,200 TWh Gartner in 2030 | Estimates/forecasts. All data centers, not AI alone; forecasts use different models. |
| Semiconductor supplier capture | NVIDIA data-center revenue $6.7B FY2021; $15.0B FY2023 | $194B FY2026 | no company forecast used | Observed. Supplier revenue is not customer value. |
| Cloud-infrastructure revenue | $129B in 2020; $227B in 2022 | ~$145B in Q2 2026; $500B trailing 12 months | AI and agents become a larger workload share | Market estimate. Includes conventional cloud as well as AI. |
| Laboratory and venture financing | GenAI VC $2.8B in 2022; AI was ~30% of global VC | OpenAI $122B committed at $852B post-money; Anthropic $65B at $965B; AI VC $258.7B in 2025 | capital likely remains concentrated around few labs and infrastructure firms | Transactions/estimate. Rounds, valuations and total VC are different objects. |
| Consumer reach | no comparable mass general-purpose assistant destination | ChatGPT >900M WAU; Gemini 950M MAU; China 602M GenAI users | assistants spread across work, devices, search, messaging and machine traffic | Company/official reports. WAU, MAU and individuals do not create valid market share. |
| Coding-agent adoption and revenue | autocomplete and narrow code generation dominated | Claude Code $1B run-rate within six months; 90% of DORA respondents used AI at work; controlled studies range from 26% more completed tasks to 19% slower in a different task population | background and parallel agents move from code suggestion toward tested changes, migrations, operations and software-created workflows | Company disclosure / surveys / experiments. Product revenue, perceived productivity and causal task outcomes are separate measures. |
| Formal business adoption | executive survey 50% in at least one function in 2020 | OECD firms 8.7% in 2023 → 20.2% in 2025; executive survey 88% in 2025 | value shifts from access toward completed, governed workflows | Official statistics/surveys. Population adoption is much lower than executive survey breadth. |
| Government adoption | fragmented pilots and narrow automation | 11 U.S. agencies: 571 → 1,110 AI cases, 2023–2024; GenAI 32 → 282 | governments become buyers, infrastructure coordinators and providers of AI-enabled capacity | Inventory. About 125 GenAI cases were beyond early stages; case count is not public value. |
| Economic outcome | no comparable GenAI market | benefits measurable but uneven | $6.42T annual enabled; $4.10T captured before cost; $3.40T net; $16.06T cumulative enabled, 2027–2031 | Article scenario. Sequential stages of one waterfall; never add them. |
Seven findings
- ChatGPT was a distribution discontinuity, not the birth of AI. Traditional ML and early foundation models already created operational value. ChatGPT joined general capability to an accessible interface and feedback loop. The resulting demand signal mattered more economically than any single benchmark result.
- AI has become a physical industry. Useful intelligence depends on the minimum throughput of tools, fabs, HBM, packaging, networks, racks, cooling, power and cloud operations. Money cannot instantly reproduce yield learning, grid connections, qualified suppliers, technicians or community consent.
- The first visible value accrued to scarcity; the next value requires utilization. Accelerators, memory, networking, cloud capacity and energized sites were paid before most enterprises redesigned workflows. The return test now shifts from announced capital to utilization, gross profit per energized megawatt, end-customer cash flow and cost per successful task.
- Model capability is diffusing faster than frontier production capacity. Closed systems remain important, but Qwen, DeepSeek, Kimi, GLM, Llama and Mistral place strong capability into downloadable, price-competitive forms. Benchmark convergence does not eliminate gaps in uptime, tools, safety or support; it reduces the likelihood that raw model access remains a durable rent.
- Coding agents are the first high-frequency market for verifiable agency. Software supplies explicit files, commands, tests, version control and rollback, so an agent can do more than draft an answer. The result is a preview of agentic work in finance, science and operations—but only when generated changes are reviewed, tested and connected to a real objective.
- Consumer adoption is far ahead of institutional absorption. Hundreds of millions of people can use an assistant immediately; an enterprise or government cannot redesign permissions, data, labor roles, accountability, procurement and appeal rights at the same speed. The principal bottleneck is increasingly the institution around the model, not access to generation.
- The 2031 prize is verified action and broad diffusion. Route each task to the smallest sufficient model, connect approved data and tools, constrain authority, verify the completed state and escalate consequential exceptions. Countries gain more from connectivity, usable compute, local context, skills and institutions than from an underused prestige model.
How to read this article
The article is designed in layers so a general reader does not need to follow every technical or financial detail.
- For the full historical and technical story: read the pre-1956 prologue and Sections 0–5.
- For business, technology and investment: begin with Sections 2–7, then the 2031 scenarios.
- For government, labor and society: begin with Section 6, continue through Sections 8–10 and finish with the conclusion.
- For a fast briefing: read the central thesis, six findings, headline cards, the four-future matrix and the conclusion.
Every quantitative claim belongs to one of four evidence classes: observed, estimated, externally forecast or an article scenario. A high benchmark score is not treated as reliable production performance; capital expenditure is not treated as supplier revenue; and either is not treated as economic value.
Plain-language glossary
| Term | Plain meaning | Why it matters economically |
|---|---|---|
| Artificial intelligence | The broad field of building machines that perceive, predict, reason, generate or act. | Includes rules, statistics, optimization, neural networks and agents—not only chatbots. |
| Machine learning | Systems that estimate patterns from data rather than receiving every rule explicitly. | Powers forecasting, risk, recommendation, recognition and many pre-ChatGPT applications. |
| Deep learning | Machine learning using multilayer neural networks that learn internal representations. | Made vision, speech, translation and foundation models far more capable. |
| Transformer | A neural architecture that uses attention to relate elements in a sequence and trains efficiently on accelerators. | Underlies most current LLMs while creating substantial memory and inference demand. |
| Foundation model | A large pretrained model adapted to many tasks and products. | Spreads one research and training investment across a broad application market. |
| Large language model | A foundation model trained primarily to process and generate token sequences such as text or code. | Provides a general interface, but does not guarantee truth, memory or authorized action. |
| Training and post-training | Training learns broad patterns; post-training shapes usefulness, reasoning, tool use, policy and safety. | The final training run is only one part of a larger research and product cost. |
| Inference | Running a trained model to answer a request or perform part of a task. | Becomes the recurring cost as users, agents, modalities and reasoning steps multiply. |
| Open weight | Model parameters can be downloaded, although data, code, license or training process may remain restricted. | Supports control and portability without automatically meeting a full open-source definition. |
| Agent | Software that interprets a goal, plans steps, uses tools and observes results. | Can complete work rather than only draft an answer, but errors and permissions compound across steps. |
| Accelerator and HBM | A processor specialized for parallel AI work and the high-bandwidth memory feeding it. | Their supply, packaging, software and energy determine useful compute—not chip count alone. |
| Capex | Money spent on long-lived assets such as fabs, servers, data centers, substations and networks. | It finances future capacity; it is not proof that the capacity will be used profitably. |
| Verified outcome | A completed result checked against evidence, rules or observable system state. | It is the proper unit for measuring AI value, reliability and cost. |
Decision rule: Do not extrapolate one layer in isolation. More GPUs do not guarantee capacity if power is delayed; cheaper tokens do not guarantee lower system cost; benchmarks do not guarantee reliable execution; exposure is not unemployment; and valuation is not cash flow. Durable winners will convert physical capacity into verified outcomes and distribute enough of the gain to sustain legitimacy.
Prologue. Before AI had a name: 1936–1955
Artificial intelligence became a named research field in 1956, but its essential ingredients were already converging. The preceding twenty years established four propositions: a general machine could execute any formally describable computation; neurons could be represented as logical units; information and feedback could be measured; and electronic switching could make computation economically scalable. Wartime operations research added a fifth: mathematical models could improve consequential decisions inside a real organization.
Alan Turing’s 1936 work defined a general model of computation and also clarified that some problems are not computable. The inheritance for AI is double-edged. A programmable machine can implement an enormous range of algorithms, but computation has formal limits; scale does not turn every question into a solvable one. Turing’s computability paper
During the Second World War, interdisciplinary operations-research teams used probability, statistics, optimization and field evidence to improve radar deployment, logistics, search and resource allocation. This decision-science lineage later supplied forecasting, queuing, simulation, dynamic programming and mathematical optimization—the hidden machinery beneath many systems now labeled AI. It remains crucial because an LLM may interpret a goal, while a solver or controller calculates the allowable action. INFORMS history
In 1943, Warren McCulloch and Walter Pitts described simplified neurons as logical elements. Their abstraction was biologically limited, but historically decisive: networks of small computational units could collectively implement complex functions. Claude Shannon’s 1948 information theory then formalized information, entropy and channel capacity; Norbert Wiener’s cybernetics connected communication with feedback and control across animals and machines. These traditions supplied concepts that reappear in loss functions, representation learning, compression, reinforcement learning, robotics and system monitoring. McCulloch and Pitts, Shannon, Wiener
The physical prerequisite arrived at Bell Labs in 1947. The transistor offered a smaller, more reliable and more energy-efficient switching element than the vacuum tube. It would take integrated circuits, manufacturing scale and decades of design progress to reach modern processors, but the industrial direction was established: intelligence implemented in software ultimately depends on devices that switch and move information. Bell Labs history
Turing’s 1950 paper Computing Machinery and Intelligence reframed an abstract argument about whether a machine can think into an observable interaction. The imitation game was not a modern benchmark or a definition of intelligence; it was a practical way to discuss machine behavior while anticipating learning, language and common objections. That product insight—capability becomes socially consequential when people can interact with it—would recur with ChatGPT seventy years later. Turing’s 1950 paper
| Pre-AI foundation | What it contributed | Physical or institutional enabler | What survives in the 2031 system |
|---|---|---|---|
| Computability | a universal formal machine and explicit limits to mechanical procedure | mathematical logic and programmable electronic computers | general software, algorithm design and recognition that some objectives remain ill-defined or undecidable |
| Operations research | optimization, simulation and evidence-based allocation under constraints | wartime interdisciplinary teams and later government and corporate planning | schedulers, supply-chain models, resource allocation, control and deterministic tools inside agents |
| Mathematical neurons | networks of simple units as a model of computation | neuroscience, logic and early electronic circuits | neural-network abstraction, distributed representations and learned functions |
| Information theory | quantification of information, uncertainty, compression and channel limits | communications engineering and Bell Labs research | data coding, loss, compression, representation and bandwidth economics |
| Cybernetics | feedback, control and communication across biological and mechanical systems | sensors, servomechanisms and interdisciplinary systems research | reinforcement learning, robotics, monitoring and closed-loop agent behavior |
| The transistor | reliable electronic switching capable of industrial miniaturization | semiconductor physics, materials, fabrication and telecommunications demand | the manufacturing ancestry of CPUs, GPUs, HBM, networks and AI accelerators |
| The imitation game | an operational test of machine behavior through human interaction | stored-program computers, language and postwar debate about automation | conversational distribution, human evaluation and the continuing difference between useful behavior and underlying understanding |
This prologue does not move AI’s formal birth before Dartmouth. It explains why Dartmouth was possible. Earlier intellectual ancestors—probability, formal logic, Babbage and Lovelace’s programmable-machine ideas—belong in the ancestry, but a detailed history of mathematics would obscure this article’s focus: when ideas became a reproducible industrial capability.
1. 2010–November 2022: analytics, deep learning and foundation models become industrial
In 2010, enterprise AI usually meant analytics, forecasting, optimization or decision support. Companies built warehouses, ETL pipelines, reports and specialist teams so a model could improve a repeated decision. Outputs entered dashboards, campaign lists, planning tools or rules; people and deterministic software usually retained action authority.
2010–2015: big data industrializes analytics
Hadoop, Spark, NoSQL systems and cloud storage made web logs, transactions and sensor data more usable. Production models remained narrow—one target, dataset, workflow and owner—but created material value.
| Enterprise decision | Typical methods | Required operating system | Primary result |
|---|---|---|---|
| Demand and operations | time series, regression and optimization | governed history, planning integration and monitoring | lower stockouts, waste and overtime |
| Credit, fraud and insurance | rules, trees, anomaly detection and scoring | feature pipelines, case queues and validation | lower loss and faster decisions |
| Marketing and recommendation | propensity, clustering and collaborative filtering | identity, campaign systems and experiments | higher conversion and retention |
| Manufacturing and maintenance | sensor classification, survival models and optimization | plant connectivity, failure labels and engineering ownership | less downtime and scrap |
| Logistics and networks | routing, geospatial data and operations research | telemetry, solvers and dispatch integration | fewer miles and higher asset use |
The expensive element was usually not model training. It was data integration, software, validation, specialist labor and process change. IDC estimated narrow big-data technology and services revenue growing from less than $5B in 2010 to roughly $15–20B in 2015. IBM reported $18B of analytics revenue in 2015; GE reported roughly $500M of annual internal digital productivity savings; UPS expected ORION route optimization to save $300–400M annually. These company cases established industrial value before generative AI. European Commission/IDC, IBM, GE, UPS
Realization was uneven. McKinsey estimated that by 2016 U.S. retail had captured about 30–40% of previously identified analytics value, manufacturing 20–30%, and healthcare and government only 10–20%. Data quality, workflow ownership and management—not algorithm availability alone—separated leaders from laggards. McKinsey
2012–2022: learned representations become reusable
AlexNet’s 2012 ImageNet result joined a large dataset, benchmark, GPU and trainable convolutional network. Deep learning then spread through vision, speech, translation, search, advertising and recommendation. LSTMs handled sequences; embeddings learned reusable representations; GANs expanded generation; and reinforcement learning plus search produced systems such as AlphaGo. Digital platforms advanced fastest because they owned large interaction streams and rapid experimental feedback. AlexNet, GANs, AlphaGo
The 2017 transformer made sequence learning more parallel and scalable. Self-supervised training extracted supervision from raw text, images, audio and code; BERT and GPT shifted development from one model per task toward reusable pretraining. Scaling laws made larger compute programs legible to laboratories and capital markets, while Chinchilla, mixture-of-experts and retrieval showed that data allocation, sparse computation and external memory could improve the return on compute. Instruction tuning and human preference learning then converted a next-token predictor into a more cooperative assistant. Transformer, Scaling laws, Chinchilla, RAG, InstructGPT
| Transition | What became possible | Economic significance | Constraint carried forward |
|---|---|---|---|
| GPU deep learning | learned visual, speech and sequence representations | displaced hand-built pipelines in data-rich markets | labels, compute, robustness and domain shift |
| Transformer attention | parallel sequence modeling and flexible token interaction | made very large reusable models practical | context cost and memory |
| Self-supervised pretraining | learning from abundant unlabeled data | reduced task-specific labeling and widened reuse | source-data errors, rights and bias |
| Scaling and compute-optimal training | more predictable capability gains from compute and data | turned training capacity into a strategic investment | lower loss did not guarantee truth or value |
| Retrieval, sparsity and external tools | specialized memory and conditional computation | improved updateability and efficiency | routing, retrieval quality and system complexity |
| Instruction and preference learning | behavior aligned more closely with user requests | prepared foundation models for mass interaction | incomplete preferences, gaming and hidden failure |
Researchers and competing hypotheses
| Research program | Representative contributors | Central idea | Relevance through 2031 |
|---|---|---|---|
| Learned representations | Geoffrey Hinton, Yoshua Bengio and Yann LeCun | useful features can be learned by gradient-based systems | supports scaling while motivating robustness and interpretation |
| Memory and recurrence | Sepp Hochreiter, Jürgen Schmidhuber and recurrent-model researchers | systems need mechanisms that preserve relevant state | informs persistent memory and alternatives to rereading context |
| Generative and adversarial learning | Ian Goodfellow and collaborators | models can generate, while small changes expose brittleness | joins creative capability to security and authenticity |
| Benchmarks and data | Fei-Fei Li and the ImageNet community | shared datasets can organize research | makes provenance, contamination and evaluation strategic |
| Attention and scaling | transformer and foundation-model teams | broad pretraining creates reusable capability | underpins LLMs and their compute or context bottlenecks |
| Search, simulation and world models | Richard Sutton, David Silver, Demis Hassabis, Yann LeCun and others | planning requires feedback and models of consequences | informs reasoning, science and embodied agents |
| Causal, symbolic and hybrid methods | Judea Pearl and knowledge-representation researchers | correlation differs from intervention and explicit constraint | remains important in science and high-stakes decisions |
The point is not to choose one famous researcher as the winner. These programs solve different problems: representation, memory, action, causal reasoning, evaluation and control. The future system is likely to combine several of them.
China before ChatGPT
Alibaba, Baidu, Tencent, ByteDance and financial platforms already operated recommendation, search, advertising, logistics, risk and computer-vision systems at immense scale. Baidu’s ERNIE, the 2020 CPM paper and the 2021 PanGu-α paper established Chinese-language foundation-model research before the global consumer breakthrough. China had capability, investment and engineering depth; it had not yet produced one global product moment that compressed public understanding and commercial urgency.
Traditional analytics versus the new value pool
| Economic stage | Dominant scope | Investment or supplier signal | Annual net direct benefit: article planning range |
|---|---|---|---|
| 2010–2015 | reporting, forecasting, scoring and optimization | big-data technology and services < $5B → ~$15–20B | ~$0.2–$0.4T |
| 2020–2022 | scaled recommendations, risk, vision, language and platform optimization | cloud about $225B in 2022; four-company broad capex about $150B | ~$0.5–$1.0T |
| 2025–2026 | traditional AI plus language, code, multimodality and early agents | cloud about $500B trailing revenue; four-company 2026 plans $735–760B | ~$1.0–$1.5T |
| 2030 | larger shares of repeatable digital work under tools and verification | infrastructure shifts toward continuous inference | ~$2.4T |
| 2031 | traditional models, foundation models and bounded agents operate together | $6.4T enabled → $4.1T captured before cost → $0.7T program cost | $3.4T |
These ranges are transparent article reconstructions, not audited global profit. The direction matters: traditional analytics improved a limited set of structured decisions; generative and agentic systems add unstructured knowledge, software and coordination. Traditional ML remains inside the new stack for forecasting, recommendation, tabular risk and optimization, while LLMs interpret intent and coordinate tools.
ChatGPT’s discontinuity came from composition: a scaled model, post-training, conversation, low-friction access and immediate feedback. Once executives perceived a platform transition, capital flowed into accelerators, clouds, laboratories, data centers and power. The central question moved from whether models could generate useful output to whether the industrial system could supply and organizations could absorb reliable intelligence.
2. December 2022–2026: mass distribution, multimodality and agents
ChatGPT joined a capable model to conversation, freemium access and rapid feedback. It sharply reduced the cost of trying AI and turned a technical frontier into a mass-market habit. The resulting demand signal pulled capital into models, chips, cloud and applications.
From product launch to global habit
| Date | Scale or event | What changed |
|---|---|---|
| November 2022 | ChatGPT launched publicly | one general interface replaced many use-case-specific demonstrations |
| Late 2023 | ChatGPT exceeded 100M weekly users | AI became a consumer habit within a year. NBER |
| 2025 | ChatGPT reached 700M weekly users; Meta AI exceeded 1B monthly users; China reached 602M GenAI users | direct destinations, embedded platforms and national ecosystems all achieved mass scale |
| May–November 2025 | Claude Code reached a company-reported $1B revenue run rate within six months | task execution became a standalone paid category. Anthropic |
| March 2026 | ChatGPT exceeded 900M weekly users and 50M subscribers | OpenAI combined reach with a large direct paid relationship. OpenAI |
| Q2 2026 | Gemini reached 950M monthly app users; Google reported more than 9M monthly developers across key model and developer products | Google joined models to search, devices, work, cloud and coding. Alphabet |
These figures are not market shares. Weekly users, monthly users, embedded assistants, subscribers, developers and API calls have different denominators. The defensible conclusion is that ChatGPT established the leading direct AI destination while Google, Meta and Chinese platforms converted existing distribution into enormous reach.
Distribution became part of capability. Usage supplied product feedback; traffic pulled forward infrastructure; APIs allowed startups and business units to experiment; and default placement inside search, work, devices and messaging became a strategic advantage. Assistants may increasingly sit upstream of websites, merchants and publishers, redirecting advertising and transaction value even while creating consumer surplus.
Coding agents: the first mass market for agentic work
Coding moved first because software provides machine-readable context, tools, tests, version control and rollback. The product evolved from autocomplete to chat, then to repository agents that plan, edit, run commands, inspect errors and return reviewable changes. The innovation is the harness around the model—not code generation alone.
| Provider and coding surface | Distribution and core capability | Current signal | Strategic reading |
|---|---|---|---|
| OpenAI — ChatGPT and Codex | app, terminal, editor and cloud; repository work, testing, review and parallel tasks | no comparable standalone Codex revenue or user disclosure | connects OpenAI’s consumer reach to execution; value depends on verified changes. OpenAI Developers |
| Anthropic — Claude and Claude Code | terminal, app, web and clouds; planning, multi-file work, tests, subagents and tools | $1B run rate within six months | strongest disclosed paid coding-agent signal; company revenue is not customer profit. Claude Code |
| Google — Gemini developer tools and Antigravity | IDE, terminal, cloud and agent platform; large-context analysis and isolated execution | more than 9M monthly developers across key products; 2.4M Antigravity weekly users | distribution and custom silicon support scale; metrics span several products. Alphabet |
| GitHub Copilot and AI-native platforms | repository and editor distribution; completion, review and issue-to-change workflows | studies range from 55% faster on one task to 26% more completed tasks across a larger field sample | installed workflow position is powerful, but completion and autonomous agents are different products. GitHub, Management Science |
| Meta — Llama and Code Llama | downloadable weights and partner clouds | no central coding-agent audience or revenue | expands local control and third-party products; users assume operating burden. Meta |
| xAI — Grok coding models and Grok Build | X, APIs and a terminal agent; fast generation and parallel subagents | no comparable standalone coding revenue or audience | latency helps tool loops, but speed must be measured at verified completion. xAI |
| Alibaba — Qwen3-Coder and Qwen Code | cloud, downloadable weights and an open terminal agent | ecosystem adoption rather than comparable revenue | lowers the cost of a capable, substitutable coding stack. Qwen |
| DeepSeek, GLM and other Chinese open models | low-cost APIs, downloadable models and third-party harness compatibility | strong vendor-reported code and agent benchmarks | their largest effect may be price competition and sovereign model choice. DeepSeek, GLM |
The evidence is deliberately mixed. Anthropic reports Rakuten reducing one feature cycle from 24 working days to five; Google reports a major internal refactor compressing a two-year schedule to three months. Yet METR found experienced maintainers 19% slower on a different set of familiar-repository tasks, and 46% of Stack Overflow respondents distrusted AI-output accuracy. The relevant metric is time and total cost per accepted, verified change, including inference, tests, review, failures and rework. Rakuten, METR, Stack Overflow
Three shifts after ChatGPT
- Models became multimodal, post-trained systems. Text expanded to images, audio, video and interfaces; preference training and safety operations shaped product behavior.
- Compute moved into inference. Reasoning, search, tool calls and verifiers made the cost of one task variable, while batching, caching, quantization and routing improved utilization.
- Answers became actions. Agents connected models to data and tools, but errors, prompt injection and excess authority made identity, permissions, tests and rollback essential.
Open weights from Llama, Mistral, Qwen, DeepSeek, GLM and Kimi made capable models cheaper and more portable, although deployment still requires compute, security, evaluation and support. By 2031, one user task may route across several closed, open, specialist and traditional models. The durable control point is therefore likely to be the system holding consented context, identity, tools, evaluation and the customer relationship.
The post-ChatGPT race is no longer only to train the strongest model. It is to assemble the most reliable and economical system for completing useful work.
3. The AI factory: chips, data centers, power and water
AI appears as software but is produced by a physical chain: semiconductor tools, leading-edge logic, HBM, packaging, networks, servers, cooling, power, water and cloud operations. A shortage or delay at one qualified link can strand investment everywhere else.
From cloud expansion to an industrial buildout
| Physical signal | Before or near ChatGPT | Current state | 2030–2031 direction | Interpretation |
|---|---|---|---|---|
| Global semiconductor sales | ~$440B in 2020; $574.1B in 2022 | $791.7B in 2025 | no comparable primary point used | AI changed the mix toward accelerators, HBM, packaging and networking faster than total-chip revenue. SIA |
| Leading foundry investment | TSMC $17.2B capex in 2020 | $40.9B in 2025; $52–56B 2026 guidance | multi-site expansion continues | tools, yield learning and advanced packaging matter as much as the factory shell. TSMC |
| Data-center investment | cloud-led expansion | about $500B in 2024, nearly double 2022 | Morgan Stanley: $2.89T cumulative 2025–2028 | bank scope overlaps hyperscaler capex and includes non-AI facilities. Morgan Stanley |
| Data-center electricity | 240–340 TWh in 2022 | 415 TWh in 2024; 485 TWh in 2025 | IEA about 950 TWh in 2030 | all data centers, not AI alone; deliverable power can lag announced demand. IEA |
| Facility concentration | enterprise and cloud capacity coexisted | 1,360 hyperscale sites; 48% of capacity at end-2025 | Synergy: 67% of capacity in 2031 | ownership continues moving toward hyperscalers and specialist providers. Synergy |
A concentrated global supply chain
| Location | Strategic role | Main constraint |
|---|---|---|
| United States | accelerator design, frontier laboratories, cloud demand, software and capital | leading fabrication and HBM remain international; power queues delay campuses |
| Taiwan | leading-edge foundry production and advanced packaging | geographic concentration and difficult-to-reproduce process knowledge |
| South Korea | HBM and advanced memory | yield, packaging, power and customer concentration |
| Netherlands, Japan and the United States | lithography, deposition, etch, materials, metrology and test | few substitutes, long qualification and export controls |
| China | large demand, cloud, open models, electronics and domestic substitution | restricted access to frontier accelerators, tools and HBM |
| Europe, India and the Gulf | industrial deployment, regional compute, talent, power and sovereign programs | utilization, equipment access and dependence on external clouds or models |
No country needs to reproduce every layer. Resilience comes from competitive model and cloud access, enough governable compute for critical work, reliable networks and power, technical talent and the ability to switch suppliers.
The bottleneck moves
Three facts organize the factory.
- Complete systems matter more than chip counts. Accelerators can wait on HBM, packaging, optical networks, switchgear, cooling or grid connections. Revenue flows first to whichever qualified component cannot expand quickly.
- Training and inference need different machines. Training emphasizes synchronized cluster throughput; interactive inference emphasizes latency, memory, batching and reliability; edge inference emphasizes privacy and low power.
- Software determines usable capacity. Compilers, kernels, distributed systems, caching, quantization and routing convert theoretical performance into cost per completed task.
The production stack remains heterogeneous: CPUs orchestrate; GPUs and custom accelerators perform parallel work; HBM feeds them; packaging and networks connect them; serving software schedules them; identity, retrieval and verification determine whether output becomes an authorized action. The economically useful unit is not a transistor, GPU or token. It is a reliable task completed by the whole system.
Power, water and materials
Power is permission to operate. A campus can be built faster than new generation or transmission, and the IEA estimates that 20% of planned data-center projects could face delays without grid action. Efficiency lowers energy per fixed task, but reasoning, video and agent loops expand demand. Operators should report verified tasks per dollar, watt and liter alongside absolute resource use.
Annual renewable matching does not prove carbon-free supply in every hour. Credible claims identify local physical supply, firming, additional generation and attributable grid cost. Large loads should fund the infrastructure they require and use flexible workloads where service and data rules permit.
Water is local. LBNL estimates about 66B liters of direct U.S. data-center consumption in 2023 and 800B liters indirectly associated with electricity, with direct use potentially reaching 150–280B liters in 2028. Withdrawal, consumption, indirect use and replenishment are different measures; disclosure should be facility- and basin-specific. LBNL
The environmental ledger also includes steel, concrete, copper, specialty chemicals, semiconductor water and frequent server replacement. Buildings, substations and fiber can serve several accelerator generations, so adaptable sites and equipment reuse reduce both capital and embodied impact.
The 2031 physical architecture
The likely system combines frontier clusters, custom silicon for stable workloads, smaller edge models and schedulers that route by quality, latency, price, residency and resource conditions. Batch work can sometimes shift in time or geography; real-time services and synchronized training cannot.
The IEA’s electricity outlook, Morgan Stanley’s investment estimate and Synergy’s capacity forecast answer different questions and must not be added. Together they show that advantage belongs to the operator that converts chips, energy, water and networks into utilized, reliable and socially permitted intelligence.
4. The model and cloud economy: laboratories, training and inference
Model laboratories convert compute into reusable capability; clouds meter and distribute capacity; applications combine models with data, tools and customers. These layers have different costs and should not be evaluated with one revenue or margin model.
From data to a production service
| Stage | Principal work | Main economic exposure |
|---|---|---|
| Data and preparation | rights, collection, filtering, provenance and reproducible datasets | licensing, privacy, creator compensation and scarce authentic data |
| Pretraining and research | experiments, large fleet optimization, checkpoints and fault recovery | cluster time, failed runs, energy and concentrated talent |
| Post-training and evaluation | instructions, reasoning, tools, safety, red teaming and deployment gates | expert feedback, inference-heavy experimentation, incidents and delayed release |
| Serving and routing | quantization, batching, caching, memory management and model selection | utilization, HBM, latency, peak capacity and price compression |
| Application and renewal | retrieval, state, identity, permissions, monitoring and workflow change | integration, security, review, switching and recurring operations |
A final training run is only one cost. Epoch AI estimates roughly $4.5M for GPT-3, $12.4M for PaLM and about $390M for the largest public 2024 estimates. It also describes a path toward $100B-plus frontier clusters by 2030. A laboratory can therefore report a training run in the hundreds of millions while relying on a many-billion-dollar reusable industrial platform. Epoch AI
Capital structures now mix equity, cloud credits, compute reservations and commercial agreements. The key question is who owns costly capacity and who owns the customer. Asset-light laboratories avoid construction and depreciation but pay a provider margin; integrated clouds capture more of the stack but must keep short-lived chips and long-lived sites utilized.
Inference and cloud become the recurring economy
Inference is continuous. Cost depends on model size, context, reasoning steps, modalities, caching, memory bandwidth, retries and service guarantees. The price of GPT-3.5-level capability fell from about $20 to $0.07 per million tokens between late 2022 and late 2024, but longer context, video, search and agents increased consumption. The useful metric is successful, verified tasks per dollar, joule and minute, not tokens alone.
Cloud infrastructure revenue grew from $129B in 2020 to $227B in 2022 and about $500B for the trailing twelve months through Q2 2026. Hyperscalers combine diversified demand, proprietary chips and product distribution; neoclouds specialize in accelerator operations; sovereign providers sell jurisdiction and continuity; enterprise or edge deployments serve private, high-volume or latency-sensitive work. Their common risk is paying for capacity before utilization arrives.
Routing is likely to become a durable control point. It can assign routine work to small or local models, difficult planning to frontier systems and exact calculations to deterministic tools, while preserving evidence and budget. If routing and data remain portable, buyers retain leverage; if hidden inside one platform, lock-in deepens.
Closed models, open weights and China’s alternative stack
Public evaluations show meaningful convergence, but benchmark proximity is not production parity. Buyers must also test uptime, tools, security, capacity, license, support and total cost per successful task.
| Family or ecosystem | Distribution and agentic position | Strategic strength | Main constraint |
|---|---|---|---|
| OpenAI — GPT, ChatGPT and Codex | leading direct assistant plus enterprise API and first-party coding agent | consumer reach, reasoning, multimodality and execution | price, policy dependence, data control and undisclosed Codex economics |
| Anthropic — Claude and Claude Code | enterprise and cloud distribution with a strong repository agent | coding, long-context work, tool use and safety engineering | compute dependence, review burden and run-rate revenue that does not prove customer return |
| Google — Gemini and developer agents | search, devices, work, cloud, IDE and terminal distribution | embedded reach, multimodality, TPUs and developer scale | ecosystem lock-in and metrics spanning several products |
| xAI — Grok and coding tools | X, subscriptions, API and terminal agent | real-time distribution, fast coding and rapid infrastructure scale | maturity, governance, capital intensity and undisclosed standalone economics |
| Meta — Meta AI and Llama | more than a billion embedded users plus downloadable weights | distribution, local adaptation and a large open ecosystem | no comparable first-party repository agent; license and operating burden |
| Mistral and European alternatives | hosted and open-weight regional options | efficiency, multilingual support and supplier diversity | smaller compute, distribution and developer-harness scale |
| Qwen | Alibaba Cloud, broad open weights and Qwen Code | multilingual breadth, coding, tool use and global derivatives | fragmented usage, support and license or version management |
| DeepSeek | low-cost service, open weights and agent-harness compatibility | price-performance, reasoning and model substitution | service reliability, hardware access, governance and support |
| Kimi and GLM | Chinese hosted and open-weight models with domestic ecosystems | long context, reasoning, coding and agentic engineering | uneven international distribution, security review and support |
Chinese open models matter because they widen supplier choice, compress prices and show how algorithmic efficiency can offset some hardware constraints. They also shift competition toward deployment engineering, local language, evaluation and integration. “Open weight” remains the accurate term when code, data, training recipe or license rights are incomplete. OSI, OECD
There may be no single most-used model in 2031. One family can lead conversation, another embedded search, another enterprise tokens and many open models can run without central measurement. The likely winner is the organization that can route among them without losing its data, permissions, evaluations or customer relationship.
5. What current AI cannot reliably do—and the research attempting to fix it
Current systems are useful but probabilistic. They can write, code, search and call tools, yet still invent facts, miss constraints, lose state and execute the wrong action. Longer context and more test-time compute improve many tasks; neither guarantees truth, memory or safe autonomy.
The practical question is not whether every limitation disappears. It is whether better algorithms, memory systems, chips and controls can lower the cost of a verified outcome enough for wider deployment.
Five limitations that shape the 2031 market
| Limitation | Production consequence | Most credible response | 2031 outlook | Economic implication |
|---|---|---|---|---|
| Grounding and confabulation | fluent but unsupported answers create legal, clinical and operational risk | retrieval, citations, calibrated uncertainty, deterministic checks and human approval | material reduction likely; elimination unlikely in open-ended work | evidence and assurance become valuable product layers. TruthfulQA, NIST |
| Uneven reasoning and planning | performance drops on unfamiliar, long-horizon or adversarial tasks | test-time search, formal tools, process supervision, specialist verifiers and smaller bounded plans | strong gains likely on checkable tasks; broad reliability remains uncertain | autonomy expands first where success can be tested cheaply. OpenAI reasoning research, METR |
| Context without durable memory | long prompts raise cost and still lose chronology, identity or relevant facts | structured state, retrieval, recurrent memory, compression and user-controlled records | hybrid memory is likely; perfect recall is neither feasible nor desirable | owners of governed context and systems of record retain leverage. Lost in the Middle |
| Agent and code verification debt | tool calls compound errors; generated changes can overwhelm reviewers | least privilege, sandboxes, tests, budgets, checkpoints, rollback and independent review | bounded digital agency should scale faster than high-stakes physical autonomy | value moves from code volume to cost per accepted, verified change. SWE-bench, Anthropic sandboxing |
| Data, security and resource intensity | synthetic feedback, prompt injection, privacy loss and inference rebound constrain scale | provenance, authenticated data, secure tool boundaries, efficient models, routing and hardware–software co-design | efficiency should improve sharply, but total demand may still rise | trusted data, cybersecurity, memory bandwidth and energy remain strategic bottlenecks. Model collapse, OWASP |
Three research and deployment priorities
- Verify the outcome, not the explanation. Use tests, retrieved evidence, formal rules and sampled human review that are independent of the generating model.
- Give agents limited authority. Identity, short-lived credentials, action budgets, checkpoints and rollback should expand only after measured reliability improves.
- Optimize the complete system. Route routine work to smaller models, preserve structured state and measure dollars, joules, latency and recovery cost per successful task.
The largest open risk is a verification bottleneck: generation becomes abundant while trustworthy review remains scarce. The largest opportunity is the same market in reverse. Organizations that combine strong models with evidence, secure tools and recoverable workflows can turn probabilistic capability into dependable service.
6. Adoption becomes application: consumers, enterprises, industries and government
Mass adoption is not the same as economic absorption. Consumers can obtain value from a useful answer immediately; companies and governments must connect models to data, permissions, workflows and accountability. The gap between those two speeds explains why user growth is spectacular while enterprise return remains uneven.
Breadth rose faster than depth
| Adoption signal | Current scale | What it means |
|---|---|---|
| ChatGPT | more than 900M weekly users and 50M subscribers in 2026 | direct global distribution; not comparable with monthly app or embedded-product metrics. OpenAI |
| Gemini and China | Gemini 950M monthly app users; China 602M GenAI users | large ecosystems can distribute AI through search, devices, cloud and domestic platforms. Alphabet, China State Council |
| Formal businesses | OECD adoption rose from 8.7% in 2023 to 20.2% in 2025; large firms reached 52% | diffusion is real but concentrated by firm size. OECD |
| Enterprise programs | 88% of surveyed organizations used AI somewhere, but about one-third had begun scaling agents | executive breadth exceeds production depth. McKinsey |
| Government inventories | eleven U.S. agencies reported 1,110 AI cases in 2024, including 282 GenAI cases | experimentation is growing faster than mature public deployment. GAO |
Three findings survive across the evidence. First, bounded work can improve materially: customer-support agents produced 13.8% more resolutions per hour, and three developer experiments found about 26% more completed tasks. Second, task boundaries matter: consultants improved inside the model’s capability frontier but performed worse outside it, while experienced maintainers in one METR study were 19% slower with early coding tools. Third, firmwide return requires complementary software, data, training and process redesign; survey evidence shows many pilots still fail to scale. NBER, Harvard Business School, Management Science, IBM
Applications and investment economics
The 2031 base case is one waterfall: $6.42T annual value enabled → $4.10T captured before cost → $700B program cost → $3.40T net direct benefit. The figures are an article scenario, not vendor revenue or GDP. They extend the economic architecture developed in Agentic AI: Transforming Industries, From Better Answers to New Industry Economics and Agentic AI: $16Tn Enabled.
| Application system | Highest-value uses | 2031 annual net direct benefit: article scenario | Main condition for capture |
|---|---|---|---|
| Software, professional services and customer operations | coding, modernization, research, service resolution and managed workflows | $900B | verified completion, outcome pricing and redesigned delivery |
| Finance, payments and insurance | research, compliance, fraud, onboarding, claims and risk operations | $420B | regulated authority, model-risk controls and integration |
| Healthcare, education and government | documentation, navigation, tutoring, case preparation and public-service capacity | $450B | evidence, privacy, appeal, workforce adoption and access |
| Commerce, logistics, agriculture and real estate | discovery, forecasting, procurement, routing and exception resolution | $830B | channel access, physical execution and reliable data |
| Energy, industry and cross-sector spillovers | engineering, maintenance, grid planning, science and constrained robotics | $800B | sensors, safety, capital cycles and simulation-to-reality performance |
| Global base case | value after adoption, realization and cost filters | $3.40T | workflow conversion, competition and institutional capacity |
Three application and investment rules
- Fund an outcome, not a feature. The strongest application owns a measurable workflow, regulated permission, system of record or customer relationship. A thin interface is vulnerable to model and platform bundling.
- Price the full operating system. Include inference, integration, review, failure recovery, security and organizational transition. Time saved becomes profit only when output, capacity, price or loss changes.
- Scale in stages. Begin with high-volume, reversible and checkable work; expand authority only when cost, quality and intervention rates remain acceptable.
Coding agents are the clearest reference market because software provides repositories, tests, version control and rollback. Claude Code’s company-reported $1B run rate shows willingness to pay, while Codex, Gemini, Grok, GitHub and open-model tools keep the category competitive. The appropriate metric is not generated code but time and total cost per accepted, verified change. Anthropic, OpenAI Developers
Government has three priority application groups: administration and revenue such as tax, procurement and benefits; human services such as health, education and workforce support; and infrastructure and security such as grids, emergency response and cyber defense. In every group, the return is reliable public capacity per dollar, while legal authority, adverse decisions and appeals remain human and institutional responsibilities.
The transition is therefore from seats and prompts to verified outcomes. Leaders should track completion, cycle time, quality, intervention, loss and resource use. Adoption is an input; retained value is the result.
7. The money flow: capex, venture capital, revenues, valuations and returns
AI finance follows a sequence: funding → equipment orders → installed capacity → usage → productivity → cash return. Suppliers are paid before buyers prove value, and the same dollar can appear as customer capex, supplier revenue and cloud cost. Those objects must not be added.
The capex reset
| Date and status | Amazon, Alphabet, Meta and Microsoft broad infrastructure signal | Interpretation |
|---|---|---|
| 2020 observed | ~$95B | cloud and platform expansion before the mass-assistant market |
| 2022 observed | $151.1B | substantial pre-ChatGPT base; broader than AI |
| 2025 observed | ~$360B | accelerators, servers, networks and data centers drove the increment |
| 2026 plans | $735–760B | Amazon about $220B; Alphabet $195–205B; Meta $130–145B; Microsoft about $190B |
| 2031 external scenario | Goldman: $1.6T annual and $7.6T cumulative, 2026–2031 | supply-side buildout scenario, not guaranteed demand |
Definitions and fiscal periods differ, and Amazon includes logistics and other assets. The direction is nevertheless clear: post-ChatGPT capital formation entered a different regime. Company filings and Goldman Sachs
Where returns appear
| Layer | Current financial signal | Return test |
|---|---|---|
| Chips, memory, foundries and equipment | NVIDIA data-center revenue reached $194B in FY2026 | demand, pricing and useful life must survive competition and chip cycles. NVIDIA |
| Cloud, data centers and power | cloud infrastructure reached about $500B trailing revenue; multi-trillion buildout scenarios are being financed | utilization and cash yield per energized megawatt must exceed depreciation and financing |
| Frontier and coding-agent platforms | enormous laboratory financing; Claude Code reported a $1B run rate | revenue and gross margin must outrun training, inference, distribution and security cost |
| Applications and AI-native services | enterprise application revenue is growing across coding, horizontal and vertical products | retention and outcome value must survive model substitution and platform bundling |
| Adopters, workers and governments | task gains exist, but enterprise return and pass-through remain uneven | productivity must become profit, wages, lower prices or better public services |
Three conclusions from venture research
| Venture conclusion | Evidence | Investment implication |
|---|---|---|
| Capital is concentrated upstream | OECD estimated $258.7B of global AI VC in 2025; about 75% went to deals above $100M, and infrastructure or hosting attracted about $110B | treat frontier labs and infrastructure as capital-intensive platforms with long payback, not ordinary software. OECD |
| Application revenue is becoming material | Menlo estimated U.S. enterprise GenAI spending rising from $11.5B in 2024 to $37B in 2025, with $19B at the application layer | back products that own distribution, workflow data, permission or an accountable result. Menlo Ventures |
| Fast growth can hide weak economics | Bessemer found some high-growth AI companies scaling quickly with roughly 25% gross margins, while more durable peers were nearer 60% | judge contribution margin after inference, implementation, review, sales and failure cost—not headline recurring revenue. Bessemer |
Andreessen Horowitz’s payment data, Sequoia’s return questions, General Catalyst’s service thesis, Coatue’s cross-stack view and Matrix’s infrastructure portfolio reinforce those conclusions from different vantage points. They are useful directional evidence, not neutral market measurement. a16z, Sequoia, General Catalyst, Coatue, Matrix
Three investment rules follow. First, match financing to asset life: accelerators, buildings and power infrastructure have different depreciation and refinancing risks. Second, require evidence of utilization and customer return before treating bookings, capacity announcements or valuation marks as demand. Third, prefer durable control points—energized sites, proprietary data, trusted distribution, regulated permission, systems of action and verified outcomes—over reproducible features.
Private valuations are claims on future cash flow, not present economic value. Paper revaluations, preferred terms and strategic compute agreements can obscure operating performance. The 2031 test is simple: does incremental gross profit and free cash flow justify capex, leases, research, inference and transition cost?
Financial verdict: Capital is committed upstream, revenue appears first at bottlenecks, and durable returns are decided downstream.
8. National and human consequences: sovereignty, work, security and trust
AI use is global, but ownership of chips, clouds, models and capital is concentrated. Capability can therefore diffuse faster than profit or policy power. The United States leads frontier laboratories, accelerator design, clouds and venture finance; China combines a vast domestic market, engineering depth, manufacturing and influential open-weight models. Both remain dependent on an international supply chain spanning Taiwan, Korea, Japan and Europe.
National strategy and local value capture
China reported 602M GenAI users, more than 6,000 AI enterprises and a core industry above RMB1.2T in 2025. U.S.-based firms received about 75% of reported global AI venture capital in the OECD dataset. These figures use different definitions, but they show two different systems: U.S. private capital and frontier platforms versus Chinese deployment scale, industrial policy and efficient-model competition.
For countries without a frontier laboratory, sovereignty should mean continuity, not prestige. The World Bank’s four priorities—connectivity, compute, context and competency—are a stronger foundation than one national chatbot. Add multi-provider procurement and an ability to evaluate, secure and switch critical systems. World Bank
| Economy or region | 2031 annual net direct benefit: article scenario | Main condition for local capture |
|---|---|---|
| United States | $1.30T | turn capex into broad business productivity while preserving competition and workforce mobility |
| China | $600B | improve domestic silicon and tools while sustaining efficient models, private demand and trusted exports |
| European Union and United Kingdom | $600B | integrate compute and procurement, diffuse into industry and retain research-linked scaleups |
| India and other Asia-Pacific | $600B | combine shared compute, talent, local-language applications and manufacturing strength |
| Rest of world | $300B | affordable access, regional procurement, skills and locally owned application ecosystems |
| Global base case | $3.40T | workflow conversion, competition and institutional capacity |
This allocation is not a GDP forecast. Value created locally can still be captured by a foreign chip, cloud or model provider. Governments should test whether infrastructure projects create reliable power, tax revenue, skills and supplier capacity rather than only exporting subsidized land, water and electricity.
Three public functions
| Public priority | Appropriate AI role | Measure of public value | Control that must remain |
|---|---|---|---|
| Administration, tax, procurement and benefits | evidence collection, correspondence, case preparation, compliance and bounded investigation | cycle time, accuracy, eligible uptake, competition and fraud avoided | legal determination, adverse-action notice and appeal |
| Health, education and workforce services | documentation, navigation, translation, tutoring and coordination | access, waiting time, learning, mobility and professional capacity | clinical, curricular, safeguarding and placement authority |
| Infrastructure, emergency and cyber resilience | forecasting, inspection, situation synthesis, maintenance and threat analysis | avoided outages, recovery speed, safety and detection quality | command authority, physical controls and lawful authorization |
Public procurement should reuse secure components, publish outcome measures and preserve portable data. More use cases are not success: the U.S. agency inventory rose from 571 to 1,110 between 2023 and 2024, while only a fraction had moved beyond early stages. Rights, security and recourse are part of the return, not obstacles outside it. GAO
Work will reorganize around tasks
Evidence supports broad exposure but not a single job-loss forecast. The ILO finds one in four jobs has some GenAI exposure and expects transformation to be more common than full replacement. The World Economic Forum’s employer survey expects large gross job creation and displacement across all modeled trends; it is not a realized AI employment count. Early studies also show uneven pressure on younger and more exposed workers. ILO, WEF
| Labor transition | Tasks under pressure | Work likely to grow | Most important response |
|---|---|---|---|
| Routine information and transaction work | drafting, search, data entry, reconciliation, routing and standard support | exceptions, relationship recovery, quality control and records governance | redeployment pathways, domain judgment and outcome-based service models |
| Software and professional work | routine coding, testing, research, diligence and document production | architecture, product judgment, security, evaluation, negotiation and accountable advice | preserve apprenticeship, independent competence and human review |
| Care, education and public service | documentation, navigation, preparation and routine monitoring | complex judgment, safeguarding, coaching and human engagement | use saved time to improve access and job quality, not only work intensity |
| Physical AI and assurance | selected inspection, planning and structured machine tasks | electricians, chip and grid specialists, robot maintainers, data stewards and incident investigators | paid vocational pathways, recognized standards and geographic mobility |
Coding agents make the transition visible first. Engineers spend less time typing every implementation and more time specifying intent, preparing context, designing tests, reviewing changes and accepting production responsibility. Demand may expand because lower software cost makes more products and modernization viable, but routine implementation and labor-billed services face pressure. The immediate risk is apprenticeship: if firms remove junior tasks and hiring, they weaken the future supply of reviewers and architects. Supervised practice, debugging, incident rotations and independent assessment should remain part of the job.
The human supply chain also includes data creators, annotators, safety raters, moderators and domain reviewers. Provenance, consent, fair working conditions and governed access to authentic scientific, operational and local-language data are economic inputs. Synthetic data helps when simulation or verification supplies ground truth; it cannot replace observation indefinitely.
Access, security and trust
Export controls and provider safeguards restrict some chips, models and uses because the same systems can support cyber defense, vulnerability discovery, science or harmful operations. A credible regime combines customer screening, secure weights and compute, least-privilege tools, monitoring, incident reporting and human authorization for consequential actions. It should also avoid turning safety into unreviewable commercial gatekeeping.
Organizations should design for capability continuity: portable data and evaluations, more than one viable model route, locally controlled identity and logs, and fallback procedures for provider or policy changes. High-impact public and commercial decisions also require notice, evidence, human authority and appeal. Trust is economic infrastructure because institutions will not delegate important work to systems they cannot challenge or recover.
The distribution test is therefore concrete: who owns the scarce assets, who receives the productivity gain, whether consumers obtain lower prices or better services, whether workers retain mobility, and whether countries can switch providers. Technology creates the opportunity; institutions distribute it.
9. 2027–2031: four futures and one operating agenda
The next five years depend on several variables moving together: model reliability, inference efficiency, energized capacity, capital cost, workflow conversion, labor response and public legitimacy. The following are conditional futures, not precise predictions.
Four plausible intelligence-age systems
| Future | What happens | Main beneficiaries | Primary risk | 2031 economic outcome |
|---|---|---|---|---|
| Managed diffusion — base case | plural models, efficient routing and bounded agents spread as infrastructure and institutions adapt | productive adopters, vertical applications, skilled workers and regions with usable capacity | uneven diffusion and slow organizational change | article base case of $3.4T annual net direct benefit remains feasible |
| Agentic productivity boom | planning, verification and tool use improve quickly; software, research and selected physical workflows are redesigned | customers, AI-native firms, science and complementary labor | transition, safety and apprenticeship lag capability | $4.2–5.0T net with stronger new demand |
| Concentrated platform abundance | intelligence becomes cheaper but a few model-cloud-distribution groups control context, identity and action | hyperscalers, frontier labs, chip suppliers and distribution owners | high switching cost and weaker pass-through to wages, prices and independent firms | $3.5–4.5T net, with more platform rent |
| Fragmented bottleneck and backlash | power delays, financing stress, incidents and geopolitical controls divide the market | secure domestic suppliers, scarce powered sites and cyber or compliance providers | duplicated stacks, write-downs, weaker hiring and capability tiers | $1.5–2.5T net |
The planning envelope
| Indicator | Current base | Managed diffusion | Upside | Downside signal |
|---|---|---|---|---|
| Formal businesses using AI | OECD 20.2% in 2025 | 40–60% by 2031 | 55–70% | below 45%, concentrated in large firms |
| Repeatable information workflows handled by bounded agents | production depth remains low | 20–40% | 35–55% | below 25%, with high intervention or failure |
| Global data-center electricity | 485 TWh in 2025 | around IEA’s 950 TWh 2030 path | approaches or exceeds 1,200 TWh without stronger efficiency | severe local scarcity, delays and duplicated capacity |
| Annual AI ecosystem capex | Goldman scenario $765B in 2026 | toward $1.6T in 2031 if utilization validates expansion | larger path supported by downstream revenue | cancellations, weak free cash flow and refinancing stress |
| Annual net direct benefit | early and uneven | $3.4T | $4.2–5.0T | $1.5–2.5T |
A simpler production architecture
The likely system is a governed portfolio rather than one universal model.
| Layer | Function | Economic purpose | Minimum control |
|---|---|---|---|
| Compute and model routing | assign work across devices, clouds, closed, open and specialist models | lower cost and preserve supplier choice | metering, portable evaluations, residency and budget rules |
| Context, memory and tools | retrieve approved data, preserve state and expose bounded actions | turn general models into organization-specific capability | provenance, consent, tool allowlists and injection defenses |
| Planning and execution | coding and workflow agents decompose goals, call tools, test and recover | convert prompts into completed work and outcome pricing | identity, least privilege, checkpoints, sandboxes and rollback |
| Verification and human command | test claims and completed state; approve consequential exceptions | reduce loss and preserve accountability | independent tests, audit logs, escalation, notice and appeal |
MCP, A2A and NIST’s agent initiative indicate a move toward interoperable tools, agent communication, identity and evaluation. The standards are early, but the direction is clear: context and authority need explicit control outside the model. MCP, A2A, NIST
Three actions through 2031
- Prove workflow conversion. Enterprises and governments should select a few measurable processes, track cost and quality per completed task, and expand autonomy only after evidence.
- Underwrite utilization, not announcements. Investors should connect chips, buildings, power and leases to energized demand, cash yield and appropriate asset lives.
- Build diffusion and exit rights. Governments, employers and buyers should support skills, apprenticeships, shared access, portable data, plural suppliers, independent evaluation and recourse.
Three indicators reveal which future is emerging: AI cash yield, or incremental gross profit and free cash flow relative to commitments; workflow conversion, or the share of important processes reliably completed; and diffusion quality, or whether smaller firms, workers, public institutions and countries without frontier labs can access and retain value.
The robust strategy does not require certainty about model progress. Modular systems, flexible infrastructure, trained people, independent verification and the ability to switch suppliers remain valuable across every scenario.
10. Beyond 2031: six questions that will define the 2035 intelligence economy
2031 is the article’s decision horizon because today’s fabs, data centers, power contracts, model investments and workforce programs can be assessed against it. Extending precise market-share or valuation forecasts to 2035 would create false confidence. A longer horizon is still useful if it is framed as questions, alternative paths and observable signals, not a single numerical destiny.
The range in energy analysis illustrates the problem. The IEA’s Energy and AI work places 2035 global data-center electricity demand across cases spanning roughly 700–1,700 TWh, with a high-efficiency pathway around 970 TWh. The width is not analytical failure. It reflects uncertainty about model architecture, task demand, chip efficiency, utilization, flexible load and policy. The same uncertainty applies to labor, market structure and national access. IEA executive summary, IEA demand analysis
| Question for 2035 | Evidence already visible | Constructive path | Adverse or limiting path | Leading signal to watch through 2031 |
|---|---|---|---|---|
| 1. Does the transformer remain the universal core? | attention, scaling and post-training still dominate, while state-space, world-model, formal and neuro-symbolic research attacks memory, grounding and verification limits | hybrid systems combine language models, persistent state, learned world models, formal tools and specialist verifiers; progress becomes less dependent on scaling one architecture | frontier progress becomes increasingly expensive, incremental and benchmark-specific; alternative architectures fail to generalize | verified success on unfamiliar, long-horizon and physical tasks at equal total compute—not parameter count or one leaderboard |
| 2. Does inference become the largest compute economy? | unit prices fall rapidly, but reasoning, video, long context and agent loops multiply work per user task | routing, small models, quantization, efficient memory and workload flexibility make ubiquitous intelligence affordable without proportional resource growth | rebound dominates efficiency; peak demand, memory and power become persistent constraints even as each operation becomes cheaper | joules, dollars, latency and failure recovery per verified outcome, alongside utilization and 2035 electricity-case movement |
| 3. Does AI cross decisively into science and the physical world? | systems already accelerate code, molecular search, weather, design and selected robotics, but physical validation remains slow and liability is real | closed loops connect models with simulation, sensors, laboratories and constrained machines, shortening discovery and production cycles | simulation gaps, scarce experimental data, safety incidents and equipment economics confine gains mostly to digital work | reproducible discoveries, safe operations over full asset lives and the share of recommendations validated in the world |
| 4. Does intelligence commoditize while control concentrates? | open weights narrow capability gaps, yet cloud, identity, distribution, systems of record and energized capacity remain concentrated | interoperable protocols, model routing, portable evaluations and competition let buyers switch while specialist firms capture workflow value | a few integrated platforms control memory, discovery, permissions, payments and enterprise action even when model prices fall | switching time and cost, independent application margins, model-routing diversity and concentration at identity and transaction layers |
| 5. Do institutions translate productivity into broad prosperity? | task studies show meaningful gains, but firmwide ROI, entry-level hiring, wages and public-service capacity remain uneven | employers preserve apprenticeship, competition passes gains into prices and wages, and governments use AI to expand reliable services | productivity raises rents and work intensity while narrowing career entry, privacy and bargaining power | median wage and price pass-through, junior hiring, hours, job quality, SME productivity, service access and appeal outcomes |
| 6. Is there one interoperable AI economy or several capability tiers? | chips, leading models, cloud capacity and venture funding are concentrated; open weights and regional strategies widen access | shared compute, efficient models, local context and reciprocal standards support plural regional ecosystems | export controls, cyber conflict, proprietary gates and duplicated stacks divide countries and sectors into privileged and restricted tiers | cross-border model and chip access, research collaboration, regional utilization, local-language quality and continuity tests for essential services |
What these questions imply
First, the 2035 economy may be more heterogeneous than today’s model race suggests. A user may experience one assistant while a router calls a small private model, a specialist predictor, a frontier reasoner, a search system, a formal solver and a human approver. The strategic asset shifts from one model checkpoint toward the institution’s context, permissions, evaluations, action surfaces and evidence of successful outcomes.
Second, physical constraints will not disappear merely because algorithms improve. Efficiency changes which projects are economical, but lower cost expands demand. Power, transmission, HBM, packaging, cooling, water and skilled construction remain important; so do the permissions of communities that host them. The best long-horizon infrastructure is therefore modular and adaptable: reusable power and networks, buildings that can accept new equipment, diversified customers and financing matched to the life of each asset.
Third, the decisive economic uncertainty is distribution, not just capability. A powerful system can create consumer surplus while reducing wages, improve a public service while weakening due process, or raise corporate output while destroying the apprenticeship that produces future experts. By 2035, the important scorecard should join technical performance with cash yield, resource intensity, competition, worker mobility, service access and recourse.
Finally, the longer horizon strengthens rather than weakens the article’s central recommendation. No government, company or investor needs to predict artificial general intelligence to act intelligently now. Build portable data and evaluations; buy plural model routes; preserve human authority for consequential decisions; finance adaptable infrastructure; train people inside redesigned roles; and measure verified outcomes. Those capabilities retain value across every plausible 2035 path.
Methods, definitions and publication package
Four evidence classes
- Observed: a disclosed financial result, completed transaction, installed asset, reported user measure or directly measured event.
- Estimate: a modeled value for a present or historical quantity that cannot be counted directly.
- External forecast: another organization’s conditional projection, retained with its scope and publication date.
- Article scenario: an internally consistent future constructed for this article. It is not a prediction or an outside organization’s forecast.
The analysis uses official statistics, filings and technical papers where possible, then triangulates them with multilateral, bank, venture-capital, consulting, market-research, think-tank and company sources. No major section rests on one institution. Sources are selected claim by claim: a company is authoritative about its own stated capital plan, but not an independent authority on market leadership; a bank forecast may be useful when its assumptions are visible, but is not an observed result.
Five ledgers that must not be confused
| Ledger | Question answered | Examples | Accounting treatment |
|---|---|---|---|
| Capital formation | What long-lived asset is financed or built? | fabs, servers, data centers, grids, generation | Count the investment once even if it appears in a developer’s project value, a hyperscaler’s capex and a financing package. |
| Supplier revenue | Which provider receives payment? | chips, cloud, model APIs, applications, services | A transfer within the economy before downstream benefit; it is not automatically GDP impact. |
| Value enabled | What improvement is technically and commercially possible? | time saved, capacity added, loss avoided, conversion improved | Apply adoption, workflow suitability, reliability and realization filters. |
| Value captured | Which share becomes a measurable benefit, and for whom? | profit, public capacity, worker time, lower prices, consumer surplus | Allocate across stakeholders; do not assume it all becomes vendor revenue. |
| Net benefit or value added | What remains after implementation, operation, failure and transition costs? | retained organizational benefit and additional output | Subtract costs once and exclude transfers and double counting. |
Audience measures retain their denominator: weekly active users, monthly active users, subscriptions, messages, API tokens and embedded feature reach are not interchangeable. “Open source” is not used as a synonym for downloadable weights; code, data, training recipe, weights and license rights are separate dimensions. Data-center capacity distinguishes announced, permitted, contracted, energized and utilized megawatts. Water withdrawal, consumption, discharge, replenishment and indirect water are separate measures. Task exposure, augmentation, displacement and unemployment are separate labor outcomes.
Companion files
- Source register
- Past–present–future quantitative scoreboard
- Economic ledger
- Country and regional impact scenario
- Industry value scenario
- Model chronology and comparison
- Limitations and research-frontier ledger
- Infrastructure stack
- Power and water evidence
- Adoption evidence
- Capital and valuation evidence
- Labor evidence
- Scenario matrix
- 2035 horizon questions and leading indicators
- Source-diversity audit
- Contradictions and reconciliation log
- Detailed methodology and update protocol
This edition extends the economic waterfall in Agentic AI: $16Tn in Global Economic Value Enabled and $8Tn Captured, the workflow architecture in Agentic AI: From Better Answers to Entirely New Industry Economics, and the industry/startup theses in Agentic AI: Transforming Industries.
Editorial disclosure
This is a research synthesis, not investment, legal, employment or engineering advice. Forecasts are conditional. Private financings may include preferences and strategic obligations hidden by headline valuations. Company usage and environmental measures use company-defined boundaries. Benchmarks can be optimized, contaminated or unrepresentative. Laws, export controls, product access and market values can change after publication.
