Skip to content
all longreads
Longread#AI#Economics#DistributedSystems

The agent economy: who benefits when machines make deals

Agents already negotiate over goods, allocate tasks through auctions, and pay for digital services. Yet a working payment system leaves questions about price, quality, power, and responsibility unresolved. What have we learned in the year since Virtual Agent Economies, and which conclusions still run ahead of the evidence?

20 September 2026≈ 24 minsources and limitations ↓

Sources checked through 19 September 2026. Experiments, protocol proposals, and company claims are distinguished throughout. Preprints are considered with their stated limitations; the author's diagrams and illustrative calculations are labeled separately.

01

What counts as an agent economy?

In October 2025, I discussed Virtual Agent Economies from Google DeepMind. I was interested in a specific transition: what happens when software can negotiate work and resources independently, without people participating in every individual transaction? A follow-up post explored how such markets might be designed: their boundaries, auctions, trust, and social objectives.

Enough has happened since then to return to the subject with more demanding questions. Who actually makes the decisions? Whose money do they spend? Is useful work being purchased from an independent provider? Who can establish that the work is inadequate? The word “economy” quickly loses explanatory power when it describes, indiscriminately, an internal loop of five models, a ticket purchase for a person, and a network of independent service providers.

For this analysis, I propose a working definition: agent economies are systems of exchange in which software has been delegated consequential decisions about counterparties, resource allocation, or transaction terms. The degree of autonomy can vary. An agent acts for an owner or organization; the technical ability to authorize a payment does not, by itself, make it an independent owner, a company, or a bearer of legal responsibility.

Three structures behind one name
Three structures behind one nameOne operatorRepresentationIndependent marketOrchestratorA123Tasks to workersa shared budgetPPrincipalASAgentSellerWhose interests are served?1234Different participantsterms of exchangeThe author's relationship diagram

The author's diagram. A shared budget, representation, and exchange among independent participants create different relationships. Each structure needs its own evidence of benefit.

System
Internal organization of computation
Where the economic choice arises
How to allocate tasks and a notional budget among one operator's models
What needs to be demonstrated
Better quality or lower total cost than a simpler approach
System
An agent representing a person
Where the economic choice arises
What to buy, whom to sell to, and which terms to accept within its authority
What needs to be demonstrated
A benefit to the principal after representation costs and preference compliance are considered
System
A market of independent providers
Where the economic choice arises
Whom to buy work from and at what price; whether to hire a subcontractor
What needs to be demonstrated
Real external demand, outcome verification, settlement, and resilient rules

This is my own classification, not an established standard. Its purpose is to prevent success in one row from being carried over into another. Effective model routing does not establish that a viable market has emerged. Large flows between wallets do not establish that the end user is better off.

02

From sandbox to market: what changed in a year

The original paper, dated 12 September 2025, describes economies along two axes: whether they emerged spontaneously or were deliberately designed, and how freely they interact with the outside human economy. The authors anticipate the growth of permeable systems and propose building governance mechanisms in advance. This is a framework for scenarios and policy choices, rather than a measurement of an existing market.

Two independent axes of the agent economy
Two independent axes of the agent economyClosedPermeableDesignedEmergentA designedtest marketPredefined rules,real providersEmergent exchangein a closed spaceEmergent networkwith externalsettlementThe author's examplesEconomic isolation does notguarantee computational security

Axes from Virtual Agent Economies; examples are the author's classification. How an economy emerges and how it connects to the outside economy are separate properties. Economic isolation does not itself guarantee computational security.

I find it useful to treat permeability as a set of specific permissions. Can the system buy data externally? Can it sell code it has produced? Can it reserve physical equipment and then transfer that commitment to another provider? Each permission allows different consequences to cross the boundary: expenditure, information leakage, production delays, or an unfulfilled agreement.

An economic sandbox therefore differs from a container for safely executing code. A process can be securely isolated from the file system while still being allowed to spend its entire budget on unnecessary services. External payments can be forbidden while processing sensitive data remains permitted. Boundaries for money, data, authority, and physical actions will have to be specified separately.

When
November 2025
Which question became easier to investigate
How agents select offers in a simulated market
When
February 2026
Which question became easier to investigate
How to connect task delegation with authority and acceptance
When
April 2026
Which question became easier to investigate
Whether agents can negotiate trades in real goods
When
June–August 2026
What appeared
Which question became easier to investigate
What prices and auctions contribute to organizing work
When
July 2026
Which question became easier to investigate
How closely payment activity corresponds to independent exchange
When
September 2026
Which question became easier to investigate
How agents' economic behavior develops over longer periods

This chronology does not describe a single path to maturity. Negotiation research, internal auctions, and payment systems are developing in parallel. Connecting them into a working market is an engineering and institutional task in its own right.

03

Real trades and the strength of a representative

The most tangible example is Anthropic's Project Deal. In December 2025, 69 employee volunteers gave agents their preferences for buying and selling personal belongings; each received a $100 budget. The results were published on 24 April 2026. In the run that was actually carried out, Opus 4.5 agents negotiated 186 deals; the report's appendix lists 206 items sold for $4,010. Deals and items are different accounting units here.

Opus and Haiku were compared in separate runs whose trades were not executed. The authors found an advantage for the stronger model. Among the 28 participants represented by different models across the two mixed runs, no statistically significant preference for one model was detected: this does not prove that people are inherently unable to notice a difference.

That leads to a more interesting question than whether trade occurred. Suppose two representatives understand their principals' preferences equally well, but one negotiates more effectively and obtains a lower price. The buyer is better off, yet society may not have gained any additional useful work: part of the effect could simply be a transfer of existing surplus from seller to buyer. Increasing total gains requires something else, such as finding a previously unachievable trade, saving time, or obtaining a more suitable product.

I would therefore measure two outcomes separately. First, does the system create more mutually beneficial exchanges? Second, who captures the gains within those exchanges? A market can become faster while becoming less favorable to participants with cheaper models, poor instructions, or limited access to data. These are different claims, and each needs its own controlled experiment.

In a practical product, user satisfaction is insufficient too. The user knows that a purchase happened, but usually cannot know what price another representative might have obtained under the same constraints. A reasonable check is to retain the original preferences and compare the outcome with an available alternative: a fixed price, a simple search rule, or another agent. The cost of that comparison must itself be included in the service's expenses.

This creates a distinct basis for competition: the quality of representation. A convenient interface and a persuasive explanation of a purchase need not coincide with effective protection of the principal's interests. If the agent receives compensation from the seller, that conflict will need to be visible and auditable. This is an implication for future product design, rather than an established finding of Anthropic's experiment.

04

Delegation: who accepts the work?

A direct continuation of the original research is Intelligent AI Delegation, published on 12 February 2026. Tomašev, Franklin, and Osindero consider task handoffs alongside authority, responsibility, monitoring, and the ability to replace a provider. Their proposed approach starts by defining verifiable outcome conditions; work that is too difficult to verify should be broken down further.

Consider an example of my own. An agent is asked to prepare market research. It finds a provider promising an attractive report for a small fee. If the agreement goes no further than “do good market research,” the economic problem begins before payment. It is unclear whether outdated data, secondary summaries, and unsupported conclusions count as acceptable work.

An actionable brief might require a list of companies, verification dates, links to primary documents, an explicit separation of facts from assumptions, and a predefined sample for independent checking. Format can be checked automatically. Establishing that claims match their sources costs more. Assessing how completely the report covers the market costs more still. These three checks have different prices; a signed file cannot turn them into one free operation.

A deal leads to acceptance—or a dispute
A deal leads to acceptance—or a disputePGoal and limitAChoose a providerTermssubcontracting rightsWExecutionsub.An illustrative sequence; orderdepends on transaction termsIndependentverificationacceptednot acceptedSettlementDisputePayment ≠ quality verification

The author's process example. Payment and outcome verification are distinct events. The order is not a property of every protocol: settlement may be upfront, staged, or after acceptance. Subcontracting rights are agreed separately.

Subcontracting adds complexity. If the first agent hires a second, and the second hires a third, the system needs to know who was allowed to pass on data, who had authority to increase the budget, and who still owes the client a result. Recording the entire chain helps, but the record alone does not create an accountable party. That obligation must be part of the system's rules and the relationships among its operators.

The ERC-8183 proposal, created on 25 February 2026, gives this problem a concrete structure: a client escrows payment, a provider submits the result, and an evaluator accepts or rejects it. The client may also act as evaluator. The specification describes state transitions and the movement of funds; the quality of the evaluator's judgment remains a property of the chosen implementation.

For me, this is an important engineering boundary. Delivering a result, deciding to accept it, and transferring money are three events with different grounds for proceeding. The system needs consistent task identifiers and result versions, protection against duplicate payments, a timeout, and a clear route through disputes. These are familiar distributed-systems problems with financial consequences.

05

Task markets: what prices convey and what negotiation costs

When providers have different skills and costs, prices can convey useful information. A central coordinator does not need to know everything about every participant in advance: each participant states the terms on which it will take the work. This argument rests on a strong assumption, however: the provider understands its own capabilities and expenses well enough.

MarketBench, published on 26 April 2026, tests precisely this economic behavior in a task market. Its authors identify problems with estimating success probabilities and execution costs. A polished bid may be another plausible model response rather than a sound estimate of the cost of doing the work.

Market design as a way for a system to learn

In Economy of Minds, dated 1 June 2026, auctions and notional wealth are used to select and modify agents. Payments, bankruptcies, and mutations are disabled for the final evaluation. The MATH ablation table separately reports a mean of 43.9% and a best result of 57.0%, against a baseline of 51.9%. The aggregation method is not described in enough detail to automatically interpret these as the mean and best final accuracy across independent runs. Equal total compute cost has not been demonstrated either.

Notional payments here help attribute the contribution of individual actions: useful agents survive, unsuccessful ones are replaced, and instructions change while model weights remain fixed. The researchers use this process to discover ways to organize computation. Its applicability to a market of independent providers requires a separate experiment.

A market with a coordinator and subcontractors

AgentLance, dated 24 August 2026, is closer to a labor-market model, with bids, reputation, provider selection, and subcontracting. A central allocator remains. The authors report improvements on their metric relative to selected alternatives, but count execution costs without market communication expenses and describe the system as a proof of concept.

That exclusion changes the practical conclusion. Suppose execution becomes 20 cents cheaper. If collecting bids, assessing reputation, and negotiating costs 30 cents, the system still loses on total cost. The same mechanism may pay for itself on an expensive task. One optimization problem, then, is deciding which tasks should be put out to bid at all.

Condition
Providers actually differ
Why it helps
Finding a suitable provider can add value
What could undermine the economics
Providers differ little in quality, data, tools, and total cost
Condition
Acceptance is inexpensive
Why it helps
The client can establish what it is paying for
What could undermine the economics
Verification requires doing all the work again
Condition
The gain exceeds selection costs
Why it helps
Negotiation pays for itself on the particular task
What could undermine the economics
A small task accumulates a long discussion
Condition
Repeated transactions expose mistakes
Why it helps
Reputation contains a useful history
What could undermine the economics
An unsuccessful participant can easily change identity

This table is my practical framework. For recurring inexpensive operations, a provider selected in advance may make more economic sense than an auction. For occasional complex tasks, comparing offers may produce a substantial gain. Markets have an appropriate scope of application too.

06

How agents pay: several different problems called a protocol

Payment infrastructure has advanced substantially, but the collection of names can create a false impression that a single standard for the agent economy has emerged. Different projects address different parts of the interaction.

Project and date
Primary purpose
Convey verifiable evidence of payment authority
What this does not establish
That the agent correctly understood all of a person's preferences
Project and date
Primary purpose
Coordinate interactions between an agent and commerce systems
What this does not establish
That sellers are equally accessible or beneficial to the buyer
Project and date
Primary purpose
Coordinate programmatic payment requests and settlement for goods or services
What this does not establish
That the purchase produced the expected outcome
Primary purpose
Develop open payments over HTTP
What this does not establish
That payment counts measure independent demand
Primary purpose
Connect identity, feedback, and validation signals
What this does not establish
That a provider's declared capability has been demonstrated

Two entries predate my October post: AP2 and ERC-8004 are included as background. The x402 foundation's launch date is also different from the protocol's original launch. Preserving these distinctions prevents the history from becoming a list of repeated “launches” of the same solution.

An April IMF Notes paper distinguishes intent, authorization, and settlement. Probabilistic agent behavior has to connect to infrastructure expected to follow definite execution rules. This is the authors' analytical framework; it does not establish a new binding legal regime for agents.

Consider a hypothetical data purchase. An agent selects a dataset, receives authority to spend ten dollars, and successfully pays. The payment service can confirm the recipient, amount, and authorization. Yet the dataset may be outdated, incomplete, or useless for the original research. Those failures require delivery terms and data-quality checks.

The reverse is possible too: useful data arrives, but payment stalls or is repeated after a network failure. The provider wants to be paid; the client wants to avoid being charged twice. That creates a need to track order state and coordinate retries. A low transfer fee helps, but does not eliminate this work.

A blockchain is one implementation option. Some open markets benefit from public records and programmable settlement; within an organization, a task registry, limited permissions, and conventional billing may suffice. The choice should follow the participants and the trust model. The word “agent” mandates neither cryptocurrency nor any particular payment method.

07

What a transaction counter measures

How Agentic Is Agentic Commerce?, published on 14 July 2026, examines x402 on Base from 17 September 2025 through 23 June 2026. The authors identify 136.7 million settlements totaling $44.1 million. Their classification labels 21.20% of the transaction count fictitious, 63.78% internal settlement within linked clusters, and leaves 15.02% unattributed.

What an x402 payment count leaves unresolved
What an x402 payment count leaves unresolvedShares of settlement count · Base17 Sep 2025 — 23 Jun 2026Fictitious21.20%Self-payments and closed loopsunder the authors' criteriaInternal63.78%A linked cluster does notprove fraud or the absenceof servicesUnattributed15.02%Independence and serviceusefulness arenot established15.02% ≠ demonstrated commerceCount share ≠ value shareHow Agentic Is Agentic Commerce? · v1

Shares of settlement count in the Base sample, not shares of monetary value. Under the authors' criteria, fictitious means self-payments and provably closed loops; internal means settlements within linked clusters, not proof of fraud or absent services. For unattributed settlements, independence and service usefulness are not established. Source: How Agentic Is Agentic Commerce?, v1, 14 July 2026.

These are shares of the transaction count, not monetary value. The blockchain does not reveal service delivery, and links between wallets do not establish malicious intent. The study covers a particular period and uses specific rules to identify transactions; a July preprint must not be presented as a September census of the market.

The measurement problem nevertheless extends beyond x402. If a counter increases with every internal movement of funds, one useful task can generate dozens of events. If the system subsidizes those events, it may encourage activity that it subsequently presents as evidence of demand. What needs examining is the path from an order to an accepted result and the ultimate beneficiary.

This recalls a subject I have discussed before: tokens as proof of work. A counter records that a system was busy. Establishing value requires another connection: a specific need, the resulting output, its use, and its consequences. Otherwise, a single organization can appear to be a huge market by repeatedly moving a budget among its own providers.

I would look at repeat purchases by external customers, the share of independent providers, acceptance and refund rates, revenue concentration, and demand after subsidies change. A zero payment does not imply zero benefit, while a positive payment does not guarantee a positive benefit. Still, a paying repeat customer provides a more meaningful signal than another registry entry or technical transfer.

The accounting unit also needs its own definition. Is an agent a model, a process, a key, a wallet, or an independent operator? A hundred copies of one process can inflate a participant count without changing competition at all. Without an answer, the “number of agents” is a poor measure of an economy's size.

08

Identity, reputation, and control over choice

ERC-8004 defines identity, reputation, and validation registries. The version checked remains a Draft; payments are outside the specification's scope. Its security section explicitly states that registration does not guarantee the operation of declared capabilities, and that multiple fabricated identities can be used to manipulate feedback.

The empirical ERC-8004 audit, v2 dated 8 July 2026, studies data through 13 May. Under its criterion, only 3%, 4%, and 15% of registrations on Ethereum, BSC, and Base, respectively, have a compliant registration file with a declared service. These are not percentages of independently verified, competent providers; the heuristics used to identify suspicious feedback do not establish an owner's culpability either.

Reputation should answer a much narrower question than whether an agent is “good.” Who checked its work? On which tasks? With what cost of error? When? After which change of model or tools? A provider with a hundred positive reviews for text translation does not automatically become qualified to assess engineering safety. An average rating hides precisely the differences that lead a client to seek a specialist.

A NIST article dated 27 August 2026 emphasizes agents having their own identities, limited and short-lived authority, and a connection to their operator. Its authors also warn about fatigue from frequent approval requests. Identity and access management is an existing engineering discipline; agents will require its careful application and further development.

There is another layer of power: who enters the candidate list at all. Microsoft Research's Magentic Marketplace simulations show buyers responding to offer order and making worse choices as the market grows. That is a reason to examine search and queue design, rather than treating them as neutral packaging around a transaction.

My interpretation is that an open message format leaves plenty of room for a highly centralized market. One operator can control the directory, ranking, quality history, and access to customers. Independent providers may speak a common language while depending on the rules of discovery. Reputation portability and the ability to change intermediaries are therefore at least as economically interesting as API interoperability.

A shared protocol can still have a gatekeeper
A shared protocol can still have a gatekeeperProviders123Open protocolIntermediaryCatalogRankingReputationVisiblechoicesABuyer's agentThe author's example ofa possible market structure

The author's illustration of a possible market structure. Message interoperability leaves open who controls the catalog, ranking, and quality history. Control over discovery affects access to the buyer.

In a July MIT Sloan article, Ramesh Raskar connects Project NANDA to open infrastructure for agent discovery, identity, and coordination. This is a research program and a position on how the market should be organized. The future number of agents and the timetable for realizing that design remain forecasts.

09

What simulations show, and what their design assumes

The recent paper But How Would AI Agents Run a Town's Economy? was published on 10 September 2026. In a closed environment with 100 agents, a twelvefold increase in tourist demand raised business revenue 4.62 times but barely changed wages; prices were revised for 0.3% of menu items. The longest horizon, 26 weeks, is represented by a single run. The authors acknowledge that price stickiness is partly built into the environment; they did not test the effect of reminding owners that raising wages was an option.

I would read this result as a diagnostic of the model and its environment. It reveals where that particular setup fails to reproduce expected behavior. It cannot establish an inevitable structure for a future machine economy. Still less can one virtual town stand in for real economies without examining how needs, prices, hiring, and the possibility of bankruptcy are defined.

Giving an agent money does not yet give it an economic motivation. Why should it spend? What does running short of funds mean? Can it succeed by doing nothing? How does the system tell the agent that its strategy has failed? The answers may determine the outcome more strongly than an impressive social simulation.

Consider a hypothetical example: if a business receives customers regardless of quality, and its owner has no reason to reinvest profits, accumulating money may be rational under the specified rules. Calling this “LLM greed” would make an appealing story but a poor explanation. The first things to examine are rewards, available actions, and constraints.

For the same reason, economic simulations cannot be evaluated solely on the plausibility of their dialogue. They need repeated runs, different models, changes to environmental rules, and comparisons with simple programmed strategies. Does the effect survive when the eloquent language is removed? Does an ordinary algorithm produce it under the same constraints? Such checks help establish what the language model actually contributes.

10

The full cost of an accepted outcome

An agent economy brings two sets of accounts into one system. Internally, a provider consumes computation and time. Externally, the client pays for an outcome. In between lie discovery, negotiation, verification, rework, downtime, and disputes. Count only the first model call, and almost any interesting approach will appear cheaper than it turns out to be in operation.

To compare approaches, I propose a simple accounting framework. Take all costs over a period, for both successful and unsuccessful attempts, and divide by the number of outcomes that passed the same predefined acceptance checks. The numerator includes execution, discovery and coordination, verification and rework, payment expenses, operations, and recognized losses from errors. Retries are included exactly once.

In an illustrative example, the first provider spends $0.20 on execution, $0.25 on coordination, and $0.80 on verification and rework. Sixty out of a hundred attempts are accepted: $125 / 60 ≈ $2.08 per outcome. For a second provider, execution costs $0.80, coordination $0.05, and verification $0.25; ninety attempts are accepted: $110 / 90 ≈ $1.22. A higher execution price can coexist with a lower cost per accepted outcome. This is an arithmetic illustration, not a measurement of a particular product.

A cheap attempt can produce an expensive outcome

Illustrative assumptions. Costs are averaged over every attempted task, including failures; the acceptance rate is measured after the included rework. No automatic retries are added.

The cost of an accepted outcomeIllustrative costs · 100 attemptsExecution$20.00+Discovery andcoordination$25.00+Review andrework$80.00$125.00spend on all100 attempts60acceptedoutcomes$2.08per acceptedoutcome÷=Accepted: 60Rejected: 40Spend includes rejected attempts.Divide by accepted outcomes.
Total spend per 100 attempts$125.00
Cost per accepted outcome$2.08

This simplified example excludes consequential losses, fixed platform costs, and the value of waiting. Add those before making an operating decision.

Reducing this metric is useful only when acceptance standards are comparable. Relaxing checks increases the number of accepted outcomes but makes the comparison meaningless. The same happens if a cheap agent gets only easy tasks while an expensive one handles the hardest, and average costs are then compared without accounting for the mix of work.

A second metric is needed: the outcome's value to the client. Ten correct reports that nobody used may cost less than one useful decision and still be a worse purchase. For a research service, value might mean reducing the time needed to test a hypothesis; for software development, accepted changes without an increase in defects; for procurement, a suitable product with appropriate delivery terms.

Finally, an ordinary average cost does a poor job of representing rare, severe consequences. Irreversible actions warrant separate limits on authority and risk exposure. Saving a few cents cannot compensate for the absence of a controllable boundary on damage.

11

How to test a market in practice

This research suggests a fairly grounded program for the next experiment. I would begin with a narrow class of tasks where quality can be defined before selecting a provider and repeat orders are possible: a verifiable data transformation, for example, or a bounded code change. An open-ended promise to “solve any business problem” makes nearly every necessary check harder.

Test
Additional value
What to compare
A person or simple process, one agent, a coordinator, and a market on the same task set
Which decision it informs
Whether a more complex organization is justified
Test
Total cost
What to compare
Every call, negotiation, check, correction, and unit of human time
Which decision it informs
Whether provider selection pays for itself
Test
Alignment with the principal's interests
What to compare
Original preferences and constraints against the actual outcome
Which decision it informs
Whether the representative acts for its principal's benefit
Test
Independent demand
What to compare
External repeat orders, separately from internal and subsidized orders
Which decision it informs
Whether a market exists beyond the demonstration
Test
Resilient rules
What to compare
Provider failure, duplicate payment, poor work, false feedback, and a model change
Which decision it informs
Whether consequences can be bounded and work can continue
Test
Distribution of gains
What to compare
Outcomes for owners of different models and providers of different sizes
Which decision it informs
Whether average growth conceals losses for some participants

Stopping conditions should be defined before such a pilot begins. For example, provider-selection expenses consistently exceed the savings; verification takes as long as doing the work; the client does not return without a subsidy; or one intermediary accounts for most of the turnover. Specific thresholds depend on the task; a single set cannot honestly be prescribed for every market.

Three possible arrangements in the near future

My baseline scenario is the spread of managed procurement among organizations that already know one another. Ownership, data access, responsibility, and acceptance are easier to establish there. Open markets of independent providers may develop alongside them, particularly for small digital services whose results are inexpensive to verify. Internal model auctions may remain a separate way to organize computation without entering an external market at all.

This is a forecast, not an established sequence of development. It should change if durable evidence emerges of repeat independent demand, competitive prices after all verification costs are included, and the portability of orders between platforms. A prominent announcement or a larger wallet count would not be sufficient.

My original discussion centered on how a governed agent economy might be organized. We can now add a testable formulation: show a chain in which an external client received an accepted outcome, an independent provider earned money, and coordination and oversight paid for themselves. Then show what happens along that same chain when the provider makes a mistake. That is where I would start the next conversation about scale.

12

Sources and research boundaries

This is a thematic review of primary research and official descriptions checked through 19 September 2026, rather than a systematic review of the entire literature. Selection follows the questions raised in the original discussion: coordination, market permeability, distribution of gains, and governance. Market-size forecasts and the prices of agent-related tokens were not used as evidence of demand for useful services.

  • Experimental findings apply to the models, environments, and baselines described. This article does not independently reproduce the experiments.
  • A preprint, a specification, a company report, and institutional analysis provide different kinds of evidence; none substitutes for measurements from operations.
  • Publication dates differ from experiment dates and the ends of data windows. Those distinctions are retained wherever they affect the conclusion.
  • The diagrams, practical testing program, cost framework, and future scenarios are the author's analysis. The calculator uses illustrative figures.

Foundations and research

  1. Book Cube · Virtual Agent Economies, part I — 1 October 2025: the author's original discussion.
  2. Book Cube · Virtual Agent Economies, part II — 1 October 2025: discussion of the scenario framework.
  3. Tomašev et al. · Virtual Agent Economies — 12 September 2025, v1. A conceptual paper from Google DeepMind.
  4. Book Cube · Token burn as proof of work — The author's discussion of computation expenditure versus value.
  5. Tomašev, Franklin, Osindero · Intelligent AI Delegation — 12 February 2026, v1. A proposed delegation framework.
  6. Qi et al. · Economy of Minds — 1 June 2026, v1. Experiments with auctions and agent selection.
  7. Liu et al. · Markets, Not Planners — 24 August 2026, v1. AgentLance; market communication costs are excluded from the main metric.
  8. MarketBench — 26 April 2026. Evaluation of model behavior in a task market.
  9. Microsoft Research · Magentic Marketplace — 5 November 2025. An open simulation environment for two-sided markets.

Measurements and limitations

  1. Anthropic · Project Deal — 24 April 2026; experiment conducted in December 2025. One executed run and three research runs.
  2. Anthropic · Project Deal: statistical appendix — Table 3: items and transaction value by run.
  3. Ling et al. · How Agentic Is Agentic Commerce? — 14 July 2026, v1. Main Base window: 17 September 2025–23 June 2026.
  4. Can Trustless Agents Be Trusted? — v2, 8 July 2026; data through 13 May. A heuristic audit of ERC-8004.
  5. Regmi et al. · But How Would AI Agents Run a Town’s Economy? — 10 September 2026, v1. A closed simulation with limited long-horizon replication.

Protocols and institutions

  1. Google · Agent Payments Protocol — 16 September 2025. The announcement predates the original post.
  2. Google · Universal Commerce Protocol — 11 January 2026. A protocol for commerce-system interoperability.
  3. Stripe / Tempo · Machine Payments Protocol — 18 March 2026. Programmatic service payments; the developers' announcement.
  4. Linux Foundation · x402 Foundation operational launch — 14 July 2026. Operational launch; not a measurement of demand.
  5. ERC-8004 · Trustless Agents — Created 13 August 2025; the version checked on 19 September 2026 is marked Draft.
  6. ERC-8183 · Agentic Commerce — Created 25 February 2026. A proposal for job escrow and outcome attestation.
  7. NIST · Back to the Future: Why Agentic AI Needs a Strong Identity Foundation — 27 August 2026. Identity, authorization scope, and delegation.
  8. Davidovic, Tourpe · How Agentic AI Will Reshape Payments — IMF Notes 2026/004. The page states 24 April 2026, despite the date in its URL.
  9. MIT Sloan · Who will own the AI agent economy? — 13 July 2026. Ramesh Raskar's perspective and the aims of Project NANDA.