A definition should explain how work gets done
By an AI-native organization, I mean a company that designs value creation around the availability of machine intelligence: people set intent and constraints, systems perform the work they are allowed to do, results are checked, and experience feeds back into processes and knowledge. This definition requires changes in work, accountability, and learning together. Buying access to a model solves only the access problem.
Three states are useful to distinguish. They can coexist within one company: support already follows the new model, procurement uses an assistant, and strategic decisions remain largely human. This is a map of differences, not a maturity certification or a mandatory ladder.
| State | What changes | What can be verified |
|---|---|---|
| AI assists a person | An employee performs an existing task faster: writing, searching, or analyzing. | The quality and duration of a particular task; the benefit may not yet reach the customer. |
| AI is embedded in a workflow | The path of work changes: classification, decision preparation, validation, and exception handling. | The time and cost of the completed process, including the workload at adjacent stages. |
| The organization operates as AI-native | Workflows, authority, the shared platform, knowledge, and incentives are designed together. | Repeatable impact across several workflows and the ability to transfer an approach that works. |
AI-first usually expresses a priority: consider an AI solution first. AI-native, as used here, describes how a company observably operates. A company selling an AI product may retain its old internal structure; a manufacturer that sells no models may substantially redesign planning, procurement, and service. Physical operations constrain automation, but do not rule out this transition.
What distinguishes this from earlier automation is the work that can be delegated to a system. Predefined rules handle known transitions well; language models add the ability to work with unstructured context and choose the next action. That flexibility also introduces probabilistic errors. Verification and feedback therefore become central to organizational design. A new name does not erase what earlier generations of automation have taught us.
The practical question is: “Which workflow now operates differently, who owns its outcome, and what demonstrates the impact?” If the answer stops at users and licenses, the operating model remains unspecified. Dependence on one model is not a sign of maturity, either: the company needs a way to continue when the provider fails.
There is evidence, but it describes different levels of impact
This research must avoid adding together percentages from incomparable studies. A randomized experiment helps estimate a causal effect on selected tasks. An adoption study within one company reveals a specific production setting. An industry survey describes participants' answers. A vendor publication offers a framework and demonstrates a case, but does not establish that its results transfer across the market.
In McKinsey's State of AI 2026, published on August 25, 80% of respondents report higher personal productivity, while 37% report a positive impact on their company's EBIT. The survey covers 1,719 participants in 97 countries. This is a strong signal of a gap between individual and corporate impact, but it is neither a measurement of profits at randomly selected businesses nor a causal estimate of adoption.
Microsoft Work Trend Index 2026 also shifts attention toward organizational conditions. Its Frontier Firm model provides useful industry language for discussing how people and agents work together. But a vendor's framework, survey responses, and analysis of their relationships do not prescribe an organizational chart. They cannot establish a universal number of agents per employee or direct reports per manager.
| Source and design | Finding | Limits of the conclusion |
|---|---|---|
| Generative AI at Work, QJE, 2025 | 5,172 support agents; introducing an assistant was associated with an average 15% increase in issues resolved per hour. | One work setting, with benefits varying by employee experience. This does not estimate whole-company productivity. |
| The Cybernetic Teammate, 2026 | In the P&G experiment, individuals using AI matched teams without AI on a product-idea development task. | A bounded task and proposal assessment; not evidence that a team can be replaced across a complete business cycle. |
| METR, early-2025 experiment | 16 experienced developers, 246 tasks in familiar projects; access to AI increased completion time by 19%. | A narrow sample and early-2025 tools. The result does not generalize to all tasks or newer models. |
| DORA, 2025 | AI's impact depends on the capabilities of the development system: its platform, feedback, quality, and clear processes. | Relationships in organizational research are not a causal estimate for a particular change in your company. |
METR published an important update on February 24, 2026: the follow-up measurement encountered changes in participant and task selection. People were reluctant to perform tasks without AI where they expected substantial gains. The original 19% therefore cannot serve as a current forecast, and later estimates are not a clean comparison between model generations. Self-reported benefits do not replace measurement either.
Together, these findings support a more modest, useful conclusion: impact depends on the task, experience, available context, tools, and how verification is organized. An adoption decision requires measurement in your own workflow. Neither an optimistic industry report nor a negative experiment relieves a leader of that work.
The Navigating the Jagged Technological Frontier experiment with 758 BCG consultants illustrates this too. Its final version was published in March 2026, but the research was first presented in 2023 and concerns GPT-4 from that period. On tasks within the model's capabilities, participants using AI worked faster. On one task outside that boundary, the share of correct solutions fell by an average of 19 percentage points. Delegation should be scoped to a validated task class. A job title and a confident answer do not establish that boundary.
As execution gets cheaper, decisions and context transfer matter more
In my article on AI-native organizations, I explored the shift from an organizational chart to a map of work. When requirements, code, tests, and documents can be prepared faster, the queues become more visible: waiting for a decision, finding an owner, explaining the task again, and coordinating across functions. Reducing the time to produce an artifact does not necessarily shorten those queues.01This series began by redesigning the path from an idea to a result. Organizational design takes the same argument to the company level.From classic PDLC to AI-native development
Consider an illustrative process: preparing a proposal takes two hours, and approval takes three days. Cut preparation to twenty minutes, and the customer still waits almost three days. Cheaper preparation may even increase the number of proposals entering the same approval queue. Local productivity then rises alongside total process time. This explains a mechanism; it is not an adoption statistic.
My conclusion is that the design unit should be a complete workflow. It exposes three constraints: access to reliable context, the ability to make decisions, and the capacity to verify results. Different processes will have different constraints. Sometimes removing a redundant approval or fixing a data interface is more useful than replacing the model.
Intent → accessible context → execution → verification → accepted result → updated knowledge
This also explains why initial adoption may require more resources. The productivity J-curve research examines the complementary intangible investments that general-purpose technologies require. For a company, it provides a useful way to understand spending on processes, knowledge, and learning. It does not promise that every prolonged pilot will eventually pay off: intermediate results and reasons to continue must remain visible.
The first project is one workflow from request to accepted result
Start with a workflow that has a clear customer, enough repetitions, and an available way to check quality. Handling a specific category of support requests is one example. “Introduce an agent into support” is too broad. “Reduce resolution time for delivery-status inquiries while maintaining quality and access to a human” defines an experiment.
Example: a support request within the product feedback loop
In the existing process, an operator reads the message, finds the order in several systems, checks the applicable rule, and drafts a reply. In the redesigned process, the system identifies an eligible category, retrieves facts through restricted tools, proposes a resolution, and checks it against the rules. A simple status confirmation can be sent automatically; a disputed compensation request goes to the responsible employee. Authority is specified for each action.
The organizational change comes next. Recurring reasons for contacting support become work for the product team, supported by examples and an estimate of their prevalence. Product improvements prevent future requests, while a newly resolved exception updates the rules and test cases. Support and product agree on a shared outcome: customers need less help, and their problems actually get resolved. The bot's message count explains very little here.
| Question | What to establish before launch |
|---|---|
| Outcome | A resolved request that is not reopened within an agreed window; verified correctness of the answer. |
| Context | The order, status, current rules, and the source and update time of each fact. |
| Action boundaries | What may be read, changed, and sent; which amounts, situations, and recipients require a human. |
| Exceptions | Who accepts the handoff, within what time, and with which supporting evidence. |
| Comparison | Comparable requests handled without the new process, costs on both sides, quality, and repeat contacts. |
The same approach applies to a sales proposal, an investigation into a production deviation, or a software change. Acceptance criteria will differ. Sales cannot be judged by attractive proposals, engineering by generated lines, or procurement by suppliers found. The result must reach its customer and pass the checks appropriate to its domain.
An industry case: what Klarna actually demonstrates
In its 2025 annual report, Klarna says AI handled 80% of support chats and estimates annual savings of approximately $59 million. The same report describes retaining human support for customers who choose it. These are company metrics and estimates, not an independent experiment. The case shows that large-scale automation can coexist with a human service channel; it neither prescribes staffing reductions nor establishes the effect across every business function.
Coordination is redistributed, beyond changes in headcount
In the old model, much coordination lives in people's heads: a manager knows where to forward a question, an experienced specialist remembers an exception, and a coordinator checks statuses. Some of this work can move into shared interfaces, rules, and agent workflows. The manager's freed attention should go to priorities, conflicting objectives, people's development, and decisions under high uncertainty.
My organizational model distinguishes research, product and platform teams, and repeatable operations. Research needs joint exploration and close discussion; product and platform work needs autonomy within agreed outcomes; repeatable operations need stable rules and reliable exception handling. The manager-to-contributor ratio should follow the nature of the work. It cannot be copied from a particular division of a large technology company.
An agent has permitted actions, but no organizational accountability to the customer. Each workflow therefore retains a person who decides on quality, budget, and acceptable risk. The platform team owns shared access and execution mechanisms; data owners are responsible for meaning and freshness; domain experts define the checks. One person may perform several of these functions, but the functions remain.
| Decision | Accountable party | The agent's role |
|---|---|---|
| Which problem to solve, and why | Business or product owner | Collects evidence and options, showing their basis. |
| How to execute an authorized task | The workflow team, within agreed boundaries | Plans and performs permitted steps. |
| Whether to accept the result | The designated quality owner | Runs checks and gathers evidence; does not change the criteria itself. |
| Who handles harm or failure | The workflow owner and the function on duty | Stops, preserves the history, and hands over control. |
Cutting roles does not, by itself, demonstrate AI-native success. Reorganizations may have financial, strategic, or market causes. Stronger evidence is that decisions move closer to the work, unnecessary handoffs decline, and both decision quality and the ability to challenge decisions are preserved.
Organizational memory becomes part of production
A person may infer that a document is outdated or that a formal rule has an exception. A system often sees several equally convincing texts. Connecting the corporate knowledge base is therefore insufficient. Critical knowledge needs an owner, a source, a scope of application, a freshness date, and a rule for resolving conflicts. Useful knowledge must be available to the workflow authorized to use it.
Separate three layers. Sources of truth hold facts: contracts, orders, code, policies, and decisions. Task context contains the minimum relevant slice of those facts. Execution memory stores observations, errors, and approved improvements. An agent's answer must not automatically become a new fact; otherwise, the system starts reinforcing its own errors.
My article on platforms makes a broader point: expanding autonomy requires an environment in which work is reproducible and verifiable. In an engineering organization, this means tests, builds, and safe release mechanisms. In other functions, it means testable rules, action logs, restricted interfaces to business systems, and a way to recover from errors.
Wider autonomy → Reliable CI and platform → Durable outcome
The shared platform should remove recurring complexity: model access, actor identities, tool permissions, execution monitoring, budgets, and version rollout. The domain team retains the knowledge of what constitutes a good result. If every change has to wait in the central AI team's queue, the new platform recreates the organization's old constraint.
My essay on the changing interface of an internal platform draws the same boundary: an agent can become the interface, while the platform retains the domain meaning of an action. “Create a service with an owner, defined constraints, and a rollback path” is more useful than exposing a set of low-level operations. Connecting tools through a protocol is insufficient: the organization must define what an authorized action means.
The build-or-buy boundary depends on what makes the workflow distinctive. An off-the-shelf service is a reasonable choice for a standard function when its quality, data access, and operating terms are suitable. Build your own logic where rules, context, and feedback differentiate the business. Being able to export history, evaluations, and rules matters more than a promise to switch effortlessly between any models.
Autonomy is granted per action and earned through results
Anthropic distinguishes workflows with predefined transitions from agents that choose their next steps during execution. For a predictable task, the first option is often easier to control. Dynamic planning is useful when the path cannot be specified in advance, but it also expands the space of possible errors, duration, and cost. This is an engineering choice, not a contest to deploy the most autonomous agents.
Useful operating modes include preparing a recommendation; executing after approval; independently performing a bounded class of reversible actions; and escalating exceptions to a person. One workflow may use several modes. An agent may find a document without permission to send it externally, or prepare a change without permission to release it before review.
- Authority is enforced through tools: a prohibition in a prompt does not replace permission checks when an action runs.
- Results are checked independently of the executor's reasoning, using rules, tests, reliable data, and domain judgment.
- Time, steps, and spending have limits; there is a stop mechanism and a working handoff to a person.
- The system version, context used, actions taken, and person responsible for incident review are known.
This applies the logic of the NIST AI RMF: risk management belongs throughout a system's life cycle. The framework is voluntary; citing it does not demonstrate legal compliance or the safety of an implementation. A workflow needs its own tests of material consequences, including data leakage, incorrect decisions, malicious instructions embedded in documents, and service failure.
Human involvement helps only when the person has the time, context, and authority to verify a decision. Hundreds of repetitive approvals can turn oversight into a formality. Track review queues and errors discovered after approval, and retain independent technical safeguards for critical actions even when a reviewer is present.
People need new skills and an honest agreement on the purpose of change
“Describe everything you do so we can automate it” may sound threatening. In those circumstances, an organization risks getting performative tool use and an incomplete account of errors. Leaders need to explain the goal, measurement approach, expected role changes, and use of freed time in advance. Their promises must match the company's actual decisions.
At least three skill groups matter. Domain knowledge helps people spot an incorrect result. Task formulation establishes intent, context, criteria, and constraints. Working with evidence makes it possible to check sources and recognize when to stop. Fluency in a particular interface is useful but ages quickly; these capabilities survive changes in models.
Teach through real work and joint review: why an answer was accepted, where an expert was needed, and which rule was missing. Newcomers still need tasks that develop domain understanding. If all learning tasks disappear, the company must provide other ways to gain experience: failure analysis, independent work on assessment tasks, and mentoring. Otherwise, it loses the means to develop its future reviewers.
In a small Anthropic experiment, 52 developers learned an unfamiliar library. The average score on the subsequent test was 50% with AI and 67% without it; the time saving was not statistically significant. This was a short-term test of a specific skill, not evidence of inevitable deskilling. It supports the practical requirement to assess execution and understanding separately. I explore this further in my longread on developing junior engineers.
My article on AI-native leadership examines leadership through the organization of work. A practical implication is to assess employees by their contribution to accepted results, collaboration, and improvements to the shared process, rather than their number of model requests. A useful discovery should become a testable team practice instead of remaining an expert user's private technique.
The plan must account for the cost of learning: expert time spent reviewing, helping colleagues, and updating knowledge. If this is invisible work added on top of existing targets, the organization will systematically underestimate transition costs. The most capable people are particularly easy to overload by making them the final destination for every difficult exception.
Measure accepted results and the full cost of the workflow
My writing on AI-native measurement argues for multiple outcomes, metrics, and methods. Licenses indicate access; activity indicates use; subjective assessments describe employee experience. A decision to scale also needs time to outcome, quality, the consequences of errors, and total costs. Each measure answers a different question.
Measurement model: Multiple outcomes · Multiple metrics · Multiple methods
Establish the baseline and task mix first. Where possible, compare randomly assigned, comparable tasks or teams. Where that is not feasible, use a phased rollout and explicitly account for seasonality, changing complexity, experience, and concurrent reforms. A simple before-and-after comparison is useful for observation but weaker for causal inference. Employees should not select only easy cases for AI if the result will then be generalized to the entire workflow.
An illustrative calculation, not a promise of returns
Suppose the original total variable cost of one accepted result is 12 units. After redesign, it is 7, including the model, human review, retries, and exception handling. Additional fixed costs for the period are 2,400, with a volume of 1,000 comparable accepted results. The estimated benefit is then 1,000 × (12 − 7) − 2,400 = 2,600. The break-even volume for those fixed costs is 2,400 / (12 − 7) = 480 results over the same period.
Period benefit = accepted-result volume × reduction in total variable cost − additional fixed costs
At 300 results, the benefit becomes negative: 300 × 5 − 2,400 = −900. If review and exceptions raise the new variable cost to 10, the threshold rises to 1,200 results. In neither case does a cheaper model call automatically solve the problem. Fixed costs must also include the implementation and shared-platform costs allocated to the period, without counting them again as variable costs.
Freed paid hours represent available capacity, not automatic cash savings. To realize financial value, the company must use that capacity for additional useful output, avoid a justified future expense, or actually change its costs. The same hours cannot be counted as both cost savings and increased output without a separate explanation.
Averages are insufficient for rare, severe errors. Monitor prohibited actions, complaints, repeat contacts, rollbacks, and expert workload separately. Acceleration that breaches these constraints does not pass acceptance. An incident-free small pilot does not yet establish safety at scale.
The first 90 days: four decisions on whether to proceed
The following is a suggested management cadence, not a speed benchmark. Organizations with long cycles, sensitive data, or complex integrations will need more time. The purpose of each stage is to ground expansion in specific evidence. A well-supported decision to abandon a weak use case can also be a successful quarterly outcome.
| Period | Work and stage outcome | Evidence needed to proceed |
|---|---|---|
| Days 1–15 | Appoint an owner. Choose one workflow, establish the baseline, and examine real cases and exceptions. Define the goal and permitted actions. | A verifiable result, accessible data, and a clear customer exist; people capable of evaluating quality are assigned. |
| Days 16–30 | Redesign the workflow. Prepare evaluations, permissions, execution history, and human handoff. Test historical cases and run in observation mode. | The system passes material checks and performs no prohibited actions; recovery and handoff have been tested. |
| Days 31–60 | Launch with a limited group and a comparable control. Track accepted work, quality, full cost, and expert workload. | The benefit repeats on the required task mix; errors and their consequences remain within predefined boundaries. |
| Days 61–90 | Expand only the validated task class. Train the next team and establish ownership, costs, and regular reassessment. | Results hold beyond the enthusiast team; there is an evidence-based decision to scale, revise, or stop. |
Record stopping conditions before the pilot: which quality losses are unacceptable, what increase in review time would destroy the economics, and which actions must never happen automatically. Workflow owners set specific thresholds according to consequences; importing percentages from another domain is risky.
Test what happens when the model or an integration fails. Who sees the queue, can finish work already underway, and explains the delay to the customer? Workflow state must survive a change of executor. Without this, a pilot demonstrates answer quality but not the readiness of the company's operating system.
Change workflow → Measure delivery → Check quality → Adjust → Change workflow
Scale the ability to change workflows
After the first success, it is tempting to multiply agents across departments. It is more useful to identify reusable capabilities: data access, authorization, action logs, quality evaluation, and cost measurement. The next workflow should reuse these capabilities while retaining its own outcome criteria and domain experts.
Giving an experimental team temporary separation can help exploration, but decide in advance how its results will return to the wider organization. Evaluations, knowledge, and operating responsibilities must be transferable. Otherwise, two permanent worlds emerge: impressive, fast pilots and ordinary teams left with the integration work and consequences. My writing on the transition from PDLC addressed precisely this divide.
Over the next three to six months, test whether the approach transfers to a neighboring team and a more demanding task mix. If everything must be rebuilt each time, the company has acquired isolated automations. If a new team receives a working foundation and can still change its own domain rules, an organizational capability is emerging. This time frame is a planning guide, not a forecast of returns.
Different organizations need different paths
A small company can bring the workflow owner, expert, and developer together in one short cycle and buy most of the infrastructure. It should take particular care not to bury all its knowledge in the founder's personal conversation history. A large company usually needs agreements on data boundaries, shared interfaces, and cost allocation. A common platform with autonomous domain teams is useful in that setting.
Where errors are costly, the starting scope may be limited to finding facts and preparing a decision for a specialist. Where results are physical and feedback takes months, fast digital indicators are insufficient. Actual consequences must be observed. Such a workflow can use AI extensively while keeping action autonomy limited.
There is also an economic boundary: an infrequent operation with expensive integration and cheap manual execution may not be worth automating. Elsewhere, the rules may be stable enough for conventional software to be more reliable and less expensive. Choosing that solution is compatible with being AI-native: the organization should choose a method that fits the task, rather than demonstrate technological allegiance through every action.
The strategic question goes beyond cutting costs: what new level of service is now possible? Examples might include refreshing an offer more frequently, serving a previously uneconomic segment, or assessing a complex request quickly. These are hypotheses to test against demand and economics. If customers do not need the additional output, producing more documents and code will not create a market by itself.
I associate long-term advantage with accumulating validated knowledge about the organization's own workflows: which decisions work, where errors occur, who can receive delegated authority, and how results should be measured. This is my hypothesis. It would be challenged if general-purpose off-the-shelf systems consistently delivered the same results without the organization's domain context and accumulated feedback.
A readiness diagnostic: present evidence, not an aggregate score
Instead of a maturity index built from dozens of subjective questions, I suggest six checks. For each, choose one state: “no evidence,” “demonstrated in a pilot,” or “repeated in routine work.” Do not average them into a score: strong training cannot compensate for a critical workflow without an owner. Address the constraint that blocks the next step first.
| What to check | Evidence required |
|---|---|
| Value creation | An end-to-end customer or business metric, a baseline, a comparable assessment, and an explanation of why the result changed. |
| Work design | Before-and-after workflow maps, removed handoffs, changed decisions, and a functioning exception path. |
| Accountability | A named outcome owner, action boundaries, authority to stop execution, and a tested failure-review process. |
| Context and knowledge | Owned sources of truth, access controls, freshness, and an approved process for feeding experience back into the system. |
| People and quality | Practical learning, independent verification, preserved expertise, and visible reviewer workload. |
| Economics and transferability | The full cost of accepted work and repeated impact when another team handles ordinary tasks. |
Red flags are often visible without a complex audit: activity rises while customer outcomes stay flat; reviewers are overloaded; the new workflow handles only the happy path; no one owns outdated knowledge; a negative experiment cannot lead to a shutdown. Each signal should prompt a change in the workflow, not another presentation about the scale of adoption.
An AI-native organization is recognizable by its ability to change how it works deliberately. It can turn a new technical capability into a verified result, preserve accountability, and learn from consequences. A sensible first step is therefore to choose one important workflow, appoint an owner, and agree on which observation a quarter from now would support or challenge the value of the change.
What to bring into your organization
- 01AI-native describes an operating model: how a company turns intent into accepted work and learns from the results. An AI product, licenses, or agents alone do not establish that model.
- 02The unit of change is an end-to-end workflow with an owner, a verifiable result, and a known cost of failure. Speeding up one task matters when it shortens the whole path to the outcome without shifting work onto another team or the customer.
- 03Coordination is redistributed among people, the platform, and agents. A system can receive authority to act; named people remain accountable for the consequences and the rules of delegation.
- 04Context, quality evaluation, and people's development become production assets. They need funding, maintenance, and regular checks, just as infrastructure does.
- 05The transition is a series of testable changes. Expand autonomy and scale after demonstrating results; negative impact, difficult exceptions, and overloaded reviewers are reasons to redesign the process.
My earlier work and external evidence
Checked as of September 24, 2026. Research versions and limitations are identified. The practical examples, calculation, and transition plan are illustrations and my recommendations; they do not describe a particular client's results.
The author's foundation
- Is IDP Dead? No, the GUI Monopoly Is Dying — 9 June 2026: agent interfaces, the business meaning of actions, and platform accountability
- Junior Engineers After Code: Developing Engineers When Agents Do the Work — 18 August 2026: independent system ownership, checking understanding, and developing expertise
- From AI-Native Development to AI-Native Organization — 21 March 2026: work types, team boundaries, and autonomy's dependence on a shared platform; the author's framework
- From AI-Native Organization to AI-Native Leadership: The CTO Role in 2026 — 22 April 2026: leadership accountability, decision context, and a changing unit of management
- From AI-Native Development to AI-Native Platform — 15 April 2026: agent workloads and a reliable shared environment; an engineering example, not a universal organizational model
- From AI-Native Development to AI-Native Measurement — 16 April 2026: multiple outcomes, metrics, and methods; the source of the measurement and feedback diagrams
- From Classic PDLC to AI-Native Development — 11 March 2026: redesigning the workflow around AI; further reading on the engineering lifecycle
- AI-native SDLC · a practical analysis — 27 August 2026: shifting bottlenecks, verification, and delivery; further reading
- The Economics of AI Development: From Tokens to Accepted Work — 21 July 2026: the full cost of accepted outcomes and shared platform capabilities; further reading
Experiments and measurement
- Brynjolfsson, Li, Raymond · Generative AI at Work — QJE, 4 February 2025: the published version reports 5,172 agents and 15% more resolved issues per hour; one company, heterogeneous effects
- Dell’Acqua et al. · The Cybernetic Teammate — Organization Science, 12 June 2026: 791 P&G participants, 776 complete surveys; a one-day ideation exercise, not a test of replacing established teams
- Dell’Acqua et al. · Navigating the Jagged Technological Frontier — Organization Science, 11 March 2026: the final study of 758 consultants; effects vary by task, with correct solutions falling by 19 percentage points on the outside-frontier task
- METR · Early-2025 AI and Experienced Open-Source Developer Productivity — 10 July 2025: 16 developers, 246 tasks, and 19% more time; a narrow sample using early-2025 tools, to be read alongside the update
- METR · We are Changing our Developer Productivity Experiment Design — 24 February 2026: 57 developers and over 800 tasks; selection effects and concurrent agents prevent a reliable estimate of the new speedup
- METR · Self-Reported Impact of Early-2026 AI on Technical Worker Productivity — 11 May 2026: a survey of 349 workers distinguishes speed from value; self-reports, not measured output gains; further reading
- Anthropic · How AI Assistance Impacts the Formation of Coding Skills — 29 January 2026: an experiment with 52 engineers; an immediate test of a new library, not evidence of long-term skill loss; further reading
Surveys and observations
- McKinsey · The State of AI in 2026: On the Road to ROI — 25 August 2026: 1,719 respondents across 97 countries; individual benefits and EBIT contributions are self-reported, not financially audited
- Microsoft · Work Trend Index 2026: Agents, Human Agency, and Opportunity — 5 May 2026: 20,000 workers already using AI across 10 markets; organizational predictors' model importance is not their causal share of the effect
- Work Trend Index 2025 · The Year the Frontier Firm Is Born — Microsoft, 23 April 2025: 31,000 workers across 31 markets; the origin of the Frontier Firm framework, not an industry standard; further reading
- DORA · State of AI-assisted Software Development 2025 — 23 September 2025: nearly 5,000 professionals; official summaries reviewed; associations between throughput, stability, and system readiness do not establish causation
- Anthropic · Economic Index: Cadences — 26 June 2026: Claude telemetry and roughly 9,700 linked survey responses; delegation patterns are not productivity or occupational automation rates; further reading
Cases and practice
- Klarna · Annual Report 2025, Form 20-F — 26 February 2026: the company reports AI handling 80% of support chats while retaining a human channel; metrics and savings are Klarna's own estimates
- DORA · AI Capabilities Model — 23 September 2025: the official overview of seven technical and cultural capabilities; an empirical software-development framework, not a universal maturity scale
- Anthropic · Building Effective Agents — the engineering distinction between predefined workflows and agents; start simply and test whether autonomy is needed
- NIST · AI Risk Management Framework — governance, context mapping, measurement, and risk management; a voluntary framework, not AI-native organization certification
- Brynjolfsson, Rock, Syverson · The Productivity J-Curve — AEJ: Macroeconomics, January 2021; the abstract on complementary intangible investment; not a forecast of generative AI payback