Skip to content
April 16, 2026

From AI-Native Development to AI-Native Measurement

Why developer productivity becomes the key to evaluating process changes

/ AI-native Measurement 2026

Slide contents

  1. 1. From AI-Native Development to AI-Native Measurement

    Why developer productivity becomes the key to evaluating process changes

  2. 2. What This Article Is About

    Measurement stack for transition

    Part 1: local acceleration misleads.

    Part 2: Google research lessons.

    Part 3: no magic metric.

    Parts 4-6: measurement layer.

  3. 3. 01. No Measurement, Only Local Optimization

    Coding speed ≠ delivery improvement

  4. 4. AI-Native Needs Developer Productivity

    Local acceleration vs delivery impact

    AI changes platform, QA, CI/CD.

    Demos and local wins are insufficient.

    Delivery needs macro-statistics.

    Developer productivity names systemic impact.

  5. 5. DORA 2024–2025: AI Adoption ≠ Delivery Improvement

    Adoption rose; throughput and stability fell

    What's growing

    +25% AI adoption year over year.

    Better docs, quality, review speed.

    What's declining

    –1.5% throughput; –7.2% stability.

    Weak foundation = instability.

  6. 6. Narrow Metrics: Convenient but Dangerous

    AI proxies mislead

    Generated code ≠ delivery.

    Accepted suggestions ≠ quality.

    Agent PR count is not value delivery.

    Local speed without downstream becomes noise.

  7. 7. 02. Developer Productivity for Humans

    Google Research series — the best framework for evaluating development process changes

  8. 8. Productivity ≠ Conveyor Belt

    Development is complex creative work

    Measuring people: technology + sociology.

    Taylorism fails for knowledge work.

    AI removes routine, leaves risk.

    The human part gets harder.

  9. 9. Measure Developer Goals, Not Tools

    Tools change

    30 stable developer goals.

    Behavioral data + user sentiment.

    Not AI in review, but code quality.

    Ideal framing for agent era.

  10. 10. Every Data Source Lies Differently

    Google uses a measurement stack, not a single dashboard

    Instrumental Data

    Cross-tool logs: code, builds, tests, reviews.

    Dozens of tools → behavioral telemetry.

    Human Data

    EngSat + diary studies for flow/friction.

    Self-assessment validates telemetry.

  11. 11. Speed, Ease, Quality and Trust

    Three productivity dimensions plus trust

    Speed without quality — Acceleration returns as debt and instability.

    Change effects are delayed — A week after rollout proves nothing.

    Trust is central — Acceptance rate means little without trust.

  12. 12. 03. No Single 'Magic' Metric

    All models are wrong, but some are useful

  13. 13. Measuring Productivity = Building a Model

    Models are partial

    Selective models miss important effects.

    Use multiple outcomes, metrics, methods.

    No single indicator explains everything.

    That is domain reality.

  14. 14. 04. Measurement Layer for SE 2.0

    Measure effect, not tool adoption

  15. 15. Unit of Analysis: Process Change

    Hypothesis about speed, ease, quality

    Agentic PR loop.

    Accelerated build/test loop.

    AI in code review and documentation.

    Onboarding for shorter batches.

  16. 16. Five Principles of the Measurement Layer

    Many domain models, not one company KPI

    1. Process change — Measure effect, not tool adoption.

    2. Speed/ease/quality — Speed cannot be bought with friction.

    3. Logs + surveys — One data source always distorts.

  17. 17. Longitudinal Measurement and Trust

    Habits change slowly

    4. Longitudinal measurement

    AI-native habits are rebuilt slowly.

    Quarterly horizon and cohort comparison.

    5. Trust and human experience

    AI can raise cognitive load.

    The model needs a subjective layer.

  18. 18. 05. What to Measure in Practice

    Four feedback loops for a large company

  19. 19. Four Measurement Loops

    From engineer loop to organizational capability

    Engineer's inner loop — Iteration time, build latency, context search, focus.

    Team delivery loop — Lead time, PR cycle, CI stability, rework.

    Quality + capability — Tech debt, defects, onboarding, documentation quality.

  20. 20. 06. Conclusions for a Large Company

    Measurement layer — the mechanism for transferring practices from frontier to core

  21. 21. Without Measurement Frontier Becomes Case Theater

    Transfer proven impact

    Frontier shines: PRs, code, docs.

    Core misses what improves delivery.

    Separate stability damage from noise.

    Measurement turns frontier into capabilities.

  22. 22. AI-Native Without Measurement: Open Loop

    Productivity closes loop

    AI helps locally, breaks systems.

    Good models use many outcomes, metrics, methods.

    DORA: weak foundation → downstream disorder.

    Winners measure delivery-changing work.

  23. 23. References and Materials

    Developer productivity and delivery outcomes

    DORA AI research: dora.dev/research/

    SPACE framework: queue.acm.org/detail.cfm?id=3454124

    DX Core 4: getdx.com/research

    Google Engineering Productivity + Accelerate.

  24. 24. Thank You!

    AI-Native Measurement

    For more materials on this topic, visit the "Book Cube" channel — all links are collected there

    Alexander Polomodov, Technical Director & Fellow, T-Technologies

    @Book_Cube