From AI-Native Development to AI-Native Measurement
Why developer productivity becomes the key to evaluating process changes
Why developer productivity becomes the key to evaluating process changes
Why developer productivity becomes the key to evaluating process changes
Measurement stack for transition
Part 1: local acceleration misleads.
Part 2: Google research lessons.
Part 3: no magic metric.
Parts 4-6: measurement layer.
Coding speed ≠ delivery improvement
Local acceleration vs delivery impact
AI changes platform, QA, CI/CD.
Demos and local wins are insufficient.
Delivery needs macro-statistics.
Developer productivity names systemic impact.
Adoption rose; throughput and stability fell
What's growing
+25% AI adoption year over year.
Better docs, quality, review speed.
What's declining
–1.5% throughput; –7.2% stability.
Weak foundation = instability.
AI proxies mislead
Generated code ≠ delivery.
Accepted suggestions ≠ quality.
Agent PR count is not value delivery.
Local speed without downstream becomes noise.
Google Research series — the best framework for evaluating development process changes
Development is complex creative work
Measuring people: technology + sociology.
Taylorism fails for knowledge work.
AI removes routine, leaves risk.
The human part gets harder.
Tools change
30 stable developer goals.
Behavioral data + user sentiment.
Not AI in review, but code quality.
Ideal framing for agent era.
Google uses a measurement stack, not a single dashboard
Instrumental Data
Cross-tool logs: code, builds, tests, reviews.
Dozens of tools → behavioral telemetry.
Human Data
EngSat + diary studies for flow/friction.
Self-assessment validates telemetry.
Three productivity dimensions plus trust
Speed without quality — Acceleration returns as debt and instability.
Change effects are delayed — A week after rollout proves nothing.
Trust is central — Acceptance rate means little without trust.
All models are wrong, but some are useful
Models are partial
Selective models miss important effects.
Use multiple outcomes, metrics, methods.
No single indicator explains everything.
That is domain reality.
Measure effect, not tool adoption
Hypothesis about speed, ease, quality
Agentic PR loop.
Accelerated build/test loop.
AI in code review and documentation.
Onboarding for shorter batches.
Many domain models, not one company KPI
1. Process change — Measure effect, not tool adoption.
2. Speed/ease/quality — Speed cannot be bought with friction.
3. Logs + surveys — One data source always distorts.
Habits change slowly
4. Longitudinal measurement
AI-native habits are rebuilt slowly.
Quarterly horizon and cohort comparison.
5. Trust and human experience
AI can raise cognitive load.
The model needs a subjective layer.
Four feedback loops for a large company
From engineer loop to organizational capability
Engineer's inner loop — Iteration time, build latency, context search, focus.
Team delivery loop — Lead time, PR cycle, CI stability, rework.
Quality + capability — Tech debt, defects, onboarding, documentation quality.
Measurement layer — the mechanism for transferring practices from frontier to core
Transfer proven impact
Frontier shines: PRs, code, docs.
Core misses what improves delivery.
Separate stability damage from noise.
Measurement turns frontier into capabilities.
Productivity closes loop
AI helps locally, breaks systems.
Good models use many outcomes, metrics, methods.
DORA: weak foundation → downstream disorder.
Winners measure delivery-changing work.
Developer productivity and delivery outcomes
DORA AI research: dora.dev/research/
SPACE framework: queue.acm.org/detail.cfm?id=3454124
DX Core 4: getdx.com/research
Google Engineering Productivity + Accelerate.
AI-Native Measurement
For more materials on this topic, visit the "Book Cube" channel — all links are collected there
Alexander Polomodov, Technical Director & Fellow, T-Technologies
@Book_Cube