What Makes a Good Productivity Metric? (Series #Management)
In “All Models Are Wrong,” I discussed why developer productivity requires several measures. I have now read “What Makes a Good Productivity Metric?” by Google's Collin Green, Ciera Jaspan, and Majed Samad, about validating an individual metric. In the margin, I marked their conclusion that schedule consistency may matter more than office attendance frequency. But that was precisely where I wanted to push back.
1️⃣ First, what are the authors actually measuring? Leaders wanted to know whether hybrid-work policy should address both attendance frequency and schedule consistency. These are different properties: two office days each week might be known in advance, or they might come as a surprise every time. The authors considered 13 metrics and selected Weekly Markov Consistency. It measures how predictable a person's status (office, home, or not working) is from their status on the same day the previous week. It is based on Markov entropy and normalized to a range of 0 to 1.
2️⃣ They propose four checks for a good metric 1) Its meaning can be explained to decision-makers. 2) It reflects expected patterns. 3) It has no properties that distort important cases. 4) It is useful for making decisions.
The most useful part comes next: testing it against examples. The researchers introduced random changes into consistent schedules, compared different hybrid arrangements, and examined seasonality. Eight participants rated 50 schedules; human ratings correlated positively with the metric. They also checked that it correlated only weakly with office attendance frequency.
In the margin, I called an understandable explanation a “shortcut for the brain.” A complex formula can have a clear meaning. But an easy explanation does not guarantee that we have understood the number correctly :)
3️⃣ Here is an example from my own examination of the formula Someone alternates a full working week in the office with a full week at home, keeping weekends unchanged. The metric returns 1: transitions are entirely predictable. Yet every working day's status changes relative to the previous week! Comparing days 28 days apart instead produces matching statuses, but the score remains the same. The metric captures predictable transitions, which do not require the schedule to repeat every week.
That alternation may be perfectly convenient for planning. It brings to mind my review of build latency and predictability from the same series: people need to understand how to organize the next part of their work.
4️⃣ What about the link to productivity? The authors report that engineers with consistent hybrid schedules involving 1–3 office days per week had CL throughput—the flow of code changes—indistinguishable from that of engineers consistently working the whole week in the office. The data cover tens of thousands of engineers, but the short article provides neither effect sizes nor sufficient analytical detail to assess a causal conclusion. Perhaps the organization of work itself makes both maintaining a schedule and delivering changes possible.
Moreover, throughput alone does not cover quality, learning, or helping colleagues. In my review of technical debt vanishing, I discussed why change counters can miss a useful effect.
The hypothesis that interests me here is that consistency helps through planning and coordination. But then we need to know whether colleagues knew the schedule in advance and whether they had overlapping time to work together. Two people can have perfectly consistent office days that never overlap. I would test the whole chain: the schedule was known in advance → the necessary interaction happened → waiting and broken arrangements decreased. That is a research proposal, not a finding of the article. If an agreed but changing schedule helps just as much as a fixed one, what exactly should company policy require?
#Management #Engineering #Productivity #Research #Metrics
Files from the post
- Annotated-What-Makes-a-Good-Productivity-Metric.pdfDownload PDF