[2/3] What’s DAT? Measuring Development Productivity at Meta with Diff Authoring Time (Category Productivity)
[2/3] What’s DAT? Three Case Studies of Measuring Software Development Productivity at Meta with Diff Authoring Time (Category #Productivity)
Continuing my discussion of this whitepaper from Meta, whose activities are banned in Russia, let’s turn to validation of DAT, or Diff Authoring Time. The authors carried out several additional studies.
1. User-experience research
- They created a ground-truth dataset by recording developers’ actual work.
- They used random sampling to minimize bias.
- Average DAT accuracy exceeded 90% when compared with the recorded data.
2. A large-scale survey
- They compared DAT with developers’ own estimates for 968 unique diffs.
- They embedded surveys in Phabricator, the code-review tool, launching each immediately after a diff was completed.
3. Descriptive statistics
- DAT covers 87% of eligible diffs.
- It proved stable, using a mean winsorized at the 99th percentile for reporting.
- The authors also validated DAT against Time Spent by Diff, defined as “averages coding time in a given period by the number of diffs published in that period.”
4. Time-series visualization
- They visualized in detail how raw telemetry becomes DAT; the image will appear in the final post.
- They cross-validated the results with the developers who authored the changes.
The paper then describes 3 experiments.
1. Typed mocking in Hack The experiment introduced typing into mocking tools for Hack, Meta’s internally modified version of PHP. The authors migrated some mocks to typed versions, left others unchanged and compared DAT for diffs in different parts of the codebase. This linked a language feature to specific productivity measures:
- A 14% improvement in DAT, presented as the first quantitative evidence of typing’s productivity impact in an industrial setting.
- Statistical significance of p < 0.001 across all diff sizes.
2. Automatic memoization in the React compiler The authors extended React for automatic memoization, then compared DAT for diffs using manual and automatic memoization.
- They used a mixed-effects regression model to account for confounders.
- For nonrandomized data, they used Wasserstein distance to measure differences between the groups.
- They found a 33% improvement in DAT, a substantial efficiency gain from automatic memoization.
3. Evaluating code reuse I found this study the most interesting: the authors assessed cross-platform technology, in their case React. They needed counterfactual analysis to estimate hypothetical development time without code reuse. Their results were:
- More than a 50% improvement over development without reuse.
- Thousands of hours of DAT saved annually through code-reuse frameworks.
Using DAT in this way led to:
- Infrastructure teams moving toward a culture of A/B experiments.
- Data-informed decisions, with DAT supporting development planning and prioritization.
- Better alignment between product and infrastructure teams through shared metrics and experiments.
The authors plan to expand the approach in two directions:
- Horizontally, by adding development artifacts beyond diffs, including documents and tasks, to create a general framework for measuring activity time in experiments.
- Vertically, by supporting more IDEs and other tools, reducing reliance on heuristics in favor of more precise measurements.
#Engineering #Software #Bigtech #Productivity #Management #Leadership #Processes