Skip to content
back to the archive page
#AI

[2/3] Resolving Code Review Comments with Machine Learning (Category AI)

Continuing my discussion of Google’s whitepaper, here is the system’s evolution in more detail: V1) It began by generating suggested changes after reviewers submitted comments. V2) It evolved to suggest code changes to reviewers while they were still writing comments. V2 + IDE integration) User feedback led to better IDE integration, including views of conflicting changes and 3-way merges.

The resulting workflow and funnel are:

  • A reviewer writes a comment, and the model generates a fix on the fly.
  • The reviewer can accept or reject it. If accepted, the change author receives the comment with the suggested edit attached; otherwise, they receive only the comment.
  • The author can then click a button to apply the suggestion.

The goal and optimization metric were defined as follows:

A primary goal for any assistance tool is to increase productivity. One metric we use to gauge the positive impact of our assistant on productivity is acceptance rate, the fraction of all code-review comments that are resolved by the assistant; this measures, out of all (non-automated) comments left by human reviewers, what fraction received an ML-suggested edit that the author accepted and applied directly to their changelist.

The stage-by-stage statistics are much more interesting than simply saying that suggestions were applied for 7.5% of all comments:

Stage -- (%) of total -- (%) of previous step Incoming comments -- 100.0% -- 100.0% Confident predictions -- 49.0% -- 49.0% Accepted by reviewer -- 33.1% -- 63.6% Previewed by authora -- 10.7% -- 34.5% Applied by author -- 7.5% -- 69.5%

The preview step matters less here than in the first version. The note below the table explains:

The concept of author preview is less significant in V2. The author automatically sees a small preview and can “click-to-view” full suggested edits. This full view either shows the “Apply“ button or informs about an edit that requires a three-way merge. Almost all not-applied previews in V2 denote an edit that required a three-way merge to be applied.

The qualitative feedback was enthusiastic:

Early feedback about the assistant in internal message boards is enthusiastic, including characterizations such as “sorcery!”, “magic!”, “impressive”. Although the new version V2, in which suggested edits are presented as the reviewer is typing a comment, has only been deployed to 100% of the population for a relatively limited time, we have received delighted reports demonstrating that just the location and the initial sentiment of the reviewer’s comment can lead to helpful suggested edits, for both parties involved.

Beyond the results, the authors explain how they improved quality through model and data tuning:

  • Model tuning: fine-tuning DIDACTR, varying parameter count, lowering precision, tuning hyperparameters and adapting to programming languages. Finally, introducing previews for reviewers allowed them to lower the precision cutoff further.
  • Data tuning: restricting the offline evaluation dataset to changes with a single comment, training on “done” comments and training on synthetic tasks.

It is an interesting whitepaper describing Google’s approach and useful practical results. The final post links interesting figures from the paper.

#Software #AI #Engineering #Process #DevEx