Two years in, and the measurement problem is unchanged

Eighteen months of trying to say something specific about assisted coding, and the same conclusion as the first attempt.

what can be measured:
  review findings by category and origin
  time on a comparable task
  defect rate, eventually, with no signal yet

what cannot:
  the counterfactual. what would have been written
  otherwise, by the same person, on the same day.

and the numbers that have held:
  convention findings ~5× higher on generated code
  test-writing ~40% faster, with a lower mutation
    score on the resulting tests
  defect rate: eight quarters, no trend

Two years and the honest position is that the tool is measurably faster at one task type, measurably worse at conventions, and unmeasurable in aggregate. Anybody quoting a percentage improvement in engineering productivity is measuring something else and has not said what.