Eighteen months of trying to say something specific about assisted coding, and the same conclusion as the first attempt.
what can be measured:
review findings by category and origin
time on a comparable task
defect rate, eventually, with no signal yet
what cannot:
the counterfactual. what would have been written
otherwise, by the same person, on the same day.
and the numbers that have held:
convention findings ~5× higher on generated code
test-writing ~40% faster, with a lower mutation
score on the resulting tests
defect rate: eight quarters, no trend
Two years and the honest position is that the tool is measurably faster at one task type, measurably worse at conventions, and unmeasurable in aggregate. Anybody quoting a percentage improvement in engineering productivity is measuring something else and has not said what.