Discussion about this post

User's avatar
Carlos Mattos's avatar

On the teams I’ve been working with, the time saved by using AI is being used to review more code rather than develop new features.

Developers are hesitating to approve and merge AI-generated PRs due to a lack of trust, even when the code has passed tests and follows the repository’s standards. This causes delays in the review process consuming work time that isn’t measured by current metrics.

Working with tier-1 banks and insurance companies, I learned that the pace of deliveries doesn’t depend only on developers' speed. There is significant portion of maintenance tasks involves legal requirements, audits, and rules from regulatory authorities that come with fixed deadlines and demands. When talking about regulated industries, delivering code faster will not reduce the amount of regulatory work because it is outside the scope of engineering.

One alternative would be to measure the average review time per approved PR as a productivity metric, along with the time spent searching for information.

Kevin Kasaei's avatar

Two things worth testing before treating the weak link as a fact about innovation.

First, the time-savings side is self-reported, and self-report on exactly this question has a measured bias. METR ran a randomised trial with 16 experienced developers working on their own repositories, screen recorded. They forecast 24% faster, came in 19% slower, and afterwards still reported a 20% speedup. If part of the 6.1 hours is belief rather than duration, then AI output explaining 63% of the variation is partly explaining variation in belief, and the failure to propagate into innovation ratio is what you would predict. The time was never freed.

Second, the ratio has a denominator problem. Maintenance load scales with merged change. Faros telemetry across 22,000 developers shows median time in review up 441.5% and incidents per pull request up 242.7% while merge rate per developer rose 16.2%. More output means more surface to maintain, so the ratio stays sticky even when absolute new-capability hours are rising. Reporting absolute new-capability hours next to the ratio would separate those two cases.

The 13% R squared is the most useful number in the piece.

No posts

Ready for more?