Engineering Enablement
Engineering Enablement by DX
How Meta reduced diff authoring time by 40%
0:00
-46:55

How Meta reduced diff authoring time by 40%

Meta researcher Moritz Beller shares how Meta measures developer productivity and how AI is changing engineers’ work, metrics, and testing.

Listen and watch now on YouTube, Apple, and Spotify.

In this episode of Engineering Enablement, I talk with Meta researcher Moritz Beller about how Meta measures developer productivity. Moritz explains diff authoring time and how the company uses it to evaluate tools and guide engineering decisions. We explore how AI is changing engineers’ work and revealing gaps in traditional metrics. We also revisit his highly influential “Mind the Gap” study and discuss the promise and risks of AI agents for testing.

Some takeaways:

Diff authoring time enables more precise productivity measurement

  • Diff authoring time measures the active work involved in creating, testing, and reviewing a specific code change. This gives Meta more precision than metrics averaged across an entire day.

  • No single diff metric provides a complete picture of productivity. Meta considers authoring time alongside throughput, diff size, and quality to make changes in the data easier to interpret.

  • Meta does not use diff authoring time to evaluate individual developers. It is used at the team, organizational, and company levels to identify trends, regressions, and opportunities for improvement.

A/B experiments help Meta prioritize engineering investments

  • Diff-level measurement allows Meta to quantify whether changes to developer tools and frameworks save engineers time. This helps leaders decide which improvements deserve further investment.

  • Automatic memoization in the React compiler reduced diff authoring time by roughly 30% compared with implementing caching manually. Results of that magnitude suggest that foundational improvements can be more valuable than small interface optimizations.

AI is changing how engineers spend their time

  • Meta’s year-over-year diff authoring time has fallen by more than 40%, while developers are producing more and larger diffs. Moritz views the company-wide pattern as a strong signal of AI’s impact.

  • It is not yet clear where all the saved authoring time goes. Moritz believes engineers are spending more time gathering context and performing work adjacent to coding, while relatively little time goes into writing prompts.

  • As code becomes cheaper to generate, intent becomes more valuable. The important question is increasingly whether the finished code accurately reflects what the developer intended to build.

Traditional telemetry misses important engineering work

  • Activities like whiteboarding, brainstorming, and architectural alignment are difficult to connect to code changes through conventional telemetry. Meeting transcripts and other AI-generated records could help make that work more visible.

  • Cheaper code generation does not eliminate the need for architecture. Skipping a design document can simply transfer the cost to reviewers, who must reconstruct the intended design from the implementation.

Developer productivity is both measured and perceived

  • Moritz’s “Mind the Gap” study connected automatically measured activity with developers’ perceptions of their own productivity. Time spent coding was an important predictor, but sleep, interruptions, and on-call responsibilities also mattered.

  • A modern version of the study would need to account for agentic work. Moritz would examine how developers interact with agents, how much they trust them, and whether managing parallel work creates cognitive overload.

  • Higher output does not necessarily improve developer experience. Brian points to research showing that output can rise while flow, cognitive load, and overall experience remain flat or worsen.

AI makes testing easier but not necessarily safer

  • Agents make it inexpensive to generate unit and end-to-end tests. This could fill gaps left by developers who previously did little testing.

  • Passing tests can create false confidence when both the implementation and tests reflect the same misunderstanding. The agent may satisfy its interpretation of an underspecified request without delivering what the developer actually intended.

  • Specification-driven development could become a new form of test-driven development. Defining correctness before implementation may matter more as agents take on more of the coding.

Collaboration remains one of the hardest parts of productivity to measure

  • Interpersonal dynamics, alignment, and knowledge sharing are central to engineering productivity but poorly captured by existing metrics. Faster implementation also makes it easier for teams to unknowingly duplicate one another’s work.

  • Reducing low-quality meetings can produce major throughput gains, but eliminating collaboration creates different problems. Teams still need enough interaction to generate ideas, share context, and stay aligned.

  • Small human interactions can have measurable value. Brian’s research found that informal conversation before and after meetings predicted self-reported productivity better than internet quality.

In this episode, we cover:

(00:00) Intro

(02:10) Moritz’s role at Meta

(04:01) Measuring diff authoring time

(08:20) Measuring A/B experiments

(12:11) What diff authoring time reveals

(14:45) Planning and leadership reporting

(18:36) Why Meta measures teams, not individuals

(22:33) Where developer time goes with AI

(26:06) AI’s impact on junior and senior developers

(27:18) Capturing invisible work with AI

(29:32) The “Mind the Gap” study

(33:40) Revisiting “Mind the Gap” in 2026

(38:32) AI agents for software testing

(41:49) What remains hard to measure

Where to find Moritz Beller:

• LinkedIn: https://www.linkedin.com/in/inventitech

• X: https://x.com/Inventitech

• GitHub: https://github.com/Inventitech

• Website: https://inventitech.com

Where to find Brian Houck:

• LinkedIn: https://www.linkedin.com/in/brianhouck

Referenced:

• State of AI Impact in Engineering Q2 Report 2026

• GitHub Copilot and Developer Productivity: An Observational Dose-Response Analysis

• From Technical Debt to Cognitive and Intent Debt

• Mind the Gap: On the Relationship Between Automatically Measured and Self-Reported Productivity

• The Impact of AI Coding Assistants on Software Engineering: A Longitudinal Study

• When, how, and why developers (do not) test in their IDEs | Proceedings of the 2015 10th Joint Meeting on Foundations of Software Engineering

• DORA, SPACE, and DevEx: Which framework should you use

• A Tale of Two Cities: Software Developers Working from Home During the COVID-19 Pandemic

Discussion about this episode

User's avatar

Ready for more?