<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Engineering Enablement]]></title><description><![CDATA[Research and perspectives on developer productivity. ]]></description><link>https://newsletter.getdx.com</link><image><url>https://substackcdn.com/image/fetch/$s_!Niij!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7dbd433b-6f11-4042-8b7d-0edb3b172966_1024x1024.png</url><title>Engineering Enablement</title><link>https://newsletter.getdx.com</link></image><generator>Substack</generator><lastBuildDate>Wed, 12 Aug 2026 13:13:52 GMT</lastBuildDate><atom:link href="https://newsletter.getdx.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Abi Noda]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[abinoda@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[abinoda@substack.com]]></itunes:email><itunes:name><![CDATA[Abi Noda]]></itunes:name></itunes:owner><itunes:author><![CDATA[Abi Noda]]></itunes:author><googleplay:owner><![CDATA[abinoda@substack.com]]></googleplay:owner><googleplay:email><![CDATA[abinoda@substack.com]]></googleplay:email><googleplay:author><![CDATA[Abi Noda]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[DX Annual 2027: San Francisco and London]]></title><description><![CDATA[Next year DX Annual is returning to San Francisco and heading to London for the first time. Here's what to expect and how to register your interest.]]></description><link>https://newsletter.getdx.com/p/dx-annual-2027-san-francisco-and</link><guid isPermaLink="false">https://newsletter.getdx.com/p/dx-annual-2027-san-francisco-and</guid><dc:creator><![CDATA[Justin Reock]]></dc:creator><pubDate>Tue, 11 Aug 2026 15:03:28 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/384d549a-7522-4640-bef4-1bcfdfe3a768_2400x1260.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!dys8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d360bc4-ecc3-4488-8542-48adb0705fba_3240x1020.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!dys8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d360bc4-ecc3-4488-8542-48adb0705fba_3240x1020.jpeg 424w, https://substackcdn.com/image/fetch/$s_!dys8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d360bc4-ecc3-4488-8542-48adb0705fba_3240x1020.jpeg 848w, https://substackcdn.com/image/fetch/$s_!dys8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d360bc4-ecc3-4488-8542-48adb0705fba_3240x1020.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!dys8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d360bc4-ecc3-4488-8542-48adb0705fba_3240x1020.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!dys8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d360bc4-ecc3-4488-8542-48adb0705fba_3240x1020.jpeg" width="1456" height="458" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0d360bc4-ecc3-4488-8542-48adb0705fba_3240x1020.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:458,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2226212,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://newsletter.getdx.com/i/210621976?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d360bc4-ecc3-4488-8542-48adb0705fba_3240x1020.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!dys8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d360bc4-ecc3-4488-8542-48adb0705fba_3240x1020.jpeg 424w, https://substackcdn.com/image/fetch/$s_!dys8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d360bc4-ecc3-4488-8542-48adb0705fba_3240x1020.jpeg 848w, https://substackcdn.com/image/fetch/$s_!dys8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d360bc4-ecc3-4488-8542-48adb0705fba_3240x1020.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!dys8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d360bc4-ecc3-4488-8542-48adb0705fba_3240x1020.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Building on this year&#8217;s </span><a href="https://getdx.com/dxannual/2026/"><span>inaugural event</span></a><span>, which brought together 500 senior engineering leaders from companies like Microsoft, Airbnb, Uber, Vanguard, Dell, and BNY, we&#8217;re expanding to two cities for 2027.</span></p><p><span>In both San Francisco and London, DX Annual will be a single-day event with keynotes, fireside chats, and peer roundtables built around how organizations are rethinking developer productivity and how software is delivered with AI. The event will once again focus on curated attendance, meaningful peer connections, and sessions led by practitioners.</span></p><p><strong><span>Save the dates &#8595;</span></strong></p><ul><li><p><span>San Francisco &#8212; May 6, 2027</span></p></li><li><p><span>London &#8212; September 23, 2027</span></p></li></ul><div class="callout-block" data-callout="true"><p><span>Attendance is curated and space will be limited. Register your interest at </span><a href="https://getdx.com/dxannual/"><span>dxannual.com</span></a><span> and we&#8217;ll be in touch with details.</span></p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://getdx.com/dxannual/&quot;,&quot;text&quot;:&quot;Register your interest&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://getdx.com/dxannual/"><span>Register your interest</span></a></p><div><hr></div><div class="image-gallery-embed" data-attrs="{&quot;gallery&quot;:{&quot;images&quot;:[{&quot;type&quot;:&quot;image/jpeg&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/df0be96f-10d1-46c7-9565-61de76b7c202_2738x1825.jpeg&quot;},{&quot;type&quot;:&quot;image/jpeg&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e3e8a85b-af79-412e-8834-a09d63aa7844_2738x1825.jpeg&quot;},{&quot;type&quot;:&quot;image/jpeg&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c68a2197-ae33-44ac-b5ff-649592780c61_2738x1825.jpeg&quot;},{&quot;type&quot;:&quot;image/jpeg&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/066d86a3-1b7f-4632-abab-81b606e1a520_2738x1825.jpeg&quot;},{&quot;type&quot;:&quot;image/jpeg&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/925dbcb9-703b-4d06-bc4f-a22e1f263ba3_2738x1825.jpeg&quot;},{&quot;type&quot;:&quot;image/jpeg&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/67c69cf7-ff87-475b-b213-8f46c4dde7a9_2738x1825.jpeg&quot;}],&quot;caption&quot;:&quot;&quot;,&quot;alt&quot;:&quot;&quot;,&quot;staticGalleryImage&quot;:{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4f1cc3cf-be1a-4284-b9cd-5d0d18910665_1456x964.png&quot;}},&quot;isEditorNode&quot;:true}"></div>]]></content:encoded></item><item><title><![CDATA[How Microsoft sees engineering bottlenecks changing with AI]]></title><description><![CDATA[Tim Bozarth, Microsoft CoreAI CVP, explains how AI is changing engineering productivity, why outcomes matter more than output, and what engineering leaders should measure instead.]]></description><link>https://newsletter.getdx.com/p/how-microsoft-sees-engineering-bottlenecks</link><guid isPermaLink="false">https://newsletter.getdx.com/p/how-microsoft-sees-engineering-bottlenecks</guid><dc:creator><![CDATA[Brian Houck]]></dc:creator><pubDate>Fri, 07 Aug 2026 15:47:36 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/209296274/1135b2a44ce3e5af0cd78de10f17d202.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>Listen and watch now on <strong><a href="https://youtu.be/fdofQkhpgDE">YouTube</a>, <a href="https://podcasts.apple.com/us/podcast/engineering-enablement-by-abi-noda/id1619140476">Apple</a>, and <a href="https://open.spotify.com/show/3NxjyIsuxeDMQtisDqBy7D">Spotify</a></strong>.</p><p><span>In this episode of Engineering Enablement, I sit down with Tim Bozarth, Corporate Vice President in Microsoft CoreAI and leader of Microsoft&#8217;s Engineering Thrive initiative. We discuss Engineering Thrive, Microsoft&#8217;s framework for measuring and improving engineering productivity, and why AI makes outcome-based metrics more important than ever.</span></p><p><span>Tim shares how AI is reshaping the software development lifecycle, where new bottlenecks are emerging, and why verification and confidence may become more valuable than code generation itself. We also explore why the purpose of engineering remains the same despite rising levels of abstraction, the skills that remain durable in an AI-driven world, and why engineering leaders should focus on outcomes rather than activity metrics.</span></p><div id="youtube2-fdofQkhpgDE" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;fdofQkhpgDE&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/fdofQkhpgDE?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h2><strong>Some takeaways: </strong></h2><h3><strong><span>Engineering Thrive gives Microsoft a common language for improving productivity</span></strong></h3><ul><li><p><strong><span>Engineering Thrive defines productivity through speed, ease, and quality.</span></strong><span> Its goal is to make it fast and easy to do great work by identifying friction across the developer experience rather than optimizing isolated activities.</span></p></li><li><p><strong><span>Speed, ease, and quality should not be treated as opposing goals.</span></strong><span> Engineering Thrive applies a &#8220;do no harm&#8221; principle: an improvement in one dimension should not be celebrated if it makes another materially worse.</span></p></li><li><p><strong><span>The framework helps Microsoft identify bottlenecks and invest where they will have the greatest impact.</span></strong><span> Tim argues that developer time is one of the company&#8217;s most valuable resources, making productivity improvements a strategic investment.</span></p></li></ul><h3><strong><span>AI makes outcome-based metrics more important, not obsolete</span></strong></h3><ul><li><p><strong><span>The industry is returning to activity metrics it rejected years ago.</span></strong><span> Tim sees renewed attention to lines of code, pull request counts, and similar measures as a search for easy answers to how AI is affecting productivity.</span></p></li><li><p><strong><span>More engineering activity does not necessarily mean more value.</span></strong><span> Engineering Thrive instead measures outcomes such as product quality, end-to-end speed, and the amount of time engineers can devote to innovation.</span></p></li><li><p><strong><span>A web of outcome metrics is harder to game than any individual measure.</span></strong><span> Looking across speed, ease, quality, and innovation time creates a more durable picture of whether an organization is actually improving.</span></p></li></ul><h3><strong><span>Planning and validation are becoming the new bottlenecks</span></strong></h3><ul><li><p><strong><span>Before AI, most engineering time was spent creating and operating software.</span></strong><span> Tim estimates that those phases historically accounted for more than 90% of engineering time, with operations alone consuming roughly 70% to 80%.</span></p></li><li><p><strong><span>On frontier teams, the bottlenecks have already shifted to planning and validation.</span></strong><span> Code creation is taking dramatically less time, while deciding what to build and determining whether the result is trustworthy are consuming a greater share of the work.</span></p></li><li><p><strong><span>The create phase may continue to shrink as models and agent harnesses improve.</span></strong><span> Tim expects planning and validation to remain durable constraints over the next several years, even as other parts of the software development lifecycle become increasingly automated.</span></p></li></ul><h3><strong><span>The code review bottleneck is really a confidence bottleneck</span></strong></h3><ul><li><p><strong><span>Higher PR throughput has increased the amount of change humans must evaluate.</span></strong><span> Enterprise software still requires a level of trust that teams cannot achieve by simply accepting AI-generated code without review.</span></p></li><li><p><strong><span>Code review is only one way to establish confidence.</span></strong><span> Testing, continuous rollouts, feature flags, canaries, and deployment practices all contribute signals about whether a change is reliable.</span></p></li><li><p><strong><span>The next generation of verification should help humans ask higher-level questions.</span></strong><span> Rather than inspecting every branch or class, engineers should be able to assess the scope, complexity, impact, and fundamental purpose of a change.</span></p></li></ul><h3><strong><span>More abstraction does not change the purpose of engineering</span></strong></h3><ul><li><p><strong><span>AI will reduce the time engineers spend on low-level implementation details.</span></strong><span> Tim sees that as another step in the long history of abstractions that allow engineers to spend more time describing system behavior and intended outcomes.</span></p></li><li><p><strong><span>Great engineers have never been defined by their ability to write a line of code.</span></strong><span> Their value comes from systems thinking, understanding objectives, and expressing functional and nonfunctional requirements coherently.</span></p></li><li><p><strong><span>Producing more software increases the need for strong engineering judgment.</span></strong><span> The faster organizations can build, the more important it becomes to ensure that their systems remain coherent, valid, and reliable.</span></p></li></ul><h3><strong><span>A maker&#8217;s mindset and ability to experiment effectively remain durable advantages</span></strong></h3><ul><li><p><strong><span>The maker&#8217;s mindset is relentlessly oriented toward producing a valuable final outcome.</span></strong><span> It combines a clear vision of what should be built with the ability to understand the needs of the customer using it.</span></p></li><li><p><strong><span>Engineers will need to continuously experiment and evaluate changes in outcomes.</span></strong><span> Tim argues that this discipline is no longer limited to model trainers; everyone using AI tools must determine whether a new approach actually improves the result.</span></p></li><li><p><strong><span>Using AI is not itself the goal.</span></strong><span> The goal is to accomplish more, and sometimes the best way to do that will be not to use AI at all.</span></p></li></ul><h3><strong><span>Engineering leaders should measure idea-to-value, not PR velocity</span></strong></h3><ul><li><p><strong><span>PR velocity says little about the success of an engineering organization.</span></strong><span> Tim urges leaders to stop evaluating teams through activity measures the industry already recognized as inadequate years ago.</span></p></li><li><p><strong><span>Idea-to-value connects engineering work to meaningful outcomes.</span></strong><span> Leaders should also examine how much time engineers spend innovating compared with keeping the lights on or handling corporate overhead.</span></p></li><li><p><strong><span>Product quality and business results should remain the ultimate measures of success.</span></strong><span> Focusing on outcomes can improve developer happiness, productivity, margins, and the value delivered by the organization.</span></p></li></ul><h2><strong>In this episode, we cover:</strong></h2><p>(<a href="https://www.youtube.com/watch?v=fdofQkhpgDE">00:00</a>) Intro</p><p>(<a href="https://www.youtube.com/watch?v=fdofQkhpgDE&amp;t=133s">02:13</a>) What Engineering Thrive is and the problem it solves</p><p>(<a href="https://www.youtube.com/watch?v=fdofQkhpgDE&amp;t=286s">04:46</a>) Why Engineering Thrive isn&#8217;t specific to Microsoft</p><p>(<a href="https://www.youtube.com/watch?v=fdofQkhpgDE&amp;t=550s">09:10</a>) The impact of Engineering Thrive at Microsoft</p><p>(<a href="https://www.youtube.com/watch?v=fdofQkhpgDE&amp;t=871s">14:31</a>) Why AI makes outcome-based productivity metrics more important</p><p>(<a href="https://www.youtube.com/watch?v=fdofQkhpgDE&amp;t=1102s">18:22</a>) Where AI is creating new bottlenecks in the SDLC</p><p>(<a href="https://www.youtube.com/watch?v=fdofQkhpgDE&amp;t=1477s">24:37</a>) Why more abstraction doesn&#8217;t change the purpose of engineering</p><p>(<a href="https://www.youtube.com/watch?v=fdofQkhpgDE&amp;t=1645s">27:25</a>) The durable skills of good engineers</p><p>(<a href="https://www.youtube.com/watch?v=fdofQkhpgDE&amp;t=1983s">33:03</a>) The changing economics of software development</p><p>(<a href="https://www.youtube.com/watch?v=fdofQkhpgDE&amp;t=2216s">36:56</a>) Advice for leaders: measure outcomes, not activity</p><p><strong><span>Where to find Tim Bozarth:</span></strong></p><p><span>&#8226; LinkedIn: </span><a href="http://linkedin.com/in/tbozarth"><span>linkedin.com/in/tbozarth</span></a></p><p><strong><span>Where to find Brian Houck:</span></strong></p><p><span>&#8226; LinkedIn: </span><a href="https://www.linkedin.com/in/brianhouck"><span>https://www.linkedin.com/in/brianhouck</span></a></p><h2><strong>Referenced:</strong></h2><p><span>&#8226; </span><a href="https://getdx.com/corefour"><span>DX Core 4 Productivity Framework</span></a></p><p><span>&#8226; </span><a href="https://spawn-queue.acm.org/doi/10.1145/3819080"><span>EngThrive: Make It Fast and Easy to Do Great Work: Building a durable model for outcome-oriented engineering measurement</span></a></p><p><span>&#8226; </span><a href="https://www.goodreads.com/quotes/10278659-tell-me-how-you-measure-me-and-i-ll-tell-you"><span>Quote by Eliyahu M. Goldratt: &#8220;Tell me how you measure me and...&#8221;</span></a></p><p><span>&#8226; </span><a href="https://en.wikipedia.org/wiki/Jevons_paradox"><span>Jevons paradox - Wikipedia</span></a></p>]]></content:encoded></item><item><title><![CDATA[What are code reviews even for?]]></title><description><![CDATA[AI didn't break code review. It just made the parts we'd been ignoring impossible to ignore.]]></description><link>https://newsletter.getdx.com/p/what-are-code-reviews-even-for</link><guid isPermaLink="false">https://newsletter.getdx.com/p/what-are-code-reviews-even-for</guid><dc:creator><![CDATA[Brian Houck]]></dc:creator><pubDate>Wed, 05 Aug 2026 10:40:18 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f4489e75-06a9-44a8-98bb-ea29694f42e9_2400x1260.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em><strong>Welcome to the latest issue of Engineering Enablement,</strong><span> a weekly newsletter sharing research and perspectives on developer productivity.</span></em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.getdx.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.getdx.com/subscribe?"><span>Subscribe now</span></a></p><p>DX&#8217;s Q2 AI Impact Report is now available with the latest research on AI&#8217;s impact across engineering organizations. <a href="https://getdx.com/resources/?utm_source=newsletter">Read the full report.</a></p><div><hr></div><p><span>Something is straining in the review queue.</span></p><p><a href="https://arxiv.org/abs/2605.30208"><span>Over the past year at Meta</span></a><span>, significant lines of code per human-landed diff increased by 106%. Diffs per developer per month rose 51%. More than 80% of that growth came from agentic AI. Meanwhile, the percentage of diffs reviewed within 24 hours is declining. In some large groups, reviewers are staring down thousands of pending reviews.</span></p><p><span>This isn&#8217;t a Meta-specific problem. Across the industry, AI coding tools are producing code faster than humans can meaningfully evaluate it. </span><a href="https://newsletter.getdx.com/p/ai-authored-code-has-nearly-doubled"><span>Our own DX analysis</span></a><span> found that AI is increasing both the number of pull requests and the size of each one (median pull request size grew by 64%). Without a corresponding increase in reviewer capacity, the review process will eventually buckle under its own weight.</span></p><p><span>Unfortunately, we don&#8217;t have more hours in the day, and even if we did, we wouldn&#8217;t want to spend them reviewing code written by AI. </span><a href="https://www.microsoft.com/en-us/research/publication/time-warp-the-gap-between-developers-ideal-vs-actual-workweeks-in-an-ai-driven-era/?msockid=15af30cf0f0662a5037e27800ec7634a"><span>In previous research</span></a><span>, we found developers ideally only want to spend about 7% of their time reviewing code. Asking developers to review more isn&#8217;t a sustainable answer.</span></p><p><span>The math doesn&#8217;t work.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Mu5k!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5710e33-dbf6-468c-8f44-b7159744f561_4200x2884.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Mu5k!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5710e33-dbf6-468c-8f44-b7159744f561_4200x2884.png 424w, https://substackcdn.com/image/fetch/$s_!Mu5k!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5710e33-dbf6-468c-8f44-b7159744f561_4200x2884.png 848w, https://substackcdn.com/image/fetch/$s_!Mu5k!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5710e33-dbf6-468c-8f44-b7159744f561_4200x2884.png 1272w, https://substackcdn.com/image/fetch/$s_!Mu5k!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5710e33-dbf6-468c-8f44-b7159744f561_4200x2884.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Mu5k!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5710e33-dbf6-468c-8f44-b7159744f561_4200x2884.png" width="1456" height="1000" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c5710e33-dbf6-468c-8f44-b7159744f561_4200x2884.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1000,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:225359,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.getdx.com/i/204342182?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5710e33-dbf6-468c-8f44-b7159744f561_4200x2884.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Mu5k!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5710e33-dbf6-468c-8f44-b7159744f561_4200x2884.png 424w, https://substackcdn.com/image/fetch/$s_!Mu5k!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5710e33-dbf6-468c-8f44-b7159744f561_4200x2884.png 848w, https://substackcdn.com/image/fetch/$s_!Mu5k!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5710e33-dbf6-468c-8f44-b7159744f561_4200x2884.png 1272w, https://substackcdn.com/image/fetch/$s_!Mu5k!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5710e33-dbf6-468c-8f44-b7159744f561_4200x2884.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>But before we ask AI to solve this problem, it&#8217;s worth asking a different question:</span></p><blockquote><p><strong><span>What problem was code review solving before AI arrived?</span></strong></p></blockquote><p><span>If the answer were as simple as &#8220;finding defects,&#8221; then a fully automated review starts to sound inevitable (and appealing).</span></p><p><span>But if code review was also how teams shared knowledge, built collective ownership, spread architectural understanding, and taught junior engineers how experienced developers think, then the answer becomes much less obvious.</span></p><p><span>That&#8217;s the mistake I think many organizations are about to make, and the reason we need to rethink what code review is actually for.</span></p><h3><span>We&#8217;ve known better for years</span></h3><p><span>Here&#8217;s an uncomfortable truth: a significant portion of the review burden we&#8217;re feeling right now is self-inflicted.</span></p><p><span>The frustrating part is that none of this is new. The research on what makes code review effective has been unambiguous. Keep changes small. Write a meaningful description of what changed and why. Run automated checks before asking a human to look. Review frequently, in bounded sessions, focused on substance over style. Select reviewers who actually know the code, but avoid concentrating review responsibility on the same small group of experts whenever possible.</span></p><p><a href="https://ieeexplore.ieee.org/document/7950877"><span>A 2016 Microsoft study</span></a><span> of 911 developers found that timely feedback, review size, and understanding the motivation for a change were the top three challenges in code review. Those challenges should sound familiar. The research had already identified many of the practices that improve review quality, yet only 26% of developers said they always wrote a detailed description of the code being reviewed. &#8220;Bikeshedding&#8221;&#8212;disputing minor issues while more serious ones went unexamined&#8212;remained one of the most common review failures. We didn&#8217;t need new guidance. We needed to consistently apply what we already knew.</span></p><p><span>AI didn&#8217;t create this situation. It inherited it, and then amplified it. Larger PRs, higher review volume, less context per change, these aren&#8217;t new symptoms. They&#8217;re old ones, scaled up.</span></p><p><span>Before asking AI to fix your code review process, ask whether your team has built the habits that make code review effective in the first place. Small, well-explained changes. Protected reviewer time. Automated routine checks so humans can focus on judgment.</span></p><p><span>AI can absolutely improve code review. But it can&#8217;t compensate for a review culture that was already struggling. It doesn&#8217;t eliminate bad review habits. It amplifies them.</span></p><h3><span>AI can help, if we use it wisely</span></h3><p><span>Once the fundamentals are in place, AI has a real role to play in code review. That&#8217;s exactly what we found in our </span><a href="https://www.microsoft.com/en-us/research/publication/ai-where-it-matters-where-why-and-how-developers-want-ai-support-in-daily-work/"><span>AI Where It Matters</span></a><span> research. Developers don&#8217;t want code review to disappear. They want AI to remove the parts of review that don&#8217;t require human judgment so reviewers can spend more time on the parts that do.</span></p><p><span>What they want AI to do: catch security and compliance issues, flag high-risk changes, generate test scaffolding, surface the impact of a change across the codebase, and handle the high-volume routine so human attention can go where it matters. As one developer put it: </span><em><span>&#8220;Should be able to detect high risk changes and derisk them.&#8221;</span></em></p><p><span>What they explicitly don&#8217;t want: AI that auto-merges, auto-commits, or takes final accountability. </span><em><span>&#8220;I don&#8217;t want AI to just act as a red-light / green-light. It should raise issues&#8230; and still require human review.&#8221;</span></em><span> Developers aren&#8217;t asking for a replacement, they&#8217;re asking for a better collaborator.</span></p><p><span>Interestingly, one of the most sophisticated production deployments I&#8217;ve seen tries to walk that line.</span></p><p><a href="https://arxiv.org/abs/2605.30208"><span>Meta&#8217;s RADAR</span></a><span> (Risk Aware Diff Auto Review) system automates review for a carefully selected subset of low-to-medium risk changes while routing higher-risk diffs to human reviewers. It combines static analysis, machine learning, LLM-based review, and deterministic validation before anything lands.</span></p><p><span>The results are striking: more than 535,000 diffs reviewed, over 331,000 landed, a revert rate roughly one-third that of non-RADAR diffs, a production incident rate one-fiftieth as high, and a 3.3x faster median time to close (roughly a 70% reduction).</span></p><p><span>RADAR isn&#8217;t simply &#8220;AI reviewing code.&#8221; It&#8217;s a carefully engineered system built around the principle that scarce human attention should be reserved for changes where human judgment and accountability matter most.</span></p><p><span>Just as importantly, the RADAR team also acknowledges a trade-off. Automated review can dramatically improve efficiency, but as automation expands, the knowledge transfer provided by human review could suffer. They identify this as something engineering organizations should actively monitor.</span></p><p><span>That&#8217;s the distinction I think many organizations miss. AI shouldn&#8217;t eliminate human review. It should make human review more valuable. AI-enabled review should have discipline around it: clear eligibility criteria, thoughtful risk stratification, and a deliberate decision about which changes deserve human attention, and why.</span></p><p><span>If you&#8217;re evaluating an AI review system, don&#8217;t start by asking, </span><em><span>&#8220;Does it work?&#8221;</span></em><span> Start by asking, </span><em><span>&#8220;How does it maximize the time and value of human judgment?&#8221;</span></em></p><h3><span>Don&#8217;t lose what review was actually doing</span></h3><p><span>Here&#8217;s the part that gets left out of the AI review conversation: code review was never just about finding defects.</span></p><p><a href="https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/ICSE202013-codereview.pdf?msockid=15af30cf0f0662a5037e27800ec7634a"><span>A landmark Microsoft study</span></a><span> found that while most developers identified defect detection as a primary motivation for code review, defect-related comments made up only 14% of actual review comments. In practice, code review serves many other purposes. More than half of developers said they use reviews to explore alternative solutions, while many also pointed to knowledge transfer and gaining awareness of what their teammates are building.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5YOe!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5d57450-a491-4dcf-a1c1-3ad706041cea_3604x2884.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5YOe!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5d57450-a491-4dcf-a1c1-3ad706041cea_3604x2884.png 424w, https://substackcdn.com/image/fetch/$s_!5YOe!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5d57450-a491-4dcf-a1c1-3ad706041cea_3604x2884.png 848w, https://substackcdn.com/image/fetch/$s_!5YOe!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5d57450-a491-4dcf-a1c1-3ad706041cea_3604x2884.png 1272w, https://substackcdn.com/image/fetch/$s_!5YOe!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5d57450-a491-4dcf-a1c1-3ad706041cea_3604x2884.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5YOe!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5d57450-a491-4dcf-a1c1-3ad706041cea_3604x2884.png" width="1456" height="1165" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a5d57450-a491-4dcf-a1c1-3ad706041cea_3604x2884.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1165,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:181227,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.getdx.com/i/204342182?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5d57450-a491-4dcf-a1c1-3ad706041cea_3604x2884.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!5YOe!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5d57450-a491-4dcf-a1c1-3ad706041cea_3604x2884.png 424w, https://substackcdn.com/image/fetch/$s_!5YOe!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5d57450-a491-4dcf-a1c1-3ad706041cea_3604x2884.png 848w, https://substackcdn.com/image/fetch/$s_!5YOe!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5d57450-a491-4dcf-a1c1-3ad706041cea_3604x2884.png 1272w, https://substackcdn.com/image/fetch/$s_!5YOe!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5d57450-a491-4dcf-a1c1-3ad706041cea_3604x2884.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">&#8220;Modern Code Review&#8221; (Bacchelli &amp; Bird, ICSE 2013)</figcaption></figure></div><p><strong><span>The visible output of code review is better code. The invisible output is a better engineering organization.</span></strong></p><p><span>Code review is how teams build shared understanding of a system. It&#8217;s how junior developers learn from experienced ones. It&#8217;s how architectural intent gets surfaced, questioned, and refined. It&#8217;s how one engineer&#8217;s mental model gradually becomes the team&#8217;s mental model. Organizations don&#8217;t become resilient because one person understands a subsystem. They become resilient because many people do.</span></p><p><span>This is why the stakes around AI review are so high. Automate the review of a diff, and you may have successfully reviewed that diff. But you haven&#8217;t transferred any knowledge. You haven&#8217;t built shared ownership. You haven&#8217;t given a newer engineer a window into how a more experienced teammate reasons about trade-offs. You haven&#8217;t surfaced the design rationale that someone will need six months from now when they&#8217;re trying to respond to customer feedback.</span></p><p><span>Developers in our </span><em><span>AI Where It Matters</span></em><span> research understood this instinctively. One participant wrote, </span><em><span>&#8220;I can&#8217;t fully delegate the final code review to AI&#8212;my approval puts my name on it.&#8221;</span></em><span> Another warned, </span><em><span>&#8220;Intellectual offloading can result in errors that eventually no one understands.&#8221;</span></em><span> That&#8217;s the slow-moving risk. The gradual erosion of a team&#8217;s ability to reason about its own software.</span></p><p><span>Margaret-Anne Storey&#8217;s </span><a href="https://queue.acm.org/detail.cfm?id=3807966"><span>recent work</span></a><span> gives this phenomenon a name. As AI accelerates software development, teams don&#8217;t just accumulate technical debt. They accumulate cognitive and intent debt&#8212;a growing gap between what the system does and what the organization collectively understands about why it does it. Those debts don&#8217;t appear on a dashboard. They surface months later, during an outage, a handoff, or a redesign, when nobody remembers the reasoning that once lived inside a code review conversation. By then, recovering that understanding is far more expensive than preserving it would have been.</span></p><p><span>This future isn&#8217;t inevitable. But it also won&#8217;t arrive all at once. It will emerge through a series of individually reasonable decisions: this change is low risk, this review can be automated, this approval can be skipped. Each decision saves a little time. Taken together, they may slowly eliminate one of the primary ways engineering teams build shared understanding.</span></p><p><span>The challenge isn&#8217;t choosing between AI and human review. It&#8217;s deciding which parts of code review are too valuable to automate away.</span></p><h2><span>What to actually do</span></h2><p><span>Three things, in order.</span></p><p><strong><span>Fix the basics first.</span></strong><span> Audit your current review process. Are pull requests small enough to review meaningfully? Do change descriptions explain </span><em><span>why</span></em><span>, not just </span><em><span>what</span></em><span>? Are reviewers protected from overload? Are automated tools already handling the routine work they should (e.g. formatting, linting, and obvious style issues)?</span></p><p><strong><span>Design AI around human judgment.</span></strong><span> Developers consistently describe code review as high-value, high-accountability work. They don&#8217;t want AI making the decision; they want AI helping them make better ones. That means risk stratification instead of blanket automation. It means AI that surfaces issues, not AI that silently resolves them. It means conservative eligibility thresholds, auditability, and clear human accountability.</span></p><p><strong><span>Protect what review is actually building.</span></strong><span> The easiest thing to measure about code review is defects. The most valuable thing it produces is shared understanding. Measure review health beyond throughput. Are junior developers learning? Is architectural knowledge spreading across the team? Are reviewers engaging with substance or simply rubber-stamping? Design your AI review strategy so automation absorbs the routine while humans spend more time on the conversations that create understanding, ownership, and better engineering judgment.</span></p><p><span>Code review is one of the highest-leverage practices in software engineering, and right now it&#8217;s under pressure from every direction. The answer isn&#8217;t to make it faster by making it shallower. It&#8217;s to get serious about doing it well&#8212;with or without AI&#8212;and then use AI deliberately, in the places where it earns trust and preserves what the practice was accomplishing all along.</span></p><p><strong><span>AI should absolutely reduce the time we spend reviewing code. It just shouldn&#8217;t reduce the amount we learn from it.</span></strong></p><div><hr></div><p>That&#8217;s it for this week. Thanks for reading.</p><p>-Brian</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.getdx.com/p/what-are-code-reviews-even-for?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.getdx.com/p/what-are-code-reviews-even-for?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p>]]></content:encoded></item><item><title><![CDATA[The unexpected things developer metrics can measure]]></title><description><![CDATA[We built metrics to evaluate developer tools. They turned out to explain everything from office design to Daylight Saving Time.]]></description><link>https://newsletter.getdx.com/p/the-unexpected-things-developer-metrics</link><guid isPermaLink="false">https://newsletter.getdx.com/p/the-unexpected-things-developer-metrics</guid><dc:creator><![CDATA[Brian Houck]]></dc:creator><pubDate>Fri, 31 Jul 2026 10:45:41 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/61bbaac1-db67-4e66-a5fa-19896e84c651_2400x1260.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em><strong>Welcome to the latest issue of Engineering Enablement</strong><span>, a weekly newsletter sharing research and perspectives on developer productivity.</span></em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.getdx.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.getdx.com/subscribe?"><span>Subscribe now</span></a></p><p><span>DX&#8217;s Q2 AI Impact Report is now available with the latest research on AI&#8217;s impact across engineering organizations. </span><a href="https://getdx.com/resources/?utm_source=newsletter">Read the full report.</a></p><div><hr></div><p><em><span>A quick heads-up: this issue is a little different from our usual format. Instead of sharing a finding from our research and conversations, this issue is more of a reframe of how to think about what developer metrics are actually for.</span></em></p><p><span>Here is a sentence I never expected to write: Developer-experience metrics can measure the impact of Daylight Saving Time.</span></p><p><span>Not a survey asking developers how they feel about the time change. An actual, measurable shift in engineering behavior, detected using the same kinds of outcome metrics we use to evaluate AI coding assistants, build systems, and code review workflows.</span></p><p><span>That sounds like an absurd thing to measure. And yet, </span><a href="https://www.linkedin.com/posts/brianhouck_developerexperience-daylightsavingtime-productivity-activity-7305671070206869504-VNgj?utm_source=share&amp;utm_medium=member_desktop&amp;rcm=ACoAAAbL-CEB_GIa71OGfWkh-nZOd86fsAhsAZc"><span>that&#8217;s exactly what we found</span></a><span>.</span></p><p><span>When Daylight Saving Time begins, developers in Seattle (one of the highest-latitude major cities in the continental United States) show a measurable shift relative to developers closer to the equator, where seasonal daylight changes are much smaller. Across three years of data, active coding time increased by roughly 6%, while pull requests completed about 8% faster, driven largely by quicker code reviews.</span></p><p><span>The most plausible explanation is also the simplest: an extra hour of evening daylight appears to keep people engaged with their work a little longer.</span></p><p><span>I&#8217;d resist reading that as purely good news. More engagement isn&#8217;t automatically healthier. If some of that extra time is coming at the expense of sleep or recovery, that&#8217;s a trade-off worth measuring, not celebrating.</span></p><p><span>At first glance, this has nothing to do with software engineering. After all, developer metrics are supposed to measure developer things: faster builds, better IDEs, improved code reviews, AI-assisted coding.</span></p><p><span>Or so I thought.</span></p><p><span>The more I&#8217;ve worked with developer-experience metrics, the more I&#8217;ve come to believe we&#8217;ve been thinking about them too narrowly. We often describe them as a way to evaluate developer tools and engineering workflows. But that&#8217;s not really what they&#8217;re measuring. They&#8217;re measuring the experience of doing software engineering. And that experience is shaped by far more than software.</span></p><h3><span>Measuring outcomes, not interventions</span></h3><p><span>When people think about developer-experience metrics, they naturally think about the interventions we introduce: a new AI coding assistant, a faster build system, a different code review process, a new deployment pipeline.</span></p><p><span>But those aren&#8217;t actually what the metrics care about.</span></p><p><span>Good developer-experience metrics measure outcomes. They tell us whether developers are able to do focused, meaningful, high-quality work. Once you measure outcomes instead of interventions, something interesting happens.</span></p><p><span>The intervention no longer has to be software. It can be an office redesign. A meeting policy. The weather. Even Daylight Saving Time.</span></p><p><span>That doesn&#8217;t mean engineering leaders suddenly own the weather, facilities, or company policy. But measurement doesn&#8217;t have to imply ownership. Sometimes it helps explain why an outcome changed. Other times it gives you evidence to influence the people who </span><em><span>do</span></em><span> own the lever.</span></p><p><span>That realization changed how I think about developer metrics. They&#8217;re still excellent for evaluating developer tools. They just turn out to be useful for much more.</span></p><p><span>Once I started looking through this lens, examples kept appearing, not just in my own research, but across entirely different disciplines. And they weren&#8217;t limited to software engineering. Researchers have found that </span><a href="https://docs.iza.org/dp12632.pdf"><span>higher indoor air pollution</span></a><span> correlates to more errors by chess players, </span><a href="https://www.aeaweb.org/articles?id=10.1257/pol.20180612"><span>warmer classrooms</span></a><span> are associated with lower student performance, and </span><a href="https://www.sciencedirect.com/science/article/abs/pii/S0272494411000429"><span>high-noise environments</span></a><span> measurably degrade memory and motivation. Different domains, different outcomes, but the same underlying lesson: our environment shapes performance.</span></p><p><span>The difference is that developer-experience metrics give us a language for asking the same kinds of questions about software engineering.</span></p><p><span>Some of those influences are things organizations can change. Others aren&#8217;t. Both leave measurable fingerprints on the developer experience.</span></p><h3><span>The spaces your developers work in</span></h3><p><span>Daylight Saving Time is an unusual example because there isn&#8217;t much an engineering leader can do about it. Office space is different. Organizations make decisions about where and how developers work all the time, yet those decisions are often driven by cost, convenience, or intuition rather than evidence.</span></p><p><a href="https://www.microsoft.com/en-us/research/publication/the-best-of-both-worlds-unlocking-the-potential-of-hybrid-work-for-software-engineers/"><span>When we asked developers</span></a><span> what they actually value about coming into the office, the answers were overwhelmingly human. The top response, by a wide margin, was other people: seeing colleagues face to face, the conversations that get sparked, the camaraderie. Office design came next (whiteboards, rooms to hash out problems), followed by food and coffee. This isn&#8217;t just a list of office perks. It&#8217;s a window into the parts of the developer experience that still depend on the physical world.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HIkS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F957195e1-477c-46bd-bb32-9ddfbd19e1ef_2788x2000.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HIkS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F957195e1-477c-46bd-bb32-9ddfbd19e1ef_2788x2000.png 424w, https://substackcdn.com/image/fetch/$s_!HIkS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F957195e1-477c-46bd-bb32-9ddfbd19e1ef_2788x2000.png 848w, https://substackcdn.com/image/fetch/$s_!HIkS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F957195e1-477c-46bd-bb32-9ddfbd19e1ef_2788x2000.png 1272w, https://substackcdn.com/image/fetch/$s_!HIkS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F957195e1-477c-46bd-bb32-9ddfbd19e1ef_2788x2000.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HIkS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F957195e1-477c-46bd-bb32-9ddfbd19e1ef_2788x2000.png" width="1456" height="1044" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/957195e1-477c-46bd-bb32-9ddfbd19e1ef_2788x2000.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1044,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:219849,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.getdx.com/i/208387583?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F957195e1-477c-46bd-bb32-9ddfbd19e1ef_2788x2000.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!HIkS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F957195e1-477c-46bd-bb32-9ddfbd19e1ef_2788x2000.png 424w, https://substackcdn.com/image/fetch/$s_!HIkS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F957195e1-477c-46bd-bb32-9ddfbd19e1ef_2788x2000.png 848w, https://substackcdn.com/image/fetch/$s_!HIkS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F957195e1-477c-46bd-bb32-9ddfbd19e1ef_2788x2000.png 1272w, https://substackcdn.com/image/fetch/$s_!HIkS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F957195e1-477c-46bd-bb32-9ddfbd19e1ef_2788x2000.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>The specifics ranged from the practical to the personal. One developer&#8217;s entire reason for coming in: &#8220;I don&#8217;t have to fight with my cats.&#8221; However they phrased it, the pattern was the same. What draws developers to the office is overwhelmingly human, and largely physical, the parts of the experience that software still hasn&#8217;t managed to replace.</span></p><p><span>Those findings explain </span><em><span>why</span></em><span> the office matters. Developer metrics can help answer the next question: </span><strong><span>how much does it matter?</span></strong></p><p><span>In a previously unpublished analysis, I looked at teams relocating between office spaces. The building itself moved the numbers.</span></p><p><span>One team relocated into a well-designed, high-performance office and saw developer engagement, measured through active coding time, increase by roughly 13% relative to comparable peer teams. The really interesting part came later. When that same team eventually moved back into their original office, their active coding time returned almost exactly to its previous level.</span></p><p><span>I also saw the opposite pattern. Teams moving into poorly designed office spaces experienced productivity declines on the order of 10%.</span></p><p><span>This is observational, not a controlled experiment, so we should be cautious about claiming causality. Office moves often coincide with organizational changes, new teammates, and countless other confounding factors. But the pattern is still striking. The same outcome metrics we might use to evaluate a new AI coding assistant also measured the impact of an office redesign. The intervention changed. The measurement didn&#8217;t.</span></p><p><span>That changes the conversation. Office space is usually treated as a facilities expense to be minimized. But if a better workspace meaningfully improves the developer experience, it becomes a productivity investment instead. A useful rule of thumb is that facilities costs are roughly 10% of payroll. That means an improvement in developer effectiveness on the order of 10% has the potential to offset the entire cost of the workspace, before considering any additional benefits such as hiring, retention, or collaboration.</span></p><p><span>Interestingly, our earlier hybrid-work research found another version of the same idea. Developers who chose whether to work from home or the office based on the type of work they planned to do reported higher productivity than those choosing primarily for personal convenience. Different environments appear to support different kinds of work, and developer metrics give us a way to test those assumptions rather than simply debate them.</span></p><h3><span>When the weather is a variable</span></h3><p><span>Daylight Saving Time at least has policy debates attached to it. Weather is even simpler. No engineering leader can change it.</span></p><p><span>And yet </span><a href="https://queue.acm.org/detail.cfm?id=3819080"><span>it still shows up in the data</span></a><span>.</span></p><p><span>In Seattle, you can often identify winter snow days simply by looking at engineering activity. Active coding time drops by roughly 18%. The reasons are easy to imagine: disrupted commutes, childcare, school closures, or simply the irresistible pull of a rare Pacific Northwest snow day. Untangling those mechanisms is difficult, and I wouldn&#8217;t claim a clean causal story.</span></p><p><span>But the mechanism isn&#8217;t really the point.</span></p><p><span>The point is that the measurement detected the change.</span></p><p><span>At first glance, measuring something you can&#8217;t control might seem pointless. I think the opposite is true. If a team&#8217;s delivery slows during a snowstorm, that isn&#8217;t necessarily a performance problem to solve. It&#8217;s the context that helps explain what happened.</span></p><p><span>That&#8217;s one of the underappreciated benefits of developer-experience metrics. Sometimes their greatest value isn&#8217;t telling you what to change. It&#8217;s telling you what not to blame.</span></p><p><span>Knowing that a metric moved because of external circumstances prevents organizations from chasing the wrong explanations, setting unrealistic expectations, or concluding that a team suddenly became less effective when nothing about the team actually changed.</span></p><h3><span>The durable part</span></h3><p><span>There is a reason I keep coming back to these examples, and it is not just that they are fun to share at a dinner party.</span></p><p><span>It is tempting to think of developer-experience metrics as a way to evaluate developer tools. But that is too narrow. Good developer metrics measure outcomes, not interventions. They tell us whether developers can do focused, meaningful, high-quality work. The intervention itself might be a faster build, a better office, a meeting policy, or even something as unexpected as Daylight Saving Time.</span></p><p><span>That distinction is what makes these measurement systems durable.</span></p><p><span>Five years ago, organizations were asking different questions than they are today. Five years from now, they&#8217;ll ask different questions again. AI is the dominant topic today, just as cloud development environments, CI/CD, or code review tooling were at other points in time. The interventions evolve. The outcomes we care about&#8212;Speed, Ease, Quality, and Thriving&#8212;do not.</span></p><p><span>That&#8217;s why I don&#8217;t think AI is a special case. It is simply the latest intervention whose impact we want to understand. The questions remain the same: Does it help developers do better work? Does it reduce friction? Does it improve quality? Does it help people thrive?</span></p><p><span>The tools will continue to change. The interventions will continue to change. The measurement doesn&#8217;t have to.</span></p><p><span>It also leaves an interesting question for another day: if these ideas apply so well to software engineering, how much further do they extend?</span></p><div><hr></div><p>That&#8217;s it for this week. Thanks for reading.</p><p>-Brian</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.getdx.com/p/the-unexpected-things-developer-metrics?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.getdx.com/p/the-unexpected-things-developer-metrics?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p>]]></content:encoded></item><item><title><![CDATA[Measuring the impact of AI coding tools: Capacity, not horsepower]]></title><description><![CDATA[Separate what to measure from how, then measure across dimensions.]]></description><link>https://newsletter.getdx.com/p/measuring-the-impact-of-ai-coding</link><guid isPermaLink="false">https://newsletter.getdx.com/p/measuring-the-impact-of-ai-coding</guid><dc:creator><![CDATA[Brian Houck]]></dc:creator><pubDate>Wed, 29 Jul 2026 10:42:16 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e0058d22-6461-4cf5-995b-bcecb2f63961_2400x1260.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em><strong>Welcome to the latest issue of Engineering Enablement</strong><span>, a weekly newsletter sharing research and perspectives on developer productivity.</span></em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.getdx.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.getdx.com/subscribe?"><span>Subscribe now</span></a></p><p><span>DX&#8217;s Q2 AI Impact Report is now available with the latest research on AI&#8217;s impact across engineering organizations. </span><a href="https://getdx.com/resources/?utm_source=newsletter">Read the full report.</a></p><div><hr></div><p><span>When an organization invests in an AI coding tool, the next question is almost always some version of &#8220;How do we prove it&#8217;s working?&#8221; It&#8217;s a fair question, and one many organizations are wrestling with today.</span></p><p><span>I was recently asked to weigh in on one such proposal: a metric called &#8220;Developer Horsepower,&#8221; defined as useful work per day and calculated by multiplying AI-assisted pull requests by an estimate of the human effort each would have required. It&#8217;s an intuitive idea, and a genuinely thoughtful attempt at a hard problem. But I think it starts one step too early.</span></p><p><span>Rather than asking how much human work the AI replaced, I&#8217;d ask whether AI has increased the organization&#8217;s capacity to deliver innovation (and whether that additional capacity is sustainable). That framing leads to a very different measurement strategy, and one that I believe is both easier to defend and more actionable.</span></p><h3><span>First, separate &#8220;what&#8221; from &#8220;how&#8221;</span></h3><p><span>Two different questions get tangled together in most measurement conversations. The first is what to measure. The second is how: not just how you collect the data (surveys versus telemetry) but how strong your evidence needs to be. Is a correlation enough, or do you need a full causal study?</span></p><p><span>Proving causation is genuinely valuable, and I&#8217;d never talk someone out of it. If you&#8217;re doing something novel, or you want to publish in a peer-reviewed journal, there are good approaches available, from difference-in-differences designs to dose-response studies. But for most organizations, a full causal study is overkill. There is now </span><a href="https://newsletter.getdx.com/p/five-studies-that-are-changing-how"><span>substantial causal evidence</span></a><span> that AI coding tools can improve coding throughput under many conditions. You don&#8217;t need to re-prove that AI increases coding throughput; the field has done that work. In most cases you can measure the correlations in your own environment and lean on the existing causal literature to interpret them. That&#8217;s a far lighter lift, and it&#8217;s honest (as long as you&#8217;re not claiming something the literature doesn&#8217;t support).</span></p><p><span>With that settled, the interesting question is what to measure.</span></p><h3><span>What to measure: innovation capacity, across dimensions</span></h3><p><span>The goal isn&#8217;t to isolate the tool. It&#8217;s to answer whether your engineering system now has more capacity to deliver innovation, and whether that capacity is sustainable. I&#8217;d build that up in layers.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!luT0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F085238b0-c071-4870-b15b-8de6ffb5c961_2400x1586.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!luT0!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F085238b0-c071-4870-b15b-8de6ffb5c961_2400x1586.png 424w, https://substackcdn.com/image/fetch/$s_!luT0!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F085238b0-c071-4870-b15b-8de6ffb5c961_2400x1586.png 848w, https://substackcdn.com/image/fetch/$s_!luT0!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F085238b0-c071-4870-b15b-8de6ffb5c961_2400x1586.png 1272w, https://substackcdn.com/image/fetch/$s_!luT0!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F085238b0-c071-4870-b15b-8de6ffb5c961_2400x1586.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!luT0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F085238b0-c071-4870-b15b-8de6ffb5c961_2400x1586.png" width="1456" height="962" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/085238b0-c071-4870-b15b-8de6ffb5c961_2400x1586.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:962,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:416952,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.getdx.com/i/208356729?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F085238b0-c071-4870-b15b-8de6ffb5c961_2400x1586.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!luT0!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F085238b0-c071-4870-b15b-8de6ffb5c961_2400x1586.png 424w, https://substackcdn.com/image/fetch/$s_!luT0!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F085238b0-c071-4870-b15b-8de6ffb5c961_2400x1586.png 848w, https://substackcdn.com/image/fetch/$s_!luT0!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F085238b0-c071-4870-b15b-8de6ffb5c961_2400x1586.png 1272w, https://substackcdn.com/image/fetch/$s_!luT0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F085238b0-c071-4870-b15b-8de6ffb5c961_2400x1586.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><span>Throughput: Is more total work happening?</span></strong><span> PR throughput is a perfectly valid system-level signal for whether more work, in aggregate, is moving through your engineering system. It gets criticized, and rightly so, as a measure of any individual developer, because not all PRs are equal in size or complexity. But at the system level it holds up well. If you can normalize for PR complexity, that&#8217;s better still (</span><a href="https://getdx.com/truethroughput/"><span>TrueThroughput</span></a><span> is one approach to doing that). But even without normalization, aggregate throughput remains a useful directional indicator. I explore this in more depth in the recent article: </span><a href="https://newsletter.getdx.com/p/revisiting-the-dx-core-4"><span>Revisiting the DX Core 4 in the age of AI</span></a></p><p><strong><span>Deployments: Does the work survive to shipped software?</span></strong><span> This is the step most teams skip, and it&#8217;s the most important one. Recent research shows that AI increases coding activity far more than it increases shipped software. In </span><a href="https://www.nber.org/papers/w35275"><span>one large study of over 100,000 developers</span></a><span>, coding agents raised commit volume by up to 180%, yet the effect on actual releases attenuated to roughly 20-30%. The authors attribute the gap to a &#8220;weak link&#8221;: the upstream speedup runs into human bottlenecks downstream, in review, integration, and release.</span></p><p><span>Counting merged PRs alone may overstate impact, because a large share of that activity attenuates before it reaches production. Tracing your throughput increase all the way through to deployments tells you how much of the upstream gain actually survives.</span></p><p><span>I&#8217;d frame the result as innovation capacity: the organization&#8217;s realized capacity to deliver innovation. Note the deliberate boundary. Whether that additional capacity translates into business value is a separate question, and often not an engineering one. Deployment frequency is the practical proxy for delivered work; business impact is the true north if you can reach it. Conflating the two is how measurement programs lose credibility.</span></p><p><strong><span>Innovation Time Ratio: Are the savings being reinvested?</span></strong><span> AI only creates organizational value if the time it saves is reinvested in higher-value work. This is exactly what a metric like </span><a href="https://queue.acm.org/detail.cfm?id=3819080"><span>Innovation Time Ratio</span></a><span> captures: the share of developer time spent creating new value versus running the business or carrying administrative load. For most organizations, survey data is the logical place to start, unless you already have robust calendar and development telemetry. And self-report is less of a concern here than it first appears, because you&#8217;re looking at change over time. Whatever reporting bias exists tends to stay consistent across measurements, so it largely cancels out in the delta.</span></p><p><strong><span>Quality: A guardrail, not an afterthought.</span></strong><span> More work moving faster is only progress if it isn&#8217;t just churn. A quality signal such as change failure rate or incident mitigation time keeps the speed gains honest. If throughput climbs while failure rates climb with it, you haven&#8217;t gained capacity. You&#8217;ve moved the cost somewhere less visible.</span></p><p><strong><span>Satisfaction: The sustainability guardrail.</span></strong><span> Finally, watch developer satisfaction. Throughput and experience can decouple, and gains bought by burning out your engineers aren&#8217;t gains you get to keep. Sustainable capacity requires healthy developers. If higher throughput comes at the expense of satisfaction, you&#8217;ve likely borrowed from future capacity rather than increased it.</span></p><p><span>Taken together, those layers answer the question leadership is actually asking: not &#8220;how much code did the tool produce?&#8221; but &#8220;does our system have more capacity to deliver innovation, and can we sustain it?&#8221;</span></p><h3><span>&#8220;Developer Horsepower&#8221;</span></h3><p><span>Which brings me back to the proposal I was asked about. I understand the appeal of a single &#8220;horsepower&#8221; number, but ultimately I&#8217;d recommend a different approach.</span></p><p><span>To be clear, my aim isn&#8217;t to single out this particular metric. It&#8217;s a reasonable attempt at a real problem, and the team that proposed it is asking exactly the right question. I want to use it as a worked example of the kind of probing I do whenever someone hands me a composite score: what is the unit actually measuring, and how much variation is it hiding? Those questions apply to any single-figure productivity metric, not just this one.</span></p><p><span>My concern isn&#8217;t that &#8220;Developer Horsepower&#8221; is impossible to calculate. It&#8217;s that I&#8217;m not sure it is answering the question that really needs answering. Leadership doesn&#8217;t ultimately care how many &#8220;human-equivalent hours&#8221; an AI system replaced. They care whether their engineering organization can deliver more innovation, more reliably, and more sustainably than before.</span></p><p><span>The core problem is that it tries to collapse a multidimensional question into one figure. That runs against a principle the field has largely converged on: </span><a href="https://www.microsoft.com/en-us/research/publication/the-space-of-developer-productivity-theres-more-to-it-than-you-think/"><span>engineering productivity can&#8217;t be captured by a single metric</span></a><span>. A single number simplifies reporting, but it obscures the tradeoffs that matter and makes it difficult to understand what&#8217;s actually driving change.</span></p><p><span>The specific construction compounds this in several ways.</span></p><p><span>First, the normalization itself is non-standard and difficult to interpret. &#8220;Human-equivalent hours&#8221; isn&#8217;t an industry term, and a figure built on a bespoke conversion is difficult for anyone outside the team to trust or reason about.</span></p><p><span>More fundamentally, it&#8217;s not obvious what those hours are supposed to represent.</span></p><p><span>One of the more surprising findings from our </span><em><a href="https://queue.acm.org/detail.cfm?id=3807961"><span>AI Native Developer</span></a></em><span> research was that developers spend only about 14% of their work week actively writing code. The rest is spread across activities like design, code review, debugging, learning, meetings, planning, documentation, and coordinating with teammates. AI doesn&#8217;t simply replace coding time; it changes how engineers spend time across many of those activities. Some work disappears, some shifts, and entirely new work (like reviewing AI-generated code or managing context) emerges.</span></p><p><span>Trying to convert an AI-assisted pull request into a fixed number of &#8220;human hours saved&#8221; assumes those relationships are stable and additive. In practice, they&#8217;re neither.</span></p><p><span>Even if &#8220;human-equivalent hours&#8221; were the right unit, a flat estimate is almost certainly inaccurate. Eight hours per PR treats every pull request as identical, but PRs vary enormously in size, complexity, and review effort.</span></p><p><span>The research also increasingly suggests that AI and human effort remain largely </span><a href="https://www.nber.org/papers/w35275"><span>complements rather than substitutes</span></a><span>. Engineers still spend substantial time reviewing, integrating, validating, testing, and coordinating around AI-generated code. A clean &#8220;hours replaced&#8221; conversion assumes substitution where the evidence points toward augmentation.</span></p><h3><span>Where to start</span></h3><p><span>If this sounds like more than you can measure today, start anyway. You do not need to be able to measure all of this with a mature telemetry stack in order to begin.</span></p><p><span>Don&#8217;t have the telemetry? Use surveys. If you have precise, telemetry-driven data for every commit, review, build, and deployment, you&#8217;re in an uncommonly good position. Most organizations aren&#8217;t, and that&#8217;s fine. Targeted surveys will get you moving quickly, and they&#8217;re better suited to some of these dimensions than telemetry is anyway (reinvested time and satisfaction, for instance). The usual objection is self-report bias, but it matters less than people expect here, because you&#8217;re watching change over time. Whatever bias exists tends to stay consistent across measurements, so it largely cancels out in the delta. The measure only has to be directionally correct to be useful.</span></p><p><span>Don&#8217;t feel lost if you can only start with a couple of metrics. You don&#8217;t need all five dimensions on day one. Pick one or two metrics per dimension, balancing an objective signal with a subjective one, and add more as your instrumentation matures. A speed metric from telemetry paired with a satisfaction question from a survey tells you more than five telemetry metrics alone. Too many metrics dilute focus and make it harder to see what&#8217;s actually changing. Start small, learn what moves, and expand deliberately.</span></p><p><span>As your instrumentation matures, work toward the full picture. The richer your measurement, the more confidently you can answer whether your capacity to deliver innovation is real and sustainable.</span></p><h2><span>The bottom line</span></h2><p><span>The pressure to produce one clean number that proves an AI tool&#8217;s horsepower is understandable, but it answers the wrong question. The better question is whether your engineering system has more capacity to deliver innovation, and whether that capacity is sustainable. Measure that across a few well-chosen dimensions, lean on the causal work the field has already done, and you&#8217;ll have something far more defensible than any single horsepower figure, and far more useful for deciding what to do next. AI shouldn&#8217;t be judged by how much human work it appears to replace. It should be judged by whether it gives your engineering organization a greater and more sustainable capacity to deliver innovation.</span></p><div><hr></div><p><span>This week&#8217;s featured DevProd job openings. See more </span><a href="https://getdx.com/resources/devex-jobs/">open roles here</a><span>.</span></p><ul><li><p><strong>Ashby</strong><span> is hiring an </span><a href="https://jobs.ashbyhq.com/Ashby/0f5dbf59-687b-4d88-88a7-73ee0a66b48d?utm_source=PRgMeEgv1Z">Staff Platform Engineer</a><span> | Remote</span></p></li><li><p><strong>Cart</strong><span> is hiring a </span><a href="https://www.linkedin.com/jobs/view/4404135082">Sr. Software Engineer II, Developer Experience</a><span> | </span>Santa Clara, CA; San Francisco, CA; New York, NY</p></li><li><p><strong>Cashea</strong><span> is hiring an </span><a href="https://cashea.na.teamtailor.com/jobs/579773-infrastructure-developer-productivity-platform-engineering-manager">Infrastructure &amp; Developer Productivity Platform Engineering Manager</a><span> | Remote</span></p></li><li><p><strong>Figma</strong><span> is hiring a </span><a href="https://job-boards.greenhouse.io/figma/jobs/5790627004?gh_jid=5790627004&amp;gh_src=db0ijm3x4us">Staff Software Engineer, Developer Experience</a><span> | Remote; US</span></p></li><li><p><strong>Morgan Stanely </strong><span>is hiring an </span><a href="https://www.linkedin.com/jobs/view/4393043964/">AI Platform Engineer - Vice President</a><span> | New York</span></p></li><li><p><strong>Notion</strong> is hiring a <a href="https://jobs.ashbyhq.com/notion/49bdf081-6e20-4323-8c73-6d6b19544ff5">Software Engineer, Developer Experience</a> | Hybrid; Hyderabad, India</p></li></ul><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.getdx.com/p/measuring-the-impact-of-ai-coding?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.getdx.com/p/measuring-the-impact-of-ai-coding?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p>]]></content:encoded></item><item><title><![CDATA[Research briefing with Brian Houck: Measuring AI agents and revisiting the Core 4]]></title><description><![CDATA[Justin Reock and Brian Houck explore how AI coding agents are reshaping engineering metrics and what leaders need to measure in the age of AI.]]></description><link>https://newsletter.getdx.com/p/research-briefing-with-brian-houck</link><guid isPermaLink="false">https://newsletter.getdx.com/p/research-briefing-with-brian-houck</guid><dc:creator><![CDATA[Justin Reock]]></dc:creator><pubDate>Fri, 24 Jul 2026 13:50:33 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/207581348/14162e401718c8fcc01f420bcb2cbc98.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>Listen and watch now on <strong><a href="https://youtu.be/U3p-rqOrAts">YouTube</a>, <a href="https://podcasts.apple.com/us/podcast/engineering-enablement-by-abi-noda/id1619140476">Apple</a>, and <a href="https://open.spotify.com/show/3NxjyIsuxeDMQtisDqBy7D">Spotify</a></strong>.</p><p><span>AI coding agents are changing how software gets built, but they&#8217;re also forcing us to rethink how we measure engineering effectiveness. Traditional developer experience metrics were designed for humans, not AI agents, so how should engineering leaders adapt?</span></p><p><span>In this webinar, I&#8217;m joined by Brian Houck, Distinguished Scientist at DX and co-author of the SPACE framework, to explore the emerging field of agent experience and how it builds on developer experience rather than replacing it. We discuss how organizations can prepare for AI-assisted software development, how the DX Core 4 applies in the age of AI, why metrics like token usage and PR throughput don&#8217;t tell the whole story, and the growing importance of documentation. We also examine the impact AI-driven pressure is having on burnout and cognitive overload.</span></p><p><span>Throughout the conversation, we share practical guidance for building engineering organizations where both developers and AI agents can do their best work.</span></p><div id="youtube2-U3p-rqOrAts" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;U3p-rqOrAts&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/U3p-rqOrAts?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h2><strong>Some takeaways: </strong></h2><h2><strong><span>Agent experience builds on developer experience</span></strong></h2><ul><li><p><strong><span>Agent experience focuses on creating the conditions for AI agents to succeed.</span></strong><span> Brian defines agent experience as the environment surrounding AI agents, including the quality of context, documentation, validation, and feedback they receive. As agents become part of software teams, improving those conditions becomes increasingly important.</span></p></li><li><p><strong><span>Model quality is only one part of successful AI adoption.</span></strong><span> Organizations often focus on choosing the best model, but Brian argues that context, clear intent, and effective workflows often have a greater impact on outcomes than incremental improvements in model capability.</span></p></li><li><p><strong><span>The same systems that help developers often help AI agents.</span></strong><span> Investments in documentation, development workflows, and engineering platforms create a stronger foundation for both humans and AI to produce high-quality work.</span></p></li></ul><h2><strong><span>Developer experience and agent experience don&#8217;t always align</span></strong></h2><ul><li><p><strong><span>Many improvements benefit both developers and AI agents.</span></strong><span> Better documentation, clearer context, and stronger engineering practices improve outcomes across the board, making existing developer experience investments even more valuable.</span></p></li><li><p><strong><span>Optimizing for one doesn&#8217;t automatically optimize for the other.</span></strong><span> Brian explains that organizations will increasingly encounter situations where workflows that help AI agents introduce friction for developers, or vice versa.</span></p></li><li><p><strong><span>Organizations should measure both independently.</span></strong><span> Rather than assuming every AI optimization improves the developer experience, engineering leaders should evaluate where the two reinforce each other and where they diverge.</span></p></li></ul><h2><strong><span>Preparing for AI requires organizational change</span></strong></h2><ul><li><p><strong><span>Successful AI adoption requires more than coding tools.</span></strong><span> Justin and Brian describe AI readiness as a combination of developer tooling, engineering platforms, and organizational practices rather than a single technology decision.</span></p></li><li><p><strong><span>The DX Core 4 still provides a useful foundation.</span></strong><span> Instead of abandoning existing engineering metrics, organizations should reinterpret them for AI-assisted development while continuing to focus on business outcomes rather than activity.</span></p></li><li><p><strong><span>Validation becomes more important as generation becomes easier.</span></strong><span> As AI produces more code, engineering organizations need stronger review, testing, and verification processes to ensure quality keeps pace with productivity.</span></p></li></ul><h2><strong><span>Documentation becomes infrastructure for AI agents</span></strong></h2><ul><li><p><strong><span>Documentation is no longer just for people.</span></strong><span> AI agents rely on high-quality documentation to understand systems, follow conventions, and complete work accurately, making documentation a core engineering asset rather than an afterthought.</span></p></li><li><p><strong><span>Not all documentation delivers equal value.</span></strong><span> Brian highlights that the biggest returns come from documenting information that helps agents understand systems, architecture, and engineering intent rather than simply producing more documentation.</span></p></li><li><p><strong><span>Capturing organizational knowledge improves both human and AI performance.</span></strong><span> Teams that make important context explicit reduce repeated questions, improve onboarding, and enable AI agents to work more effectively.</span></p></li></ul><h2><strong><span>Engineering metrics need to evolve with AI</span></strong></h2><ul><li><p><strong><span>Token usage is a cost metric, not a productivity metric.</span></strong><span> Brian cautions against treating token consumption as a measure of engineering effectiveness because it reflects AI usage rather than business value or software quality.</span></p></li><li><p><strong><span>PR throughput tells only part of the story.</span></strong><span> Larger pull requests and faster code generation may indicate increased AI adoption, but they can also increase review complexity and cognitive load if organizations measure throughput in isolation.</span></p></li><li><p><strong><span>Outcome metrics matter more than activity metrics.</span></strong><span> Justin emphasizes measuring whether engineering teams deliver value, improve quality, and create better developer experiences instead of rewarding raw AI utilization.</span></p></li></ul><h2><strong><span>AI changes how engineering work feels&#8212;not just how it&#8217;s done</span></strong></h2><ul><li><p><strong><span>AI pressure is contributing to burnout and cognitive overload.</span></strong><span> Brian describes growing pressure to move faster with AI while simultaneously reviewing larger code changes and maintaining confidence in increasingly AI-generated systems.</span></p></li><li><p><strong><span>Software engineering is much more than writing code.</span></strong><span> Even as AI accelerates code generation, engineers remain responsible for judgment, communication, system design, validation, and building trust in what gets shipped.</span></p></li><li><p><strong><span>The long-term challenge is balancing speed with confidence.</span></strong><span> Organizations that move faster than their ability to verify AI-generated work risk increasing technical debt, developer stress, and uncertainty rather than creating sustainable productivity gains.</span></p></li></ul><h2><strong>In this episode, we cover:</strong></h2><p>(<a href="https://www.youtube.com/watch?v=U3p-rqOrAts">00:00</a>) Intro</p><p>(<a href="https://www.youtube.com/watch?v=U3p-rqOrAts&amp;t=86s">01:26</a>) Justin&#8217;s new role at DX</p><p>(<a href="https://www.youtube.com/watch?v=U3p-rqOrAts&amp;t=233s">03:53</a>) What agent experience is and why engineering leaders should care</p><p>(<a href="https://www.youtube.com/watch?v=U3p-rqOrAts&amp;t=535s">08:55</a>) How to improve agent experience at the platform level</p><p>(<a href="https://www.youtube.com/watch?v=U3p-rqOrAts&amp;t=685s">11:25</a>) How agent experience and developer experience influence each other</p><p>(<a href="https://www.youtube.com/watch?v=U3p-rqOrAts&amp;t=881s">14:41</a>) Preparing engineering teams for agentic work</p><p>(<a href="https://www.youtube.com/watch?v=U3p-rqOrAts&amp;t=1298s">21:38</a>) Why the DX Core 4 still matters in the age of AI</p><p>(<a href="https://www.youtube.com/watch?v=U3p-rqOrAts&amp;t=1643s">27:23</a>) What PR throughput actually measures</p><p>(<a href="https://www.youtube.com/watch?v=U3p-rqOrAts&amp;t=1967s">32:47</a>) The limits of token metrics</p><p>(<a href="https://www.youtube.com/watch?v=U3p-rqOrAts&amp;t=2230s">37:10</a>) What the data shows about documentation and developer experience</p><p>(<a href="https://www.youtube.com/watch?v=U3p-rqOrAts&amp;t=2372s">39:32</a>) Improving documentation for AI agents</p><p>(<a href="https://www.youtube.com/watch?v=U3p-rqOrAts&amp;t=2443s">40:43</a>) AI-washing, burnout, and cognitive overload</p><p>(<a href="https://www.youtube.com/watch?v=U3p-rqOrAts&amp;t=2735s">45:35</a>) Brian&#8217;s upcoming research on agent experience</p><p><strong><span>Where to find Brian Houck:</span></strong></p><p><span>&#8226; LinkedIn: </span><a href="https://www.linkedin.com/in/brianhouck"><span>https://www.linkedin.com/in/brianhouck</span></a></p><p><strong><span>Where to find Justin Reock:</span></strong></p><p><span>&#8226; LinkedIn: </span><a href="https://www.linkedin.com/in/justinreock"><span>https://www.linkedin.com/in/justinreock</span></a></p><h2><strong>Referenced:</strong></h2><p><span>&#8226; </span><a href="https://getdx.com/corefour"><span>DX Core 4 Productivity Framework</span></a></p><p><span>&#8226; </span><a href="https://spawn-queue.acm.org/doi/10.1145/3807964"><span>The SPACE of AI | Queue</span></a></p><p><span>&#8226; </span><a href="https://www.linkedin.com/in/sarachizari/"><span>Sara Chizari on LinkedIn</span></a></p><p><span>&#8226; </span><a href="https://annievella.com/posts/the-middle-loop/"><span>The Middle Loop - Annie Vella</span></a></p><p><span>&#8226; </span><a href="https://github.com/gastownhall/gastown"><span>gastownhall/gastown: Gas Town - multi-agent workspace manager &#183; GitHub</span></a></p><p><span>&#8226; </span><a href="https://getdx.com/guide/dora-space-devex/"><span>DORA, SPACE, and DevEx: Which framework should you use?</span></a></p><p><span>&#8226; </span><a href="https://getdx.com/blog/ai-impact-report-q1-2026/"><span>AI Impact Report: Q1 2026</span></a></p><p><span>&#8226; </span><a href="https://newsletter.getdx.com/p/ai-authored-code-has-nearly-doubled"><span>AI-authored code has nearly doubled, but so has PR size</span></a></p><p><span>&#8226; </span><a href="https://spawn-queue.acm.org/doi/10.1145/3819080"><span>EngThrive: Make It Fast and Easy to Do Great Work: Building a durable model for outcome-oriented engineering measurement</span></a></p><p><span>&#8226; </span><a href="https://psychsafety.com/googles-project-aristotle/"><span>Google&#8217;s Project Aristotle - Psychological Safety</span></a></p><p><span>&#8226; </span><a href="https://zapier.com"><span>Zapier</span></a></p>]]></content:encoded></item><item><title><![CDATA[The State of AI Impact in Engineering: Q2 2026]]></title><description><![CDATA[Data from 500+ teams reveals that AI is delivering measurable velocity gains, but velocity alone isn't the story.]]></description><link>https://newsletter.getdx.com/p/the-state-of-ai-impact-in-engineering</link><guid isPermaLink="false">https://newsletter.getdx.com/p/the-state-of-ai-impact-in-engineering</guid><dc:creator><![CDATA[Justin Reock]]></dc:creator><pubDate>Wed, 22 Jul 2026 10:02:58 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/28554c65-1e7d-432a-9e62-b017f371f383_2400x1254.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em><strong><span>Welcome to the latest issue of Engineering Enablement,</span></strong><span> a weekly newsletter sharing research and perspectives on developer productivity.</span></em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.getdx.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.getdx.com/subscribe?"><span>Subscribe now</span></a></p><p><span>&#128467; </span><a href="https://getdx.com/webinar/ai-in-engineering-q2-2026-benchmarks-research-readout/?utm_source=newsletter"><span>Join me and Brian Houck on July 23</span></a><span> for a readout of this report, where we&#8217;ll discuss new findings from DX&#8217;s data on AI tool usage, spend, and impact across 500+ organizations. Register </span><a href="https://getdx.com/webinar/ai-in-engineering-q2-2026-benchmarks-research-readout/?utm_source=newsletter"><span>here.</span></a></p><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!i6Pp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F471c6aab-292d-4e15-b164-5be4355550ce_2400x1254.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!i6Pp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F471c6aab-292d-4e15-b164-5be4355550ce_2400x1254.png 424w, https://substackcdn.com/image/fetch/$s_!i6Pp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F471c6aab-292d-4e15-b164-5be4355550ce_2400x1254.png 848w, https://substackcdn.com/image/fetch/$s_!i6Pp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F471c6aab-292d-4e15-b164-5be4355550ce_2400x1254.png 1272w, https://substackcdn.com/image/fetch/$s_!i6Pp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F471c6aab-292d-4e15-b164-5be4355550ce_2400x1254.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!i6Pp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F471c6aab-292d-4e15-b164-5be4355550ce_2400x1254.png" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/471c6aab-292d-4e15-b164-5be4355550ce_2400x1254.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1634649,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://newsletter.getdx.com/i/205958887?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F471c6aab-292d-4e15-b164-5be4355550ce_2400x1254.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!i6Pp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F471c6aab-292d-4e15-b164-5be4355550ce_2400x1254.png 424w, https://substackcdn.com/image/fetch/$s_!i6Pp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F471c6aab-292d-4e15-b164-5be4355550ce_2400x1254.png 848w, https://substackcdn.com/image/fetch/$s_!i6Pp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F471c6aab-292d-4e15-b164-5be4355550ce_2400x1254.png 1272w, https://substackcdn.com/image/fetch/$s_!i6Pp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F471c6aab-292d-4e15-b164-5be4355550ce_2400x1254.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>We are excited to announce our Q2 2026 AI impact report.</span></p><p><span>When we first began tracking the impact of AI on engineering teams, our primary goal was to measure AI cohorts against historical baselines to answer the question of what happens to software output after adoption. With industry-wide AI adoption exceeding 90%, comparing AI users against a non-user control group is no longer a viable measurement strategy.</span></p><p><span>Engineering leaders are now under immense pressure to justify exponentially-increasing AI budgets. The data from our Q2 report reveals that while AI is delivering objective gains in velocity, those gains are highly uneven.</span></p><p><span>Download the full analysis </span><strong><a href="https://getdx.com/report/state-of-ai-impact-in-engineering-q2-report/?utm_source=newsletter"><span>here.</span></a></strong></p><p><span>In the new report, we&#8217;ve uncovered a number of critical trends, including:</span></p><p><strong><span>1. Over 50% of code is now generated by AI. </span></strong><span> This metric has accelerated rapidly, increasing from 34% in Q1 2026 to 52% in Q2 2026. This steep trajectory indicates that once AI tools are deployed, the code they generate rapidly scales across codebases, frequently moving through reviews, dependencies, and shared workflows.</span></p><p><strong><span>2. Quality may be declining. </span></strong><span>During the same period that AI adoption has increased, median pull request sizes have nearly doubled. Increases in PR size can serve as an early indicator of technical debt, as higher code volumes generally correlate with increased complexity and potential for bugs. This trend can also introduce additional friction in the review process, as more lines of code generated means more lines of code to review.</span></p><p><strong><span>3. Some aspects of developer experience are declining.</span></strong><span> The Developer Experience Index (DXI) dropped from 67 to 65 over four quarters. AI is improving some aspects of the developer experience&#8212;documentation quality, code maintainability, onboarding speed&#8212;while creating new friction in others: larger PRs, slower reviews, less incremental delivery. In aggregate, the net effect is currently negative. Velocity metrics alone will tell you things are improving. Developer experience metrics will tell you whether that&#8217;s actually true</span><strong><span>.</span></strong></p><p><strong><span>4. AI is making codebases easier to understand, but it&#8217;s also making the code it generates harder to trust. </span></strong><span>The Q2 data highlights a striking divergence between two historically correlated software quality metrics. Specifically, from Q1 2026, Code Maintainability improved by 3.8%, whereas Change Confidence decreased by 6.1%. Code Maintainability indicates how easily developers can understand the codebase, while Change Confidence measures their trust that modifications won&#8217;t cause production failures. Traditionally, highly maintainable code results in higher confidence when making changes. However, this data reveals a new tension: although AI helps developers understand the code in front of them, they exhibit less trust in the code they are pushing to production.</span></p><p><strong><span>5. Saved time isn&#8217;t converting into innovation. </span></strong><span>AI users are now saving an estimated 4 to 6 hours per week. However, the innovation ratio, defined as the percentage of time spent on building new features versus maintenance and overhead, has remained flat over the same period of study. This flat trend indicates that the time saved by AI is not currently converting into increased capacity for creating new value. Leaders should keep a close eye on this metric over time. Ideally, innovation ratio will increase as AI frees up engineers to work on more new features.</span></p><p><strong><span>6. AI spend is accelerating faster than outcomes.</span></strong><span> Median quarterly organizational AI spend climbed from ~$1.5K to ~$44K over four quarters. Tech-sector spend increased nearly 28x. These numbers will draw scrutiny. Leaders who cannot connect this investment to downstream outcomes (feature velocity, innovation ratio, quality) may face increasingly difficult budget conversations in the back half of 2026.</span></p><h3><span>What this means for leaders</span></h3><p><span>The Q2 2026 data indicates that the industry is shifting from base AI deployment to evaluating concrete return on investment. As AI expenditures accelerate, engineering leaders must shift their focus from simply acquiring AI tools to optimizing the surrounding development pipelines and resolving systemic bottlenecks. To achieve true ROI, leaders must ensure that saved hours are reinvested into product innovation rather than absorbed by existing organizational friction.</span></p><p><span>To explore the full data and benchmark your team against 500+ organizations on measures of throughput, quality, and AI tooling cost, </span><strong><a href="https://getdx.com/report/state-of-ai-impact-in-engineering-q2-report/"><span>download the full report here.</span></a></strong></p><div><hr></div><p>That&#8217;s it for this week. Thanks for reading.</p><p>-Justin</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.getdx.com/p/the-state-of-ai-impact-in-engineering?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.getdx.com/p/the-state-of-ai-impact-in-engineering?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p>]]></content:encoded></item><item><title><![CDATA[UKG’s system for driving effective AI use]]></title><description><![CDATA[How they created internal scorecards that managers could use to coach and guide their teams&#8217; AI use.]]></description><link>https://newsletter.getdx.com/p/ukgs-system-for-driving-effective-ai-use</link><guid isPermaLink="false">https://newsletter.getdx.com/p/ukgs-system-for-driving-effective-ai-use</guid><dc:creator><![CDATA[Abi Noda]]></dc:creator><pubDate>Wed, 15 Jul 2026 10:03:30 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/cb3e2bb4-e068-4186-8d86-44ec3911e33a_2400x1260.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em><strong><span>Welcome to the latest issue of Engineering Enablement,</span></strong><span> a weekly newsletter sharing research and perspectives on developer productivity.</span></em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.getdx.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.getdx.com/subscribe?"><span>Subscribe now</span></a></p><p><span>&#128467; Join DX Deputy CTO, Justin Reock and Distinguished Scientist, Brian Houck on July 23 for a readout of the Q2 State of AI Impact in Engineering Report. Register </span><a href="https://getdx.com/webinar/ai-in-engineering-q2-2026-benchmarks-research-readout/?utm_source=newsletter"><span>here.</span></a></p><div><hr></div><p><span>I recently sat down with </span><a href="https://www.linkedin.com/in/thomas-newton-6821985/"><span>Thomas Newton</span></a><span>, VP of Engineering at UKG, to discuss how his team built a system to guide AI adoption and assess whether it&#8217;s translating into meaningful engineering outcomes. What was particularly interesting was how UKG provides teams with a metrics dashboard that managers can use to have better coaching conversations, and ultimately help developers become more effective with AI tools.</span></p><p><span>For this week&#8217;s newsletter, Thomas is describing how their approach works.</span></p><p><span>Here&#8217;s Thomas.</span></p><div><hr></div><p><strong><span>Thomas: </span></strong><span>As a leader in our industry, we previously faced a question that most engineering organizations are grappling with right now: how do you measure whether AI usage is translating into more meaningful work shipped?</span></p><p><span>This was a leadership priority with one goal from the start: make sure that the right data, in a digestible format, landed with the people closest to the work&#8212;the managers.</span></p><p><span>The result is what we call the Manager AI Adoption Dashboard: a set of visuals that combine AI usage patterns, delivery outcomes, and spend into a single picture that engineering managers can use to coach their teams, guide adoption, and have better conversations about how work is getting done.</span></p><h3><span>Starting with experimentation and adoption</span></h3><p><span>Our journey using AI tools in product development started the way most do. We experimented with several tools (GitHub Copilot, Windsurf, and others) before making a significant push toward Claude.</span></p><p><span>Giving our engineers access to these tools was a good place to start, but our managers needed visibility. They could feel the productivity gains anecdotally, but the data was missing. We wanted to better understand where AI was helping teams, and which workflows were creating the most impact.</span></p><p><span>The shift to a consumption-based model made this need even more urgent. Unlike fixed-cost seat licenses, consumption pricing means the meter is always running. Leaders needed to understand not just whether teams were using AI, but whether the investment was producing returns.</span></p><h3><span>Deciding what to measure</span></h3><p><span>One of the earliest decisions we made was to anchor the dashboard around TrueThroughput, a metric developed by DX that goes beyond raw pull request counts to account for the relative complexity and size of work delivered.</span></p><p><span>TrueThroughput uses AI to classify the complexity of different tasks, giving you a size-adjusted throughput number. Think of it like a weighted GPA versus an unweighted GPA. Both are useful, but the weighted version tells you whether someone is delivering meaningful work or just merging a lot of five-second fixes.</span></p><p><span>Instead of indexing on consumption, the dashboard creates a balanced view of AI impact by correlating consistent AI usage, measured in days of use, with TrueThroughput. These metrics help us answer whether consistent use of AI is helping teams deliver more meaningful work.</span></p><p><span>The dashboard combines three signals:</span></p><ul><li><p><span>AI usage: How consistently someone is using AI tools in their workflow.</span></p></li><li><p><span>TrueThroughput: The volume and complexity of work delivered.</span></p></li><li><p><span>Spend: Awareness of AI investment and consumption patterns.</span></p></li></ul><h3><span>Inside the dashboard</span></h3><p><span>The dashboard was built around a quadrant view. The horizontal axis tracks consistent days of AI use over a 30-day window: fewer than 15 days puts you on the left, more than 15 on the right. The vertical axis tracks TrueThroughput.</span></p><p><span>Example for illustration purpose only:</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OXcN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26504830-3b37-4cee-b44f-06f81eef7f26_2048x1611.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OXcN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26504830-3b37-4cee-b44f-06f81eef7f26_2048x1611.png 424w, https://substackcdn.com/image/fetch/$s_!OXcN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26504830-3b37-4cee-b44f-06f81eef7f26_2048x1611.png 848w, https://substackcdn.com/image/fetch/$s_!OXcN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26504830-3b37-4cee-b44f-06f81eef7f26_2048x1611.png 1272w, https://substackcdn.com/image/fetch/$s_!OXcN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26504830-3b37-4cee-b44f-06f81eef7f26_2048x1611.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OXcN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26504830-3b37-4cee-b44f-06f81eef7f26_2048x1611.png" width="1456" height="1145" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/26504830-3b37-4cee-b44f-06f81eef7f26_2048x1611.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1145,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!OXcN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26504830-3b37-4cee-b44f-06f81eef7f26_2048x1611.png 424w, https://substackcdn.com/image/fetch/$s_!OXcN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26504830-3b37-4cee-b44f-06f81eef7f26_2048x1611.png 848w, https://substackcdn.com/image/fetch/$s_!OXcN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26504830-3b37-4cee-b44f-06f81eef7f26_2048x1611.png 1272w, https://substackcdn.com/image/fetch/$s_!OXcN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26504830-3b37-4cee-b44f-06f81eef7f26_2048x1611.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>That creates four zones:</span></p><ul><li><p><strong><span>Exploring (bottom left):</span></strong><span> Lower AI adoption, developing throughput. Teams discovering what problems AI might solve.</span></p></li><li><p><strong><span>Learning (bottom right):</span></strong><span> High AI adoption, building throughput. Expected pattern as teams explore different applications over 2-3 months.</span></p></li><li><p><strong><span>Efficient (top left):</span></strong><span> High throughput without heavy AI adoption. Some roles and workflows don&#8217;t need AI.</span></p></li><li><p><strong><span>Amplified (top right):</span></strong><span> High AI adoption, high throughput. Patterns here show which applications create real impact.</span></p></li></ul><p><span>Each dot on the chart represents an individual. Bubble size reflects spend. And the views are drillable: from the entire organization down to a business unit, a team, and ultimately an individual manager&#8217;s direct reports.</span></p><h3><span>What the data revealed</span></h3><p><span>Since rolling out the dashboard, the results have been striking. Initially, as one might expect with newly introduced tooling, the bottom-left quadrant was packed, meaning that a large portion of our engineering organization hadn&#8217;t touched AI tools at all. Within four months,, that quadrant was nearly empty; less than 1% of the organization remained in the low-AI, high-throughput zone. It&#8217;s been exciting to watch the whole organization shift on this.</span></p><blockquote><p>&#8220;Less than 1% of the organization remained in the low-AI, high-throughput zone.&#8221;</p></blockquote><p><span>We also saw some things in the data that confirmed what we&#8217;d hypothesized. For example, the group seeing the biggest throughput gains were our senior and principal engineers, at ~20-30% above what other roles saw. That&#8217;s intuitive, but it was interesting to see it in the data. Senior engineers already have strong instincts for where AI can help and where it can&#8217;t. They were often quicker to identify high-leverage opportunities and incorporate the tools into existing workflows.</span></p><p><span>One pattern caught us off guard: managers and directors started writing code again. Our leadership population started doing more direct hands-on coding. They have more assistants, they can multitask better, whatever the reason, it&#8217;s a clear trend. That pattern extends beyond engineering: our product managers and designers have started leaning into AI tools too, getting comfortable with the terminal, creating digital artifacts, and contributing in ways that show up in delivery metrics. When we started this, it wasn&#8217;t what we anticipated, but it&#8217;s a trend that is now more commonly observed and discussed.</span></p><h3><span>How managers use the dashboard</span></h3><p><span>The real value of the dashboard is the quality of the conversations the data enables.</span></p><p><span>For managers, the first question often focused on adoption: &#8220;How do we get you from left to right?&#8221; Are you using it daily, is it part of your workflow, or are you still finding your footing with it? But over time, the conversation became less about adoption itself and more about impact. Where is AI helping? Which workflows are working well? Where is it creating leverage, and where is it not?</span></p><p><span>When a manager saw an engineer with high AI usage but flat throughput, the right response wasn&#8217;t to question the spend. It was to ask what they were working on; maybe they were ramping up on new AI workflows, or in a role where their AI-assisted work didn&#8217;t produce code commits.  One example is heavy operation roles where productivity didn&#8217;t always result in commits into GitHub.</span></p><p><span>The best way to head off gaming or surveillance concerns is proactive communication. Don&#8217;t let people fill in the blanks on what the dashboards are for or why they exist. Be very clear. The dashboard is a conversation starter, not a performance management system. If teams interpret the dashboards as a tool for punishment or reward, they&#8217;ll have every incentive to game the numbers.</span></p><blockquote><p>&#8220;The dashboard is a conversation starter, not a performance management system.&#8221;</p></blockquote><p><span>At UKG, the framing has been consistent from day one: the dashboard exists to help managers empower their team, understand the work, understand how AI is being applied to that work, and help people lean into a new way of working.</span></p><h3><span>Keeping AI spend in check without leading with cost</span></h3><p><span>AI spend is visible on the dashboard, but it&#8217;s deliberately not the headline metric. It&#8217;s an awareness layer, something managers should be conscious of, not something that drives the conversation.</span></p><p><span>We use what I call &#8220;circuit breakers&#8221;; daily budget controls that flag when usage spikes above a threshold. But we haven&#8217;t yet landed on a firm benchmark for what reasonable per-engineer spend looks like.</span></p><p><span>The range is just too wide right now. Some teams are considering multi-agent orchestration and long-running loops that transform an entire codebase overnight. Others are using AI to finish a feature today. At this stage, it&#8217;s really hard to generalize.</span></p><p><span>I expect spending patterns will stabilize as the initial learning ramp flattens. Your first couple of prompts will probably be expensive. You&#8217;re still learning how to use the models correctly, you&#8217;re playing, you&#8217;re understanding how to use this radical new thing. But in as little as a few months, you will start to see averages form. We&#8217;re still in an experimentation phase, and for now, the approach is pragmatic: watch carefully, trust manager judgment, and make sure the investment is going toward the right outcomes.</span></p><h3><span>What comes next</span></h3><p><span>We&#8217;re already thinking about the limits of what the current dashboard captures. TrueThroughput is powerful for teams that ship code, but it misses productivity gains in operations, SRE, and other roles where AI is being used to correlate incidents, search logs, and accelerate incident resolution, work that never ends up in a pull request.</span></p><p><span>We have teams where an engineer uses AI to search past incidents during a live outage and correlate similar patterns to get to a resolution faster. That&#8217;s enormously productive, but it won&#8217;t show up in throughput. We&#8217;re trying to think through what the next level of digital footprints looks like, the metrics that capture the full picture of AI-enabled productivity, not just the code-commit slice of it.</span></p><p><span>The dashboard is a living system designed to evolve as we learn more about what effective AI-enabled engineering looks like.</span></p><h2><span>Final thoughts:</span></h2><p><span>For engineering leaders at other organizations who haven&#8217;t yet started measuring AI adoption, my advice is simple: Measure something. Metrics you have access to might differ, but collect some form of data, figure out what makes sense for your organization, and don&#8217;t make it a binary decision based on the data itself. The data should enable leaders to have further conversations.</span></p><p><span>That&#8217;s an important takeaway: The tools are powerful, and the data is illuminating, but the transformation happens in the conversations between managers and their teams.</span></p><div><hr></div><p><em><span>If you have questions about UKG&#8217;s approach, or just want to hear more from Thomas, make sure to follow or </span><a href="https://www.linkedin.com/in/thomas-newton-6821985/"><span>connect with him on LinkedIn</span></a><span>.</span></em></p><div><hr></div><p><span>That&#8217;s it for this week. Thanks for reading.</span></p><p><span>-Abi</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.getdx.com/p/ukgs-system-for-driving-effective-ai-use?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.getdx.com/p/ukgs-system-for-driving-effective-ai-use?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p>]]></content:encoded></item><item><title><![CDATA[Adopting the product operating model at Priceline]]></title><description><![CDATA[How Priceline used developer experience metrics, organizational change, and a product operating model to improve engineering effectiveness and prepare for AI-driven software development.]]></description><link>https://newsletter.getdx.com/p/adopting-the-product-operating-model</link><guid isPermaLink="false">https://newsletter.getdx.com/p/adopting-the-product-operating-model</guid><dc:creator><![CDATA[Justin Reock]]></dc:creator><pubDate>Fri, 10 Jul 2026 15:38:54 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/205653265/2adc0bae43ca43be00e7327ced84f67e.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>Listen and watch now on <strong><a href="https://youtu.be/c-O1wrEjx6w">YouTube</a>, <a href="https://podcasts.apple.com/us/podcast/engineering-enablement-by-abi-noda/id1619140476">Apple</a>, and <a href="https://open.spotify.com/show/3NxjyIsuxeDMQtisDqBy7D">Spotify</a></strong>.</p><p><span>In this episode of the Engineering Enablement podcast, I sit down with Sejal Amin, Chief Technology Officer at Priceline, and Pedro Gutierrez, Senior Director of Software Engineering, to discuss how Priceline adopted a product operating model and the role developer experience played in making that transformation successful.</span></p><p><span>We explore why the company moved away from a project-based approach, how DX metrics and developer feedback helped uncover organizational bottlenecks, and why a phased rollout, clear communication, and empowered engineering managers were critical to building trust and improving developer experience. We also discuss creating a dedicated developer experience team, lessons learned throughout the transformation, and how Priceline&#8217;s product operating model has helped the organization adapt to AI-driven software development.</span></p><div id="youtube2-c-O1wrEjx6w" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;c-O1wrEjx6w&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/c-O1wrEjx6w?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h2><strong>Some takeaways: </strong></h2><p><strong><span>Developer experience as a driver of organizational change</span></strong></p><ul><li><p><strong><span>Developer experience data can reveal organizational problems that traditional engineering metrics miss.</span></strong><span> At Priceline, DX signals uncovered organizational bottlenecks&#8212;including handoffs, dependencies, and team friction&#8212;that ultimately led the company to adopt a product operating model.</span></p></li><li><p><strong><span>Developer experience should be treated as a strategic capability, not just an engineering metric.</span></strong><span> Rather than measuring developer satisfaction in isolation, Priceline used DX insights to guide structural changes that improved autonomy, delivery, and engineering culture.</span></p></li></ul><p><strong><span>Adopting a product operating model</span></strong></p><ul><li><p><strong><span>Reducing dependencies gives teams greater ownership.</span></strong><span> Priceline shifted from a project-based organization to cross-functional product teams, reducing handoffs and giving teams the people and capabilities needed to own outcomes end to end.</span></p></li><li><p><strong><span>Autonomy requires visibility into team health.</span></strong><span> DX metrics gave engineering managers a clear view of the obstacles affecting their teams, allowing them to improve local workflows while staying aligned with broader organizational goals.</span></p></li></ul><p><strong><span>Turning developer feedback into action</span></strong></p><ul><li><p><strong><span>Developer experience surveys should lead to action&#8212;not just measurement.</span></strong><span> Managers reviewed survey results, completed a triage process, created quarterly action plans, and measured whether those improvements had an impact in the next survey cycle.</span></p></li><li><p><strong><span>Small workflow improvements can have an outsized impact.</span></strong><span> DX data helped teams reclaim focus time, identify tooling regressions after migrations, surface cross-team dependencies, and address day-to-day friction before it became systemic.</span></p></li></ul><p><strong><span>Building trust in developer experience metrics</span></strong></p><ul><li><p><strong><span>Clear communication is essential for adoption.</span></strong><span> Leaders consistently reinforced that DX metrics existed to improve teams rather than evaluate individuals, helping build confidence in the process from the outset.</span></p></li><li><p><strong><span>Trust grows when developers see meaningful change.</span></strong><span> Acting on feedback quarter after quarter encouraged greater participation, strengthened psychological safety, and made developer experience part of the organization&#8217;s culture.</span></p></li></ul><p><strong><span>The evolving role of engineering managers</span></strong></p><ul><li><p><strong><span>Engineering managers became owners of developer experience.</span></strong><span> Managers were expected to understand DX data, improve their team&#8217;s DXI each quarter, and make developer experience part of their regular operating rhythm.</span></p></li><li><p><strong><span>Developer experience became part of everyday engineering leadership.</span></strong><span> DX metrics were discussed openly in all-hands meetings and other forums, making developer experience a visible measure of organizational health rather than a one-time initiative.</span></p></li></ul><p><strong><span>Preparing engineering organizations for AI</span></strong></p><ul><li><p><strong><span>AI changes where bottlenecks occur&#8212;not whether they exist.</span></strong><span> As AI accelerated code generation, Priceline used its product operating model and developer experience data to identify where constraints had shifted and respond accordingly.</span></p></li><li><p><strong><span>A strong operating model helps organizations adapt to AI.</span></strong><span> Autonomous teams, continuous measurement, and visibility into developer workflows allowed Priceline to embrace AI while continuing to improve flow across the software development lifecycle.</span></p></li></ul><h2><strong>In this episode, we cover:</strong></h2><p>(<a href="https://www.youtube.com/watch?v=c-O1wrEjx6w">00:00</a>) Intro</p><p>(<a href="https://www.youtube.com/watch?v=c-O1wrEjx6w&amp;t=67s">01:07</a>) Meet Sejal Amin and Pedro Gutierrez</p><p>(<a href="https://www.youtube.com/watch?v=c-O1wrEjx6w&amp;t=107s">01:47</a>) How Priceline&#8217;s developer experience journey began</p><p>(<a href="https://www.youtube.com/watch?v=c-O1wrEjx6w&amp;t=295s">04:55</a>) Lessons from Priceline&#8217;s first developer experience surveys</p><p>(<a href="https://www.youtube.com/watch?v=c-O1wrEjx6w&amp;t=415s">06:55</a>) How DX improved Priceline&#8217;s developer experience surveys</p><p>(<a href="https://www.youtube.com/watch?v=c-O1wrEjx6w&amp;t=587s">09:47</a>) Identifying the causes of organizational slowness</p><p>(<a href="https://www.youtube.com/watch?v=c-O1wrEjx6w&amp;t=753s">12:33</a>) How the product operating model changed the way Priceline works</p><p>(<a href="https://www.youtube.com/watch?v=c-O1wrEjx6w&amp;t=850s">14:10</a>) Priceline&#8217;s phased rollout with DX</p><p>(<a href="https://www.youtube.com/watch?v=c-O1wrEjx6w&amp;t=1094s">18:14</a>) How DX insights drove organizational changes</p><p>(<a href="https://www.youtube.com/watch?v=c-O1wrEjx6w&amp;t=1173s">19:33</a>) Why Priceline improved developer experience before org change was complete</p><p>(<a href="https://www.youtube.com/watch?v=c-O1wrEjx6w&amp;t=1338s">22:18</a>) How clear communication builds trust</p><p>(<a href="https://www.youtube.com/watch?v=c-O1wrEjx6w&amp;t=1465s">24:25</a>) Early results from Priceline&#8217;s Core Four</p><p>(<a href="https://www.youtube.com/watch?v=c-O1wrEjx6w&amp;t=1538s">25:38</a>) Creating a culture of continuous feedback to build trust</p><p>(<a href="https://www.youtube.com/watch?v=c-O1wrEjx6w&amp;t=1660s">27:40</a>) What has changed in the engineering manager role</p><p>(<a href="https://www.youtube.com/watch?v=c-O1wrEjx6w&amp;t=1810s">30:10</a>) Resources for learning about the product operating model</p><p>(<a href="https://www.youtube.com/watch?v=c-O1wrEjx6w&amp;t=1960s">32:40</a>) What Pedro learned from implementing DX</p><p>(<a href="https://www.youtube.com/watch?v=c-O1wrEjx6w&amp;t=2091s">34:51</a>) The developer experience team</p><p>(<a href="https://www.youtube.com/watch?v=c-O1wrEjx6w&amp;t=2159s">35:59</a>) How AI tools have impacted Priceline&#8217;s teams</p><p>(<a href="https://www.youtube.com/watch?v=c-O1wrEjx6w&amp;t=2240s">37:20</a>) How the product operating model supports AI-driven development</p><p>(<a href="https://www.youtube.com/watch?v=c-O1wrEjx6w&amp;t=2353s">39:13</a>) Final advice for engineering leaders</p><p><strong><span>Where to find Sejal Amin:</span></strong></p><p><span>&#8226; LinkedIn: </span><a href="https://www.linkedin.com/in/sejal-amin"><span>https://www.linkedin.com/in/sejal-amin</span></a></p><p><strong><span>Where to find Pedro Gutierrez:</span></strong></p><p><span>&#8226; LinkedIn: </span><a href="https://www.linkedin.com/in/pedro-gutierrez-b6605422"><span>https://www.linkedin.com/in/pedro-gutierrez-b6605422</span></a></p><p><strong><span>Where to find Justin Reock:</span></strong></p><p><span>&#8226; LinkedIn: </span><a href="https://www.linkedin.com/in/justinreock"><span>https://www.linkedin.com/in/justinreock</span></a></p><h2><strong>Referenced:</strong></h2><p><span>&#8226; </span><a href="https://getdx.com/research/measuring-developer-productivity-with-the-dx-core-4/"><span>Measuring developer productivity with the DX Core 4</span></a></p><p><span>&#8226; </span><a href="https://www.amazon.com/dp/1119697336?lv=shuf&amp;channelId=500&amp;plpRedirect=mhFallback"><span>Transformed: Moving to the Product Operating Model (Silicon Valley Product Group)</span></a></p><p><span>&#8226; </span><a href="https://teamtopologies.com/"><span>Team Topologies</span></a></p><p><span>&#8226; </span><a href="https://flowframework.org/"><span>Flow Framework</span></a></p><p><span>&#8226; </span><a href="https://www.amazon.com/dp/1942788398?lv=shuf&amp;channelId=500&amp;plpRedirect=mhFallback"><span>Project to Product: How to Survive and Thrive in the Age of Digital Disruption with the Flow Framework</span></a></p>]]></content:encoded></item><item><title><![CDATA[Five studies changing how I think about AI in software engineering]]></title><description><![CDATA[AI compressed the upstream work. What does that mean for everything downstream?]]></description><link>https://newsletter.getdx.com/p/five-studies-that-are-changing-how</link><guid isPermaLink="false">https://newsletter.getdx.com/p/five-studies-that-are-changing-how</guid><dc:creator><![CDATA[Brian Houck]]></dc:creator><pubDate>Fri, 10 Jul 2026 13:04:17 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/7a798890-bc91-4dc8-87f9-656e7f2f5f13_2400x1260.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em><strong><span>Welcome to the latest issue of Engineering Enablement,</span></strong><span> a weekly newsletter sharing research and perspectives on developer productivity.</span></em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.getdx.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.getdx.com/subscribe?"><span>Subscribe now</span></a></p><p><span>&#128467; </span><a href="https://getdx.com/webinar/ai-in-engineering-q2-2026-benchmarks-research-readout/?utm_source=newsletter"><span>Join me on July 23</span></a><span> for a readout of the upcoming </span>State of AI Impact in Engineering: Q2 Report<span>. We&#8217;ll discuss new findings from DX&#8217;s data on AI tool usage, spend, and impact across 500+ organizations. Register </span><a href="https://getdx.com/webinar/ai-in-engineering-q2-2026-benchmarks-research-readout/?utm_source=newsletter"><span>here.</span></a></p><div><hr></div><p><span>Every once in a while, several independent papers arrive at roughly the same time and collectively tell a bigger story than any one of them does alone. This week, I&#8217;m sharing five recent papers that have significantly influenced how I&#8217;m thinking about AI and software engineering.</span></p><p><span>Each paper tackles a different question. Some measure the productivity impact of AI coding assistants. Others examine how those gains propagate through the software delivery process, explore what developers actually want from future AI systems, or reconsider the kinds of debt we should be paying attention to in an AI-assisted world.</span></p><p><span>Despite coming from different research groups and using very different methodologies, they all seem to be converging on the same underlying story.</span></p><p><span>AI is compressing the upstream work of software engineering. The more I sat with these papers, the less I found myself asking, &#8220;Is AI making developers faster?&#8221; and the more I found myself asking, &#8220;What happens after the code is written?&#8221; Are we actually shipping more value? Where do the new bottlenecks emerge? And what are the costs if understanding can&#8217;t keep pace with generation?</span></p><p><span>After reading these five papers, I came away with one overarching conclusion: we&#8217;re generating code faster than we&#8217;re generating the systems needed to safely understand, verify, and deliver it.</span></p><p><span>A quick note on disclosure: three of these papers come from people I know and work with extensively. None of the papers are mine.</span></p><p><span>Here they are, in the order I&#8217;d recommend reading them.</span></p><h3><span>1. GitHub Copilot and Developer Productivity</span></h3><p><em><span>Paper: Heilman, A., Kyllo, A., Murphy-Hill, E. </span><a href="https://arxiv.org/abs/2606.00438"><span>GitHub Copilot and Developer Productivity: An Observational Dose-Response Analysis.</span></a></em></p><p><a href="https://arxiv.org/abs/2606.00438"><span>The first paper</span></a><span> I want to highlight tackles the familiar question of whether GitHub Copilot makes developers more productive, but it does so with one of the more clever research designs I&#8217;ve seen.</span></p><p><span>Rather than simply comparing Copilot users to non-users (which are getting harder and harder to find), the authors control for Active Coding Time (i.e., how much time developers spend actively engaging with development tools) and examine how productivity changes within the same engineer over 43 weeks across a population of 16,223 developers.</span></p><p><span>The payoff of this design is that it compares engineers to themselves rather than to one another. Using that approach, the authors found that weeks with the highest Copilot usage were associated with ~40% more completed PRs per hour of coding time than weeks with no usage.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!aTuc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14c97306-fdec-49e1-9741-6da9e90a0745_2048x1142.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!aTuc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14c97306-fdec-49e1-9741-6da9e90a0745_2048x1142.png 424w, https://substackcdn.com/image/fetch/$s_!aTuc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14c97306-fdec-49e1-9741-6da9e90a0745_2048x1142.png 848w, https://substackcdn.com/image/fetch/$s_!aTuc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14c97306-fdec-49e1-9741-6da9e90a0745_2048x1142.png 1272w, https://substackcdn.com/image/fetch/$s_!aTuc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14c97306-fdec-49e1-9741-6da9e90a0745_2048x1142.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!aTuc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14c97306-fdec-49e1-9741-6da9e90a0745_2048x1142.png" width="1456" height="812" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/14c97306-fdec-49e1-9741-6da9e90a0745_2048x1142.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:812,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!aTuc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14c97306-fdec-49e1-9741-6da9e90a0745_2048x1142.png 424w, https://substackcdn.com/image/fetch/$s_!aTuc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14c97306-fdec-49e1-9741-6da9e90a0745_2048x1142.png 848w, https://substackcdn.com/image/fetch/$s_!aTuc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14c97306-fdec-49e1-9741-6da9e90a0745_2048x1142.png 1272w, https://substackcdn.com/image/fetch/$s_!aTuc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14c97306-fdec-49e1-9741-6da9e90a0745_2048x1142.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>The relationship showed a clear dose-response pattern (a way to do a causal analysis, once everyone is already using the tools). More Copilot engagement was associated with more PR throughput, although the gains appeared to level off at very high usage.</span></p><p><span>The authors ran seven robustness and falsification tests to rule out alternative explanations (team-level effects, generic AI engagement, PR slicing, shifts toward easier work). The positive association remained remarkably consistent.</span></p><p><span>Interestingly, the gains were not concentrated in tiny PRs. The strongest effects were observed for larger PRs (7+ files), arguing against the idea that developers are simply breaking work into smaller units.</span></p><p><span>It&#8217;s a thoughtful analysis and shows that we&#8217;re not just coding more, we&#8217;re increasing coding efficiency as well. These findings anchor many of the studies that follow in this roundup.</span></p><h3><span>2. Writing Code vs. Shipping Code: Productivity Effects Across Generations of AI Coding Tools</span></h3><p><em><span>Paper: Demirer, M., Musolff, L., Yang, L. </span><a href="https://www.nber.org/papers/w35275"><span>Writing Code vs. Shipping Code: Productivity Effects Across Generations of AI Coding Tools.</span></a></em></p><p><a href="https://www.nber.org/papers/w35275"><span>The next paper</span></a><span> I&#8217;m highlighting was published by the National Bureau of Economic Research. It analyzes AI adoption across 100,000+ GitHub developers and asks a more nuanced question than Heilman&#8217;s: when AI makes individual coding steps faster, how much of that gain actually survives all the way to shipped software?</span></p><p><span>The authors examine how AI productivity gains propagate through a hierarchy of software development: lines of code &#8594; files &#8594; commits &#8594; pull requests &#8594; projects/repos &#8594; releases.</span></p><p><span>They found that AI is clearly increasing coding activity, and the gains grow with each generation of tools. They estimate roughly +40% more commits from autocomplete, growing to +140% from interactive coding agents, and finally +180% from autonomous agents.</span></p><p><span>However, those gains fall off significantly as work moves through the software delivery process. The largest effects are seen in code generation, but smaller effects appear in repos touched, small still in releases shipped, and ultimately software consumed by users. Even with very large increases in coding activity, the effect on shipped software is much smaller, topping out at roughly +30% more releases. This is illustrated in figure 2 below.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pzuI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59f761e2-1b48-472f-8aa8-8a3b33c0a4a0_2048x1330.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pzuI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59f761e2-1b48-472f-8aa8-8a3b33c0a4a0_2048x1330.png 424w, https://substackcdn.com/image/fetch/$s_!pzuI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59f761e2-1b48-472f-8aa8-8a3b33c0a4a0_2048x1330.png 848w, https://substackcdn.com/image/fetch/$s_!pzuI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59f761e2-1b48-472f-8aa8-8a3b33c0a4a0_2048x1330.png 1272w, https://substackcdn.com/image/fetch/$s_!pzuI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59f761e2-1b48-472f-8aa8-8a3b33c0a4a0_2048x1330.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pzuI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59f761e2-1b48-472f-8aa8-8a3b33c0a4a0_2048x1330.png" width="1456" height="946" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/59f761e2-1b48-472f-8aa8-8a3b33c0a4a0_2048x1330.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:946,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!pzuI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59f761e2-1b48-472f-8aa8-8a3b33c0a4a0_2048x1330.png 424w, https://substackcdn.com/image/fetch/$s_!pzuI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59f761e2-1b48-472f-8aa8-8a3b33c0a4a0_2048x1330.png 848w, https://substackcdn.com/image/fetch/$s_!pzuI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59f761e2-1b48-472f-8aa8-8a3b33c0a4a0_2048x1330.png 1272w, https://substackcdn.com/image/fetch/$s_!pzuI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59f761e2-1b48-472f-8aa8-8a3b33c0a4a0_2048x1330.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>One of the findings I found most interesting is that they estimate a low elasticity of substitution (~0.25) between AI-generated output and human effort. That&#8217;s an economics concept that measures how replaceable human work is with AI output. As a methodology nerd, and someone with an economics degree, I found this particularly clever &#8212; they infer this elasticity from how AI productivity gains attenuate across the delivery process. Their estimate suggests AI and human work are still largely complements rather than substitutes, with substantial human effort still required to review, integrate, validate, and ship software.</span></p><p><span>One open question is whether the observed fall-off through the delivery process is some fundamental limit of software engineering, or simply the fact that organizations have not yet adapted their processes to an agentic world.</span></p><p><span>If Heilman tells you Copilot is making engineers measurably faster, this paper asks the harder question: faster at what, exactly?</span></p><h3><span>3. The Impact of AI Coding Assistants on Software Engineering</span></h3><p><em><span>Paper: Vella, A., Blincoe, K. </span><a href="https://arxiv.org/abs/2605.23135"><span>The Impact of AI Coding Assistants on Software Engineering: A Longitudinal Study.</span></a></em></p><p><span>The next study I want to highlight is unique because it isn&#8217;t just a snapshot in time, it&#8217;s a six-month </span><a href="https://arxiv.org/abs/2605.23135"><span>longitudinal study</span></a><span> of 95 professional software engineers. It also calls into question a relationship that we&#8217;ve long believed to be a bedrock of developer experience.</span></p><p><span>The study was done using two questionnaires six months apart, mixed-methods, with reflexive thematic analysis on the open-ended responses.</span></p><p><span>Vella found that productivity perceptions were stable and strongly positive over time. 84% of study participants reported improvement at both time points. Consistent with the first two studies in this round-up, the story of accelerated throughput is real and persistent.</span></p><p><span>The really striking finding is what the authors call the productivity-experience paradox. Among the matched cohort, the share of engineers reporting worse DevEx on at least one dimension nearly doubled in just six months, from 14% to 27%. Flow state was the most vulnerable; cognitive load eroded modestly; feedback loops actually improved.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tdUU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43b1042a-4a2c-405f-9a4d-cb3fde714fb5_2048x1065.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tdUU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43b1042a-4a2c-405f-9a4d-cb3fde714fb5_2048x1065.png 424w, https://substackcdn.com/image/fetch/$s_!tdUU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43b1042a-4a2c-405f-9a4d-cb3fde714fb5_2048x1065.png 848w, https://substackcdn.com/image/fetch/$s_!tdUU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43b1042a-4a2c-405f-9a4d-cb3fde714fb5_2048x1065.png 1272w, https://substackcdn.com/image/fetch/$s_!tdUU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43b1042a-4a2c-405f-9a4d-cb3fde714fb5_2048x1065.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tdUU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43b1042a-4a2c-405f-9a4d-cb3fde714fb5_2048x1065.png" width="1456" height="757" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/43b1042a-4a2c-405f-9a4d-cb3fde714fb5_2048x1065.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:757,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tdUU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43b1042a-4a2c-405f-9a4d-cb3fde714fb5_2048x1065.png 424w, https://substackcdn.com/image/fetch/$s_!tdUU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43b1042a-4a2c-405f-9a4d-cb3fde714fb5_2048x1065.png 848w, https://substackcdn.com/image/fetch/$s_!tdUU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43b1042a-4a2c-405f-9a4d-cb3fde714fb5_2048x1065.png 1272w, https://substackcdn.com/image/fetch/$s_!tdUU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43b1042a-4a2c-405f-9a4d-cb3fde714fb5_2048x1065.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>More importantly: while the cross-sectional correlations between DevEx and productivity were strong, the change scores didn&#8217;t correlate. Productivity and developer experience appear to be decoupling over time in AI-assisted workflows. For those of us who&#8217;ve spent years working with the SPACE and DevEx frameworks, that&#8217;s worth sitting with.</span></p><p><span>While this study didn&#8217;t have a particularly large population, the findings were significant and rigorously validated, proving that a study doesn&#8217;t have to be massive if the strength of results is strong enough. This longitudinal design is rare and valuable, and the productivity-experience decoupling is the kind of finding worth replicating in larger populations.</span></p><h3><span>4. To Copilot and Beyond: 22 AI Systems Developers Want Built</span></h3><p><em><span>Paper: Choudhuri, R., Badea, C., Bird, C., Butler, J., DeLine, R., Houck, B. </span><a href="https://arxiv.org/abs/2510.00762"><span>AI Where It Matters: Where, Why, and How Developers Want AI Support in Daily Work.</span></a></em></p><p><span>Last year I published a paper called </span><a href="https://arxiv.org/abs/2510.00762"><span>AI Where It Matters</span></a><span>, and my co-authors ended up writing a 2nd paper based on the original survey responses (860 Microsoft developers across roles, domains, and geographies). The paper outlines a roadmap of 22 AI tools that developers want beyond just code generation, centered around a concept they call &#8220;bounded delegation.&#8221; A lot of this echoes what the rest of this round-up is circling:</span></p><ul><li><p><span>The &#8220;right-shift&#8221; burden. Because AI is speeding up code generation, it&#8217;s creating a massive bottleneck downstream. Devs are getting flooded with more code to review, more production incidents to debug, and documentation that falls behind faster than ever.</span></p></li><li><p><span>The move to verification. Developers don&#8217;t want more code-generation assistants; they want AI embedded into verification tasks &#8212; tools that automatically assemble log/trace &#8220;case files&#8221; for on-call incidents, PR reviewers that catch complex business logic flaws before human review, change-aware test generation that knows which assertions actually matter.</span></p></li><li><p><span>&#8220;Bounded delegation.&#8221; There is a strict boundary around where developers want AI to stop. Developers want AI to absorb the tedious &#8220;assembly work&#8221; surrounding their craft (updating docs, writing edge-case unit tests), but never the core logic, architecture, or critical decision-making. Notably, developers drew this line even for tasks they acknowledged AI could plausibly handle &#8212; suggesting it&#8217;s not just about capability gaps and won&#8217;t move just because models improve.</span></p></li><li><p><span>Four non-negotiable guardrails. For future AI tools to be adopted, developers say they must enforce explicit authority scoping (no auto-approvals), clear data provenance, explicit uncertainty signaling (the AI must admit when it doesn&#8217;t know something), and least-privilege security access.</span></p></li></ul><p><span>You can check out both papers and an interactive website here: </span><a href="http://aka.ms/ai-where-it-matters"><span>aka.ms/ai-where-it-matters</span></a></p><h3><span>5. From Technical Debt to Cognitive and Intent Debt</span></h3><p><em><span>Paper: Storey, M. </span><a href="https://queue.acm.org/detail.cfm?id=3807966"><span>From Technical Debt to Cognitive and Intent Debt: Rethinking software health in the age of AI</span></a></em></p><p><span>I&#8217;ve saved this for last because I think this is the </span><a href="https://queue.acm.org/detail.cfm?id=3807966"><span>most important paper I&#8217;ve read in a long time.</span></a><span> Margaret-Anne Storey makes a generational argument: the metaphor we&#8217;ve used for decades to think about software health&#8212;technical debt&#8212;is no longer sufficient. AI is reducing technical debt (through refactoring, test generation, automated review) while quietly accelerating the accumulation of two other forms of debt that matter more in this era.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7Aa3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b7a2f5d-cca4-4a2e-b3de-06fe4f27af36_2048x1069.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7Aa3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b7a2f5d-cca4-4a2e-b3de-06fe4f27af36_2048x1069.png 424w, https://substackcdn.com/image/fetch/$s_!7Aa3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b7a2f5d-cca4-4a2e-b3de-06fe4f27af36_2048x1069.png 848w, https://substackcdn.com/image/fetch/$s_!7Aa3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b7a2f5d-cca4-4a2e-b3de-06fe4f27af36_2048x1069.png 1272w, https://substackcdn.com/image/fetch/$s_!7Aa3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b7a2f5d-cca4-4a2e-b3de-06fe4f27af36_2048x1069.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7Aa3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b7a2f5d-cca4-4a2e-b3de-06fe4f27af36_2048x1069.png" width="1456" height="760" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5b7a2f5d-cca4-4a2e-b3de-06fe4f27af36_2048x1069.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:760,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!7Aa3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b7a2f5d-cca4-4a2e-b3de-06fe4f27af36_2048x1069.png 424w, https://substackcdn.com/image/fetch/$s_!7Aa3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b7a2f5d-cca4-4a2e-b3de-06fe4f27af36_2048x1069.png 848w, https://substackcdn.com/image/fetch/$s_!7Aa3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b7a2f5d-cca4-4a2e-b3de-06fe4f27af36_2048x1069.png 1272w, https://substackcdn.com/image/fetch/$s_!7Aa3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b7a2f5d-cca4-4a2e-b3de-06fe4f27af36_2048x1069.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><span>Technical debt</span></strong><span> lives in code. It accumulates when implementation decisions compromise future changeability. AI is genuinely helping here.</span></p><p><strong><span>Cognitive debt </span></strong><span>lives in people. It accumulates when a team&#8217;s shared understanding of a system erodes faster than it&#8217;s replenished. When AI generates the code, developers may accept it without building the same mental model they would have built by writing it themselves. Multiply that across a team and over time, and you get &#8220;an accumulation of not knowing.&#8221;</span></p><p><strong><span>Intent debt</span></strong><span> lives in artifacts. It accumulates when the goals, constraints, and rationale that guide a system&#8212;the things both humans and AI agents need to work safely&#8212;are unclear, unwritten, or forgotten. As more development is AI-assisted, intent debt becomes a first-order constraint on what AI can actually do for you.</span></p><p><span>The three debts interact and compound. Intent debt causes cognitive debt; cognitive debt causes technical debt; technical debt amplifies cognitive debt. Managing software system health requires attention to all three layers, not just the one easiest to measure.</span></p><p><span>The four practical implications Storey draws are worth reading in full, but the headline is: treat understanding as a deliverable. Just as working code is a product of software development, shared understanding should be treated as a first-class deliverable, not something that happens as a side effect of writing code.</span></p><h2><span>Final thoughts</span></h2><p><span>Read together, these five papers say something stronger than any of them say individually. AI is genuinely making code generation faster, and the per-engineer efficiency gains are real (Heilman). But those gains don&#8217;t survive the trip to shipped software at anywhere near the same magnitude (Demirer). The bottleneck has moved downstream  to review, integration, verification, and understanding. Developers feel it, they&#8217;re explicitly asking for tools to address those bottlenecks while refusing to delegate the parts of the job they consider craft (Choudhuri). The lived experience of working this way is more uneven than the productivity numbers suggest, with flow and cognitive load eroding even as throughput holds (Vella). And the deepest cost may be one we don&#8217;t yet measure: the slow erosion of shared understanding, which is what makes any system safe to change (Storey).</span></p><p><span>The bottleneck has moved. Our tools, metrics, and team designs haven&#8217;t moved with it yet. That&#8217;s where the next several years of work in our field are going to happen.</span></p><div><hr></div><p><span>That&#8217;s it for this week. And make sure to sign up for my </span><a href="https://getdx.com/webinar/ai-in-engineering-q2-2026-benchmarks-research-readout/?utm_source=newsletter"><span>upcoming live research readout</span></a><span> covering new findings on AI&#8217;s impact, where we&#8217;ll discuss data from both DX and the broader industry.</span></p><p><span>-Brian</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.getdx.com/p/five-studies-that-are-changing-how?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.getdx.com/p/five-studies-that-are-changing-how?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p>]]></content:encoded></item><item><title><![CDATA[From dashboards to decisions]]></title><description><![CDATA[What three real-world case studies taught us about productivity measurement that works.]]></description><link>https://newsletter.getdx.com/p/from-dashboards-to-decisions</link><guid isPermaLink="false">https://newsletter.getdx.com/p/from-dashboards-to-decisions</guid><dc:creator><![CDATA[Brian Houck]]></dc:creator><pubDate>Wed, 08 Jul 2026 10:03:57 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/71b865b4-6cb2-4173-ae09-2899db922b0a_2400x1148.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em><strong>Welcome to the latest issue of Engineering Enablement,</strong> a weekly newsletter sharing research and perspectives on developer productivity.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.getdx.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://newsletter.getdx.com/subscribe?"><span>Subscribe now</span></a></p><p>&#128467; <a href="https://getdx.com/webinar/ai-in-engineering-q2-2026-benchmarks-research-readout/?utm_source=newsletter">Join me on July 23</a> for a readout of the upcoming State of AI Impact in Engineering: Q2 Report. We&#8217;ll discuss new findings from DX&#8217;s data on AI tool usage, spend, and impact across 500+ organizations. Register <a href="https://getdx.com/webinar/ai-in-engineering-q2-2026-benchmarks-research-readout/?utm_source=newsletter">here.</a></p><div><hr></div><p><span>I&#8217;ve spent a lot of time over the years thinking about developer productivity metrics. The longer I do this work, the more convinced I become that the hardest part isn&#8217;t measuring software engineering. It&#8217;s improving it.</span></p><p><span>That&#8217;s where I think the best measurement systems distinguish themselves. They don&#8217;t just tell you whether things are getting better or worse. They guide you to interventions that actually improve how developers work. If your dashboard isn&#8217;t changing decisions, it&#8217;s not creating value.</span></p><p><span>That idea sits at the heart of </span><em><a href="https://queue.acm.org/detail.cfm?id=3819080"><span>EngThrive: Make It Fast and Easy to Do Great Work</span></a></em><span>, a paper I recently co-authored with Tim Bozarth, David Liu, and Dean Carignan. The paper describes the measurement and improvement system I helped build at Microsoft, but what has stuck with me most aren&#8217;t the dashboards or the framework. They&#8217;re the stories.</span></p><p><span>One team intentionally &#8220;gamed&#8221; a productivity metric and improved onboarding for an entire year. Another protected developers&#8217; focus time and discovered that the biggest gains came from work they weren&#8217;t even trying to improve. A third gave every developer two unexpected days off and found that the lost &#8220;productivity&#8221; disappeared within weeks while the wellbeing benefits lasted for months.</span></p><p><span>Three different organizations. Three different problems. Three different interventions. Yet they all taught the same lesson: the best measurement systems don&#8217;t just tell you what&#8217;s happening. They help you figure out what to do next.</span></p><p><span>Before we get to those stories, though, it&#8217;s worth remembering how easy it is to measure confidently and be wrong.</span></p><p><span>In the first two months of mandatory remote work at Microsoft in early 2020, pull requests per developer jumped more than 20%, and the company&#8217;s stock price rose more than 15%. By those measures, things looked great. During that same period, however, 78% of developers reported feeling burned out. Three signals from the same quarter, pointing in two different directions. Any one of them on its own would have told a confident but completely misleading story.</span></p><p><span>That is the trap EngThrive was built to avoid. Organizations often measure activities&#8212;pull requests, commits, tasks completed&#8212;and quietly treat them as proxies for outcomes like delivery speed, software quality, or developer effectiveness. The two are not the same.</span></p><p><span>EngThrive instead organizes measurement around outcome dimensions like Speed, Ease, and Quality, with Thriving serving as a guardrail: if an intervention makes developers faster but leaves them burned out, we don&#8217;t consider it a success. We use a handful of outcome-oriented North Star metrics supported by diagnostic metrics that help explain why those outcomes move.</span></p><p><span>With that frame in mind, here are three case studies that fundamentally changed how I think about measuring, and improving, developer productivity.</span></p><h3><span>Case one: Improving one outcome changed four</span></h3><p><span>&#8220;Too many meetings&#8221; is one of the most frequently cited workplace challenges among software engineers. So in late 2025, Microsoft&#8217;s CoreAI organization launched an initiative to protect developers&#8217; focus time. Rather than simply banning meetings, leaders set an explicit target: lift the bottom 20% of developers to at least 25 hours of focus time per week.</span></p><p><span>To achieve this goal, teams did a handful of sensible things. They removed low-quality meetings, they clustered meetings together to form larger uninterrupted blocks of time, and they explicitly blocked time on their calendars to do focus work.</span></p><p><span>Within eight weeks, the results showed up across multiple dimensions. Focus time increased by 2.1 hours per developer per week, roughly twice the improvement seen in the control group. Bad Developer Days, a composite measure of daily developer friction, fell by 25%. PR velocity increased by 13%, about four times the control group. Taken together, the productivity gains were roughly equivalent to adding the output of 350 developers.</span></p><p><span>What surprised me wasn&#8217;t that focus time improved. It was that improving focus time seemed to improve things we weren&#8217;t directly trying to change.</span></p><p><span>Only about half of the reduction in Bad Developer Days could be explained by the additional focus time itself. Teams appeared to be using their newly protected capacity to pay down technical debt and eliminate other sources of friction that had been generating bad days in the first place. Creating focus time didn&#8217;t just help developers concentrate. It gave teams the space to improve the system that had been interrupting them.</span></p><p><span>That&#8217;s exactly why I think outcome-oriented measurement matters. If we had only measured focus time, we would have concluded that developers gained two extra hours each week. Looking across multiple dimensions revealed something much more interesting: the intervention triggered improvements well beyond its original goal.</span></p><p><span>One team lead summarized the lesson better than I could:</span><em><span> </span></em></p><blockquote><p><em><span>Metrics led to questions, questions led to improvements, improvements reinforced the metric&#8217;s value.</span></em></p></blockquote><h3><span>Case two: When gaming the metric is the right answer</span></h3><p><span>One of the most common objections to productivity metrics is that people will game them. I worry about that too, but I increasingly think it&#8217;s also one of the best tests of whether a metric is well designed. The best metrics are ones where gaming them is indistinguishable from genuine improvement.</span></p><p><span>Time-to-First-PR is my favorite example.</span></p><p><span>One organization of roughly 4,000 developers decided to &#8220;game&#8221; the metric on purpose by assigning every new hire a trivial pull request on their first day. As a gaming exercise, it worked exactly as intended. Time-to-First-PR improved by 30%.</span></p><p><span>The surprise came later.</span></p><p><span>Those same developers went on to complete 23% more pull requests over their first year than the control group. Interviews explained why. The first pull request was never really about the code. It was about setting up the development environment, learning the team&#8217;s tools and review process, and becoming comfortable contributing. By forcing all of that to happen in the first week, the organization didn&#8217;t just improve a metric. It accelerated onboarding.</span></p><p><span>A separate AI-assisted onboarding tool, FirstMate, arrived at the same conclusion from a different direction. By automating environment setup and helping new hires navigate an unfamiliar codebase, it reduced Time-to-First-PR by 65%.</span></p><p><span>That&#8217;s the lesson I keep coming back to:</span></p><blockquote><p><em><span>A well-designed metric shouldn&#8217;t be easy to game. It should be difficult to improve without doing something genuinely valuable. When that happens, gaming the metric and improving the system become the same thing.</span></em></p></blockquote><h3><span>Case three: A cost that turned out to be free</span></h3><p><span>During the burnout crisis of 2020, one organization tried something that looked reckless on a Speed dashboard: it gave every developer two unexpected days off, called Health Days.</span></p><p><span>At first, the metric behaved exactly as you would expect. Pull request output dropped during those two days.</span></p><p><span>Then something surprising happened.</span></p><p><span>Within two weeks, the &#8220;lost&#8221; pull requests had all been made up. The apparent productivity cost disappeared. The burnout relief, however, lasted another 14 weeks.</span></p><p><span>To me, this is the clearest example of why we treat Thriving as a guardrail rather than just another metric. If we had only looked at Speed, Health Days would have appeared to be a costly intervention and might never have been attempted again. Looking across multiple dimensions told a completely different story. The intervention was effectively free from a Speed perspective while delivering a sustained improvement in developer wellbeing.</span></p><p><span>I&#8217;ll also acknowledge an important limitation. This wasn&#8217;t a controlled experiment; it was one organization&#8217;s experience. But it&#8217;s consistent with a pattern we&#8217;ve seen repeatedly: leaders often assume wellbeing and productivity exist in tension, when in practice the trade-off is frequently much smaller than expected (or doesn&#8217;t exist at all).</span></p><blockquote><p><em><span>The lesson isn&#8217;t that every organization should schedule Health Days. It&#8217;s that looking across outcome dimensions gives leaders permission to try interventions that a single productivity metric would immediately reject.</span></em></p></blockquote><h2><span>Why this matters for engineering leaders</span></h2><p><span>The common thread across all three stories isn&#8217;t focus time, onboarding, or Health Days. It&#8217;s that none of those interventions came from optimizing a single activity metric. They came from measuring outcomes, looking across dimensions, and using supporting metrics to understand </span><em><span>why</span></em><span> those outcomes changed.</span></p><p><span>That&#8217;s the distinction I hope readers take away from the EngThrive work. A good measurement system isn&#8217;t just a reporting system. It&#8217;s a learning system. It doesn&#8217;t simply tell leaders whether things are getting better or worse. It helps them discover interventions they wouldn&#8217;t have tried otherwise, understand why they worked, and build confidence in repeating them.</span></p><p><span>Activity metrics still have an important role to play, but not as the destination. They&#8217;re clues. They help explain </span><em><span>why</span></em><span> an outcome changed, not whether it mattered in the first place.</span></p><p><span>Ultimately, I think that&#8217;s the shift engineering organizations need to make. Stop asking, </span><em><span>&#8220;What should we measure?&#8221;</span></em><span> Start asking, </span><em><span>&#8220;What decisions are we trying to make, and what measurements would help us make them better?&#8221;</span></em><span> The metrics should serve the intervention, not become the intervention.</span></p><p><span>None of this is the work of one person, or even four. EngThrive is the product of a large team that has spent years building the platform, the research, the surveys, and the discipline behind these results. If these stories are useful, the credit belongs to them.</span></p><div><hr></div><p><span>This week&#8217;s featured DevProd job openings. See more </span><a href="https://getdx.com/resources/devex-jobs/">open roles here</a><span>.</span></p><ul><li><p><strong>Ashby</strong><span> is hiring an </span><a href="https://jobs.ashbyhq.com/Ashby/0f5dbf59-687b-4d88-88a7-73ee0a66b48d?utm_source=PRgMeEgv1Z">Staff Platform Engineer</a><span> | Remote</span></p></li><li><p><strong>Carta</strong> is hiring a <a href="https://www.linkedin.com/jobs/view/4404135082">Senior Software Engineer II, Developer Experience</a> | Santa Clara, CA; San Francisco, CA; New York, NY</p></li><li><p><strong>Figma</strong><span> is hiring a </span><a href="https://job-boards.greenhouse.io/figma/jobs/5790627004?gh_jid=5790627004&amp;gh_src=db0ijm3x4us">Staff Software Engineer, Developer Experience</a><span> | Remote; US</span></p></li><li><p><strong>GM</strong> is hiring a <a href="https://generalmotors.wd5.myworkdayjobs.com/Careers_GM/job/Austin-Technical-Center---Austin-Technical-Center/Principal-Software-Engineer---Developer-Experience_JR-202610217">Principal Software Engineer, Developer Experience</a>  | Austin, Texas</p></li><li><p><strong>Morgan Stanley </strong><span>is hiring an </span><a href="https://www.linkedin.com/jobs/view/4393043964/">AI Platform Engineer - Vice President</a><span> | New York</span></p></li><li><p><strong>Notion</strong> is hiring a <a href="https://jobs.ashbyhq.com/notion/49bdf081-6e20-4323-8c73-6d6b19544ff5">Software Engineer, Developer Experience</a> | Hybrid; Hyderabad, India</p></li><li><p><strong>Vercel </strong>is hiring a<strong> </strong><a href="https://vercel.com/careers/sr-engineering-manager-platform-5461002004">Sr. Engineering Manager, Platform</a> | New York City, San Francisco</p></li></ul><div><hr></div><p>That&#8217;s it for this week. Thanks for reading.</p><p>-Brian</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.getdx.com/p/from-dashboards-to-decisions?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.getdx.com/p/from-dashboards-to-decisions?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p>]]></content:encoded></item><item><title><![CDATA[AI and engineering productivity: Debating the headlines]]></title><description><![CDATA[Listen now | Leaders from Etsy, Twilio, GitHub, Google, and Microsoft debate how AI is changing engineering productivity, technical debt, developer roles, and the future of software teams.]]></description><link>https://newsletter.getdx.com/p/ai-and-engineering-productivity-debating</link><guid isPermaLink="false">https://newsletter.getdx.com/p/ai-and-engineering-productivity-debating</guid><dc:creator><![CDATA[Justin Reock]]></dc:creator><pubDate>Mon, 29 Jun 2026 14:07:18 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/203305239/a5a8cc5da73a044fc35877683971ba8c.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>Listen and watch now on <strong><a href="https://youtu.be/BcnqmcgScgM">YouTube</a>, <a href="https://podcasts.apple.com/us/podcast/engineering-enablement-by-abi-noda/id1619140476">Apple</a>, and <a href="https://open.spotify.com/show/3NxjyIsuxeDMQtisDqBy7D">Spotify</a></strong>.</p><p><span>In this closing panel from DX Annual, Rafe Colburn, Chief Product and Technology Officer at Etsy; Jesse Adametz, Senior Director of Engineering, Platform Engineering at Twilio; Eirini Kalliamvakou, Research Advisor at GitHub; Collin Green, Senior Staff UX Researcher at Google; and Brian Houck, Senior Principal Applied Scientist at Microsoft debate some of the biggest questions surrounding AI and engineering productivity.</span></p><p><span>They discuss whether AI will reduce the need for engineers, how AI is affecting technical debt, the future role of software engineers in an agentic world, and whether organizations should mandate AI adoption. They also explore how bottlenecks are shifting across the software development lifecycle, the challenges facing junior engineers, and why learning, culture, and change management may ultimately matter more than the tools themselves.</span></p><div id="youtube2-BcnqmcgScgM" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;BcnqmcgScgM&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/BcnqmcgScgM?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h2><strong>Some takeaways: </strong></h2><p><strong><span>AI is changing software engineering, but not eliminating the need for engineers</span></strong></p><ul><li><p><strong><span>The panel largely rejected the idea that an AI-first SDLC means dramatically fewer engineers.</span></strong><span> As the cost of building software decreases, demand for software is likely to increase, creating new opportunities rather than eliminating the need for technical talent.</span></p></li><li><p><strong><span>Several panelists argued that the role of software engineers will evolve rather than disappear.</span></strong><span> The tasks that make up the job may change, but organizations will continue to need people who can solve problems, make decisions, and build systems.</span></p></li></ul><p><strong><span>Technical debt remains a tradeoff, not just an AI problem</span></strong></p><ul><li><p><strong><span>Panelists disagreed on whether AI is creating technical debt faster than it can remove it.</span></strong><span> Some argued that AI is accelerating both code generation and technical debt, while others believed the underlying business pressures that create technical debt remain largely unchanged.</span></p></li><li><p><strong><span>The discussion also introduced the idea of cognitive debt.</span></strong><span> As engineers rely more heavily on AI-generated code, understanding and maintaining systems may become more difficult even if development velocity increases.</span></p></li></ul><p><strong><span>The future engineer may work at a higher level of abstraction</span></strong></p><ul><li><p><strong><span>Several panelists predicted that engineers will spend less time writing code directly and more time defining intent, setting constraints, providing context, and validating results.</span></strong><span> Rather than replacing engineering work, AI may shift it to a different level of abstraction.</span></p></li><li><p><strong><span>The panel also pushed back on the idea that engineers will simply become managers of agents.</span></strong><span> Effective AI use still requires technical judgment, communication skills, and careful oversight.</span></p></li></ul><p><strong><span>Mandates rarely create meaningful AI adoption</span></strong></p><ul><li><p><strong><span>Most panelists opposed the idea that organizations should mandate AI usage.</span></strong><span> Instead, they emphasized enablement, reducing friction, and helping developers discover value through their own workflows.</span></p></li><li><p><strong><span>Usage metrics can easily become the wrong goal.</span></strong><span> The group cautioned against treating AI usage itself as a performance metric, arguing that outcomes matter more than activity.</span></p></li></ul><p><strong><span>Junior engineers remain essential to the future of the profession</span></strong></p><ul><li><p><strong><span>The panel strongly rejected the idea that organizations will no longer need junior engineers.</span></strong><span> Today&#8217;s junior engineers become tomorrow&#8217;s senior engineers, making talent development critical to the long-term health of the industry.</span></p></li><li><p><strong><span>Several speakers also noted that newer engineers may bring valuable AI-native perspectives.</span></strong><span> Just as previous technology shifts rewarded developers who grew up with new tools, the next generation may help shape how AI is used in practice.</span></p></li></ul><p><strong><span>The biggest AI adoption challenges are human, not technical</span></strong></p><ul><li><p><strong><span>While tooling matters, the panel repeatedly returned to learning, culture, incentives, and change management as the biggest barriers to successful AI adoption.</span></strong><span> Engineers are navigating rapid technological change, shifting workflows, and new expectations about their role.</span></p></li><li><p><strong><span>Organizations that create space for learning appear to see stronger results.</span></strong><span> The panel highlighted examples where teams learned together, experimented together, and achieved better adoption outcomes than individuals working in isolation.</span></p></li></ul><h2><strong>In this episode, we cover:</strong></h2><p>(<a href="https://www.youtube.com/watch?v=BcnqmcgScgM">00:00</a>) Intro</p><p>(<a href="https://www.youtube.com/watch?v=BcnqmcgScgM&amp;t=76s">01:16</a>) Why an AI-first SDLC doesn&#8217;t mean fewer engineers</p><p>(<a href="https://www.youtube.com/watch?v=BcnqmcgScgM&amp;t=189s">03:09</a>) The debate over AI and technical debt</p><p>(<a href="https://www.youtube.com/watch?v=BcnqmcgScgM&amp;t=460s">07:40</a>) AI-generated code and the future role of engineers</p><p>(<a href="https://www.youtube.com/watch?v=BcnqmcgScgM&amp;t=856s">14:16</a>) Why mandating AI use doesn&#8217;t necessarily lead to better outcomes</p><p>(<a href="https://www.youtube.com/watch?v=BcnqmcgScgM&amp;t=1243s">20:43</a>) Predictions for the future of junior engineers</p><p>(<a href="https://www.youtube.com/watch?v=BcnqmcgScgM&amp;t=1402s">23:22</a>) Where the bottlenecks are in the SDLC now</p><p>(<a href="https://www.youtube.com/watch?v=BcnqmcgScgM&amp;t=1705s">28:25</a>) How risk influences AI use</p><p>(<a href="https://www.youtube.com/watch?v=BcnqmcgScgM&amp;t=1958s">32:38</a>) Why the human side is the biggest AI adoption challenge</p><h2><strong>Referenced:</strong></h2><p><span>&#8226; </span><a href="https://www.etsy.com/"><span>Etsy</span></a></p><p><span>&#8226; </span><a href="https://github.com/"><span>GitHub</span></a></p><p><span>&#8226; </span><a href="https://www.microsoft.com/en-us"><span>Microsoft</span></a></p><p><span>&#8226; </span><a href="https://www.twilio.com/"><span>Twilio</span></a></p><p><span>&#8226; </span><a href="https://www.google.com/"><span>Google</span></a></p><p><span>&#8226; </span><a href="http://linkedin.com/in/stewartreichling"><span>Stewart Reichling</span></a></p><p><span>&#8226; </span><a href="https://getdx.com/blog/space-metrics/"><span>What is the SPACE framework and when should you use it?</span></a></p>]]></content:encoded></item><item><title><![CDATA[2x the power users: How structured AI training scaled developer productivity]]></title><description><![CDATA[How Indeed drove AI coding tool adoption from 25% to 97% across 2,000 engineers, and what it learned about training, enablement, and preparing for the next phase of AI-assisted development.]]></description><link>https://newsletter.getdx.com/p/2x-the-power-users-how-structured</link><guid isPermaLink="false">https://newsletter.getdx.com/p/2x-the-power-users-how-structured</guid><pubDate>Mon, 29 Jun 2026 14:04:35 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/203305672/3f998975f9d9868650338b9a7537c14f.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>Listen and watch now on <strong><a href="https://youtu.be/iomiGESWxMg">YouTube</a>, <a href="https://podcasts.apple.com/us/podcast/engineering-enablement-by-abi-noda/id1619140476">Apple</a>, and <a href="https://open.spotify.com/show/3NxjyIsuxeDMQtisDqBy7D">Spotify</a></strong>.</p><p><span>Indeed increased AI coding tool adoption from roughly 25% to 97% across its engineering organization, but getting engineers to use the tools was only part of the challenge.</span></p><p><span>In this session from DX Annual, Michael Redding, Principal Product Manager, and Jeff Davis, VP of Core Infrastructure at Indeed, explain how the company used structured training, leadership support, and ongoing community engagement to help more than 2,000 engineers build practical AI skills. They share why an early train-the-trainer model fell short, how they redesigned their approach around hands-on learning, and what they learned about balancing adoption, measurement, and psychological safety.</span></p><p><span>They also discuss the impact of the program on coding time, the role of continuous enablement after formal training ended, and how Indeed is preparing for the next phase of AI adoption, including agentic workflows and AI-powered coaching.</span></p><div id="youtube2-iomiGESWxMg" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;iomiGESWxMg&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/iomiGESWxMg?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h2><strong>Some takeaways: </strong></h2><p><strong><span>Indeed started with a productivity problem, not an AI problem</span></strong></p><ul><li><p><strong><span>At the beginning of 2025, Indeed&#8217;s DX survey showed that only about half of developer time was being spent on new features and innovation.</span></strong><span> The remaining 48% was consumed by maintenance, upgrades, incident response, and other forms of engineering overhead.</span></p></li><li><p><strong><span>The company&#8217;s AI strategy focused on two goals: reducing overhead work and increasing output during coding time.</span></strong><span> The long-term objective was to double engineering productivity by shrinking non-value-added work while helping engineers produce more during the time they spend building.</span></p></li></ul><p><strong><span>AI Coding Essentials succeeded where AI Coding Ambassadors fell short</span></strong></p><ul><li><p><strong><span>Indeed&#8217;s first enablement effort, AI Coding Ambassadors, used a train-the-trainer model built around roughly 60 AI champions across the organization.</span></strong><span> While ambassadors maintained high levels of engagement, adoption among their teammates declined after the program ended.</span></p></li><li><p><strong><span>The company responded by launching AI Coding Essentials (AICE), a structured training program designed for all engineers.</span></strong><span> The experience convinced the team that direct, hands-on learning was far more effective than relying on knowledge to spread organically through teams.</span></p></li></ul><p><strong><span>Indeed treated AI upskilling as a company-wide investment</span></strong></p><ul><li><p><strong><span>Training more than 2,000 engineers required significant organizational commitment and leadership support.</span></strong><span> Michael estimated the investment at roughly $3&#8211;4 million in engineering time across the company.</span></p></li><li><p><strong><span>Rather than mandating AI usage, Indeed strongly encouraged completion of the training itself.</span></strong><span> Managers were given visibility into participation, while engineers retained flexibility in how and whether they ultimately incorporated AI into their workflows.</span></p></li></ul><p><strong><span>AI adoption increased from 25% to 97%</span></strong></p><ul><li><p><strong><span>Despite offering AI tools, training resources, and executive support, weekly AI usage remained stuck around 25% at the start of 2025.</span></strong><span> The challenge was not tool access but helping engineers develop practical skills and confidence.</span></p></li><li><p><strong><span>By the time of the presentation, weekly AI tool usage had reached approximately 97%.</span></strong><span> The company also successfully navigated multiple tool transitions, moving from Cody and Copilot to newer agentic tools such as Claude Code, Cursor, Windsurf, and Amp.</span></p></li></ul><p><strong><span>Structured training produced measurable results</span></strong></p><ul><li><p><strong><span>Engineers who completed AI Coding Essentials reduced coding time by roughly 35&#8211;36%, while engineers who did not complete the training saw little change.</span></strong><span> Across the broader organization, coding time decreased by roughly 20%.</span></p></li><li><p><strong><span>Indeed measured coding time as the period between a developer picking up a Jira ticket and opening a diff in GitLab.</span></strong><span> The company continued to see benefits months after training ended, especially as newer frontier models became available.</span></p></li></ul><p><strong><span>Community and continuous enablement kept momentum going</span></strong></p><ul><li><p><strong><span>Indeed reinforced learning through coding forums, office hours, hackathons, Slack communities, and its AI Showcase recognition program.</span></strong><span> More than 100 unique community posts were being shared monthly in the company&#8217;s primary AI channel.</span></p></li><li><p><strong><span>The goal was to make AI learning continuous rather than event-based.</span></strong><span> Engineers had multiple ways to share discoveries, get help, and learn from peers long after formal training concluded.</span></p></li></ul><p><strong><span>The next challenge is moving beyond coding</span></strong></p><ul><li><p><strong><span>Indeed is now focused on agentic workflows, AI coaching, and expanding enablement beyond software engineering.</span></strong><span> Product managers, designers, researchers, and other R&amp;D functions are becoming part of the company&#8217;s AI adoption strategy.</span></p></li><li><p><strong><span>As coding becomes faster, bottlenecks are beginning to shift elsewhere in the development lifecycle.</span></strong><span> The team is already monitoring signs that code review and other downstream activities may become the next constraints on engineering throughput.</span></p></li></ul><h2><strong>In this episode, we cover:</strong></h2><p>(<a href="https://www.youtube.com/watch?v=iomiGESWxMg">00:00</a>) Intro</p><p>(<a href="https://www.youtube.com/watch?v=iomiGESWxMg&amp;t=65s">01:05</a>) Indeed&#8217;s DX survey from January 2025</p><p>(<a href="https://www.youtube.com/watch?v=iomiGESWxMg&amp;t=150s">02:30</a>) The two-part strategy to double engineering productivity</p><p>(<a href="https://www.youtube.com/watch?v=iomiGESWxMg&amp;t=261s">04:21</a>) How Indeed increased AI adoption from 25% to 97%</p><p>(<a href="https://www.youtube.com/watch?v=iomiGESWxMg&amp;t=940s">15:40</a>) Results from Indeed&#8217;s AI training program</p><p>(<a href="https://www.youtube.com/watch?v=iomiGESWxMg&amp;t=1113s">18:33</a>) How Indeed sustains AI adoption and learning</p><p>(<a href="https://www.youtube.com/watch?v=iomiGESWxMg&amp;t=1386s">23:06</a>) What&#8217;s next for AI enablement at Indeed</p><p>(<a href="https://www.youtube.com/watch?v=iomiGESWxMg&amp;t=1481s">24:41</a>) Q&amp;A: How coding time was calculated</p><p>(<a href="https://www.youtube.com/watch?v=iomiGESWxMg&amp;t=1525s">25:25</a>) Q&amp;A: How Indeed uses AI playbooks</p><p>(<a href="https://www.youtube.com/watch?v=iomiGESWxMg&amp;t=1600s">26:40</a>) Q&amp;A: Balancing asynchronous and live AI training</p><p>(<a href="https://www.youtube.com/watch?v=iomiGESWxMg&amp;t=1702s">28:22</a>) Q&amp;A: Psychological safety during AI adoption</p><p>(<a href="https://www.youtube.com/watch?v=iomiGESWxMg&amp;t=1904s">31:44</a>) Q&amp;A: Why AI adoption spikes after the holidays</p><p>(<a href="https://www.youtube.com/watch?v=iomiGESWxMg&amp;t=2000s">33:20</a>) Q&amp;A: The metrics Indeed tracked</p><p>(<a href="https://www.youtube.com/watch?v=iomiGESWxMg&amp;t=2122s">35:22</a>) Q&amp;A: Where the time savings are going</p><p>(<a href="https://www.youtube.com/watch?v=iomiGESWxMg&amp;t=2214s">36:54</a>) Q&amp;A: Reaching engineers who skipped the training</p><p>(<a href="https://www.youtube.com/watch?v=iomiGESWxMg&amp;t=2288s">38:08</a>) Closing thoughts</p><h2><strong>Referenced:</strong></h2><p><span>&#8226; </span><a href="https://www.indeed.com/"><span>Indeed</span></a></p><p><span>&#8226; </span><a href="https://www.anthropic.com/product/claude-code"><span>Claude Code | Anthropic&#8217;s agentic coding system</span></a></p><p><span>&#8226; </span><a href="https://cursor.com/"><span>Cursor</span></a></p><p><span>&#8226; </span><a href="https://www.windsurf.dev/"><span>Windsurf</span></a></p><p><span>&#8226; </span><a href="https://ampcode.com/"><span>Amp Code</span></a></p><p><span>&#8226; </span><a href="https://resources.anthropic.com/hubfs/The-Complete-Guide-to-Building-Skill-for-Claude.pdf"><span>The Complete Guide to Building Skills for Claude | Anthropic</span></a></p><p><span>&#8226; </span><a href="https://getdx.com/report/dx-core-4/"><span>Measuring developer productivity with the DX Core 4</span></a></p>]]></content:encoded></item><item><title><![CDATA[From PR throughput to product velocity: How Dropbox is rethinking productivity in the agentic era]]></title><description><![CDATA[How Dropbox is adapting its engineering systems, workflows, and metrics for the agentic era as AI shifts bottlenecks beyond code generation.]]></description><link>https://newsletter.getdx.com/p/from-pr-throughput-to-product-velocity</link><guid isPermaLink="false">https://newsletter.getdx.com/p/from-pr-throughput-to-product-velocity</guid><pubDate>Mon, 29 Jun 2026 13:59:38 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/203305440/dc35434f3b2719fdc32ad6787e8d8f75.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>Listen and watch now on <strong><a href="https://youtu.be/w0kHCjTOvyo">YouTube</a>, <a href="https://podcasts.apple.com/us/podcast/engineering-enablement-by-abi-noda/id1619140476">Apple</a>, and <a href="https://open.spotify.com/show/3NxjyIsuxeDMQtisDqBy7D">Spotify</a></strong>.</p><p><span>In this session from DX Annual, Uma Namasivayam, Senior Director of Engineering Productivity at Dropbox, shares how the company&#8217;s developer productivity efforts evolved from improving developer experience to preparing for the agentic era.</span></p><p><span>He explains how Dropbox approached AI adoption across its engineering organization, the impact it had on developer productivity, and why faster code generation is creating new bottlenecks in areas such as code review, validation, and CI/CD. He also discusses Dropbox&#8217;s efforts to rethink engineering systems, measurement, and workflows, including the development of agentic tooling and new metrics designed to move beyond PR throughput and toward product velocity.</span></p><div id="youtube2-w0kHCjTOvyo" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;w0kHCjTOvyo&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/w0kHCjTOvyo?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h2><strong>Some takeaways: </strong></h2><p><strong><span>Dropbox&#8217;s productivity journey started before AI</span></strong></p><ul><li><p><strong><span>DXI helped Dropbox identify productivity problems as system problems rather than talent problems.</span></strong><span> When the company began measuring developer experience in 2023, it found significant variation across teams in DXI scores, PR throughput, and cycle time.</span></p></li><li><p><strong><span>Measuring developer experience created a framework for prioritizing investments.</span></strong><span> The team used DXI to identify friction across areas such as debugging, documentation, and build systems while giving leadership a common language for discussing productivity.</span></p></li></ul><p><strong><span>AI adoption required more than access to tools</span></strong></p><ul><li><p><strong><span>Dropbox combined executive support, developer segmentation, enablement, and strong guardrails to drive adoption.</span></strong><span> Different teams and developer roles were matched with different tools and workflows based on their needs.</span></p></li><li><p><strong><span>The approach helped Dropbox increase AI adoption from roughly 30% to 100% within six months.</span></strong><span> During the same period, PR throughput doubled and developer satisfaction with AI tools increased significantly.</span></p></li></ul><p><strong><span>Engineers used their extra capacity to tackle neglected work</span></strong></p><ul><li><p><strong><span>As AI increased throughput, engineers naturally pulled maintenance work, migrations, and technical debt from the backlog.</span></strong><span> Dropbox saw significant growth in these categories without any specific direction from leadership.</span></p></li><li><p><strong><span>The additional capacity was often reinvested into engineering health.</span></strong><span> Teams used the opportunity to address long-standing issues that had accumulated over time rather than focusing exclusively on new feature development.</span></p></li></ul><p><strong><span>The next challenges are scale, trust, and measurement</span></strong></p><ul><li><p><strong><span>Dropbox believes the move to agentic engineering creates three major challenges: scale, validation and trust, and measurement.</span></strong><span> Existing development systems were not designed for a world where AI dramatically increases code throughput.</span></p></li><li><p><strong><span>As code generation accelerates, bottlenecks are shifting toward code review, validation, and CI/CD systems.</span></strong><span> The company is already seeing pressure move downstream in the software development lifecycle.</span></p></li></ul><p><strong><span>Agentic engineering requires redesigning the entire system</span></strong></p><ul><li><p><strong><span>Uma compared the transition to the shift from steam-powered factories to electric factories.</span></strong><span> The biggest gains came from redesigning the entire system rather than simply replacing one technology with another.</span></p></li><li><p><strong><span>Dropbox is investing in agentic workflows across the SDLC and building Nova as an orchestration layer.</span></strong><span> The company is evaluating roughly 30 development steps, and one in twelve pull requests is already being generated by Nova.</span></p></li></ul><p><strong><span>PR throughput is becoming a less useful measure of productivity</span></strong></p><ul><li><p><strong><span>Dropbox believes traditional engineering metrics need to evolve alongside AI.</span></strong><span> As agentic workflows become more common, measuring productivity through pull request volume alone provides an incomplete picture of engineering output.</span></p></li><li><p><strong><span>The company is increasingly focused on metrics such as AI contribution, loaded cost per PR, agentic workflow coverage, work distribution, and time to ship.</span></strong><span> The goal is to better connect engineering activity to customer value and business outcomes.</span></p></li></ul><h2><strong>In this episode, we cover:</strong></h2><p>(<a href="https://www.youtube.com/watch?v=w0kHCjTOvyo">00:00</a>) Intro</p><p>(<a href="https://www.youtube.com/watch?v=w0kHCjTOvyo&amp;t=57s">00:57</a>) The beginning of Dropbox&#8217;s DX journey</p><p>(<a href="https://www.youtube.com/watch?v=w0kHCjTOvyo&amp;t=154s">02:34</a>) AI adoption at Dropbox: what made it work</p><p>(<a href="https://www.youtube.com/watch?v=w0kHCjTOvyo&amp;t=286s">04:46</a>) The results of Dropbox&#8217;s AI adoption efforts</p><p>(<a href="https://www.youtube.com/watch?v=w0kHCjTOvyo&amp;t=339s">05:39</a>) What the results mean for the business</p><p>(<a href="https://www.youtube.com/watch?v=w0kHCjTOvyo&amp;t=415s">06:55</a>) The phases of AI adoption and where they are now</p><p>(<a href="https://www.youtube.com/watch?v=w0kHCjTOvyo&amp;t=480s">08:00</a>) The new bottlenecks</p><p>(<a href="https://www.youtube.com/watch?v=w0kHCjTOvyo&amp;t=556s">09:16</a>) Three challenges Dropbox faces moving into agentic engineering</p><p>(<a href="https://www.youtube.com/watch?v=w0kHCjTOvyo&amp;t=605s">10:05</a>) How Dropbox is redesigning the SDLC for agentic engineering</p><p>(<a href="https://www.youtube.com/watch?v=w0kHCjTOvyo&amp;t=946s">15:46</a>) The new metrics that matter</p><p>(<a href="https://www.youtube.com/watch?v=w0kHCjTOvyo&amp;t=1156s">19:16</a>) Final takeaways</p><h2><strong>Referenced:</strong></h2><p><span>&#8226; </span><a href="https://www.dropbox.com/"><span>Dropbox</span></a></p><p><span>&#8226; </span><a href="https://getdx.com/developer-experience-index/"><span>Developer Experience Index (DXI) | DX</span></a></p><p><span>&#8226; </span><a href="https://getdx.com/corefour"><span>DX Core 4 Productivity Framework</span></a></p><p><span>&#8226; </span><a href="https://cursor.com/"><span>Cursor</span></a></p><p><span>&#8226; </span><a href="https://www.anthropic.com/product/claude-code"><span>Claude Code | Anthropic&#8217;s agentic coding system</span></a></p><p><span>&#8226; </span><a href="https://www.jetbrains.com/"><span>JetBrains</span></a></p><p><span>&#8226; </span><a href="https://code.visualstudio.com/"><span>Visual Studio Code</span></a></p><p><span>&#8226; </span><a href="https://www.atlassian.com/software/jira"><span>Jira | Project Management for the AI Era | Atlassian</span></a></p><p><span>&#8226; </span><a href="https://github.com/"><span>GitHub</span></a></p>]]></content:encoded></item><item><title><![CDATA[Revisiting the DX Core 4 in the age of AI]]></title><description><![CDATA[Why the dimensions that matter most for engineering productivity remain stable, and how to interpret them as AI reshapes work.]]></description><link>https://newsletter.getdx.com/p/revisiting-the-dx-core-4</link><guid isPermaLink="false">https://newsletter.getdx.com/p/revisiting-the-dx-core-4</guid><dc:creator><![CDATA[Brian Houck]]></dc:creator><pubDate>Wed, 24 Jun 2026 10:00:51 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/a844724d-d2cc-4c90-9033-b5139cd0a03e_2400x1260.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em><strong>Welcome to the latest issue of Engineering Enablement,</strong><span> a weekly newsletter sharing research and perspectives on developer productivity.</span></em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.getdx.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.getdx.com/subscribe?"><span>Subscribe now</span></a></p><p><span>&#128467; </span><a href="https://getdx.com/webinar/ai-in-engineering-q2-2026-benchmarks-research-readout/?utm_source=newsletter"><span>Join me on July 23</span></a><span> for a readout of the upcoming Q2 2026 AI Impact Report. We&#8217;ll discuss new findings from DX&#8217;s data on AI tool usage, spend, and impact across 500+ organizations. Register </span><a href="https://getdx.com/webinar/ai-in-engineering-q2-2026-benchmarks-research-readout/?utm_source=newsletter"><span>here.</span></a></p><div><hr></div><p><span>When AI coding tools started delivering meaningful results, a predictable question followed from CTOs and engineering leaders: how do we measure the impact? There is a strong instinct to assume that the frameworks built over the last decade no longer apply, and that the age of AI demands a fundamentally different measurement architecture.</span></p><p><span>I&#8217;d push back on that instinct. The evidence suggests the opposite is closer to the truth.</span></p><p><span>While AI represents a massive paradigm shift in how software is built, it does not alter what engineering organizations are ultimately trying to accomplish. Foundational engineering principles still map to high-level outcomes. How quickly is value delivered? How easy is it for developers to do their work effectively? How stable are the systems? And, ultimately, what is the business impact of the work? Rather than rendering these categories obsolete, the introduction of AI makes anchoring to a stable, outcome-oriented framework more critical than ever.</span></p><p><span>Engineering leaders are under unprecedented pressure to justify the massive budgets being poured into AI tooling. When executives demand proof that an AI investment is paying off, the immediate temptation is to reach for a shiny new metric that isolates the tool itself. But that is exactly where the risk lies.</span></p><p><span>The </span><a href="https://getdx.com/research/measuring-developer-productivity-with-the-dx-core-4/"><span>DX Core 4</span></a><span> framework (speed, effectiveness, quality, and business impact) is built around answering these persistent questions. It was designed to give engineering leaders a durable measurement architecture that survives new technology cycles. AI is a significant shift in workflow, but because the framework anchors to macro outcomes rather than the mechanics of coding, it remains stable. If anything, the rise of AI makes this type of durable framework more important, not less.</span></p><p><span>This article makes three related arguments:</span></p><ol><li><p><span>First, the high-level dimensions of engineering productivity remain remarkably stable, even as AI transforms how software is built.</span></p></li><li><p><span>Second, AI-specific telemetry should be treated as diagnostic context rather than a replacement for outcome-oriented measurement.</span></p></li><li><p><span>Finally, while many traditional engineering metrics remain valuable, the behaviors that generate them are changing, and disentangling those signals requires triangulating across the layered structure of diagnostic, system, and outcome metrics.</span></p></li></ol><h2><span>The Core 4 holds (and here&#8217;s why that matters)</span></h2><p><span>The value of anchoring to these four overarching dimensions&#8212;speed, effectiveness, quality, and business impact&#8212;is that they synthesize key principles from </span><a href="https://dora.dev/capabilities/"><span>DORA</span></a><span>, </span><a href="https://queue.acm.org/detail.cfm?id=3454124"><span>SPACE</span></a><span>, and </span><a href="https://queue.acm.org/detail.cfm?id=3595878"><span>DevEx</span></a><span> into a unified methodology. Core 4 inherits DORA&#8217;s focus on delivery outcomes, SPACE&#8217;s insistence that productivity is multidimensional, and DevEx&#8217;s emphasis on the lived experience of developers&#8212;and combines them into a four-dimension framework optimized for executive decision-making.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!UNOC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cf31595-88c2-4428-9c52-76eac609dd09_8763x3629.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!UNOC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cf31595-88c2-4428-9c52-76eac609dd09_8763x3629.jpeg 424w, https://substackcdn.com/image/fetch/$s_!UNOC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cf31595-88c2-4428-9c52-76eac609dd09_8763x3629.jpeg 848w, https://substackcdn.com/image/fetch/$s_!UNOC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cf31595-88c2-4428-9c52-76eac609dd09_8763x3629.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!UNOC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cf31595-88c2-4428-9c52-76eac609dd09_8763x3629.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!UNOC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cf31595-88c2-4428-9c52-76eac609dd09_8763x3629.jpeg" width="1456" height="603" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5cf31595-88c2-4428-9c52-76eac609dd09_8763x3629.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:603,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2591128,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.getdx.com/i/203146039?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cf31595-88c2-4428-9c52-76eac609dd09_8763x3629.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!UNOC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cf31595-88c2-4428-9c52-76eac609dd09_8763x3629.jpeg 424w, https://substackcdn.com/image/fetch/$s_!UNOC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cf31595-88c2-4428-9c52-76eac609dd09_8763x3629.jpeg 848w, https://substackcdn.com/image/fetch/$s_!UNOC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cf31595-88c2-4428-9c52-76eac609dd09_8763x3629.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!UNOC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cf31595-88c2-4428-9c52-76eac609dd09_8763x3629.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>AI doesn&#8217;t change what engineering organizations are trying to accomplish. What it does is make the signals noisier.</span></p><p><span>As AI coding assistants become standard and agentic workflows begin handling multi-step tasks autonomously, traditional activity metrics shift in ways that can easily mislead. Pull request counts spike, cycle times compress, and code volumes bloat. Engineering leaders who chase these surface-level fluctuations without anchoring to a balanced, outcome-oriented framework risk optimizing for sheer motion rather than actual progress.</span></p><p><span>This is precisely where a high-level outcome framework proves its utility. I&#8217;m using the Core 4 as the specific example here, but the same logic applies to any mature measurement framework aligned to the principles of </span><a href="https://queue.acm.org/detail.cfm?id=3454124"><span>SPACE</span></a><span>. By focusing on outcomes that matter, regardless of how code gets written, the model remains insulated from technology disruptions. This structural design looks increasingly necessary as developer workflows continue to evolve away from manual synthesis and toward intent-driven architecture.</span></p><h3><span>Activity vs. outcome: The role of AI telemetry</span></h3><p><span>To be clear, focusing on measuring stable macro outcomes does not mean engineering leaders should ignore AI adoption and usage. Tracking how developers engage with AI tools is incredibly valuable, but it is critical to understand </span><em><span>what</span></em><span> those metrics are telling us.</span></p><p><span>AI adoption, token usage, and the number of tasks assigned to agents are examples of diagnostic telemetry. Like more traditional operational metrics such as pull request size, build duration, or meeting load, they provide visibility into how work is being performed rather than whether it is producing better outcomes.</span></p><p><span>One way to think about this distinction is illustrated in the image below, whether AI-specific or traditional, helps explain the mechanics of software delivery and the dynamics of the engineering system. By contrast, outcome-oriented frameworks evaluate whether those operating patterns are ultimately translating into better engineering results.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pkQA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76804bca-cf2f-41ea-8e36-c8b1e12f6d5e_2048x1209.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pkQA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76804bca-cf2f-41ea-8e36-c8b1e12f6d5e_2048x1209.png 424w, https://substackcdn.com/image/fetch/$s_!pkQA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76804bca-cf2f-41ea-8e36-c8b1e12f6d5e_2048x1209.png 848w, https://substackcdn.com/image/fetch/$s_!pkQA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76804bca-cf2f-41ea-8e36-c8b1e12f6d5e_2048x1209.png 1272w, https://substackcdn.com/image/fetch/$s_!pkQA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76804bca-cf2f-41ea-8e36-c8b1e12f6d5e_2048x1209.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pkQA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76804bca-cf2f-41ea-8e36-c8b1e12f6d5e_2048x1209.png" width="1456" height="860" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/76804bca-cf2f-41ea-8e36-c8b1e12f6d5e_2048x1209.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:860,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!pkQA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76804bca-cf2f-41ea-8e36-c8b1e12f6d5e_2048x1209.png 424w, https://substackcdn.com/image/fetch/$s_!pkQA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76804bca-cf2f-41ea-8e36-c8b1e12f6d5e_2048x1209.png 848w, https://substackcdn.com/image/fetch/$s_!pkQA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76804bca-cf2f-41ea-8e36-c8b1e12f6d5e_2048x1209.png 1272w, https://substackcdn.com/image/fetch/$s_!pkQA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76804bca-cf2f-41ea-8e36-c8b1e12f6d5e_2048x1209.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Specialized measurement frameworks can help organize these diagnostic signals. For example, </span><a href="https://getdx.com/research/measuring-ai-code-assistants-and-agents/"><span>DX&#8217;s AI Measurement Framework</span></a><span> combines AI-specific telemetry around utilization and cost with outcome-oriented metrics to evaluate AI&#8217;s overall impact on engineering organizations. These two classes of measurement answer fundamentally different questions: &#8220;How is work being performed?&#8221; versus &#8220;Is the engineering organization delivering better outcomes?&#8221;</span></p><p><span>The value of tracking AI activity is that it helps us understand the shifting </span><em><span>patterns</span></em><span> that lead to our outcomes. For example, if a team&#8217;s AI adoption spikes to 90%, that metric alone doesn&#8217;t prove success. Instead, it serves as a lens to interpret changes in the Core 4: did that spike in adoption correlate with an increase in speed? Did it negatively impact quality via a higher change failure rate? Or did it inadvertently degrade developer effectiveness by introducing new code-review bottlenecks?</span></p><p><span>Tracking AI telemetry tells us how the work is changing. Tracking the core dimensions tells us if that change is actually delivering results.</span></p><p><span>When leaders are tasked with proving AI investment ROI, they cannot do it by pointing to adoption spikes or token volume. A high utilization rate means nothing if software delivery stalls or system stability crashes. Outcome-based developer experience metrics aren&#8217;t just a way to measure engineering anymore, they may be the most reliable ledger for proving AI value.</span></p><h3><span>PR throughput in the AI era</span></h3><p><span>Of the key metrics within the Core 4, PR throughput has attracted the most debate, both before and after the arrival of AI.</span></p><p><span>The criticism of PR throughput is entirely fair at the individual level. Not all PRs are created equal in terms of size, complexity, or value. DX developed a methodology called </span><a href="https://getdx.com/truethroughput/"><span>TrueThroughput</span></a><span>, which uses AI to normalize these variations by weighting PRs based on actual complexity. Yet, even with that kind of normalization in place, the metric is a poor instrument for evaluating any individual developer&#8217;s contribution. I&#8217;ve argued this myself, and I&#8217;d stand by it. Using PR throughput to assess individuals is the wrong application of the metric.</span></p><p><span>At the system level, though, it remains one of the most useful signals available. The reason is that it doesn&#8217;t just measure output, it measures engineering flow. Whether code in a pull request was written by a human or generated by an AI agent, if it&#8217;s moving through review, CI, and deployment without friction, the metric reflects that. If it&#8217;s stalling&#8212;because review is bottlenecked, builds are flaky, or deployment processes are slow&#8212;the metric surfaces that too. PR throughput is a signal for whether an engineering system can move work through, regardless of where that work originates.</span></p><p><span>It also occupies a unique position among the Core 4 metrics. Unlike measures such as Change Failure Rate or DXI, which continue to evaluate enduring organizational outcomes, PR throughput is directly tied to the mechanics of software delivery. As workflows evolve from code-first to intent-first development, the role of the pull request itself may change substantially, making PR throughput more susceptible to reinterpretation than most other metrics in the framework.</span></p><p><span>In </span><a href="https://newsletter.getdx.com/p/ai-productivity-gains-more-modest-than-expected"><span>our own longitudinal research at DX,</span></a><span> we found that AI coding tools produced roughly a 7.8% increase in PR throughput across organizations that had adopted them. That&#8217;s a real and meaningful signal. It&#8217;s also a useful corrective to more optimistic claims about AI&#8217;s productivity impact. The gains are real; they tend to be more modest than headline figures suggest, and they vary considerably across different types of work.</span></p><p><span>The majority of code shipped in production &#9;is still written by humans, though that share is shifting. </span><a href="https://newsletter.getdx.com/p/ai-generated-merged-code-holds-steady"><span>Our research</span></a><span> showed that during the first quarter of 2026, the percentage of code generated by AI that reaches production is 27.4% of production code on average. For most engineering organizations today, pull requests remain the primary unit of software delivery, making PR throughput one of the clearest indicators of engineering system flow.</span></p><p><span>If, and when, the transition to intent-first workflows materializes, the field will likely need a metric that captures innovation velocity as a higher level of abstraction. The </span><strong><span>Idea-to-Customer</span></strong><span> velocity metric introduced in the recent </span><a href="https://arxiv.org/abs/2605.04259"><span>EngThrive framework paper</span></a><span> is one implementation worth watching as a future key metric for the speed dimension. But even in that future, PR throughput will likely remain a crucial secondary metric for diagnosing system flow.</span></p><h3><span>Evolving the interpretation, not the framework</span></h3><p><span>To recap, the top-level dimensions of the DX Core 4 are stable and as meaningful as ever. The key metrics that support them also continue to hold.</span></p><p><span>What is changing is the diagnostic layer beneath them, the operational signals that have always helped explain how engineering systems produce those outcomes. AI doesn&#8217;t change what good looks like at the outcome level, but it does change the mechanisms that generate many of our familiar diagnostic metrics. The same number can now be produced by very different combinations of human and AI behavior, which means individual diagnostic metrics are noisier than they used to be, and the signals they do provide may relate to outcomes in different ways than they used to.</span></p><p><span>Take, for example:</span></p><ul><li><p><strong><span>PR Merge Rate:</span></strong><span> Historically, a high merge rate signaled a highly aligned team shipping clean, uncontroversial work. In an agentic workflow, does a 95% merge rate mean the AI is flawless? Or does it mean your human developers are rubber-stamping machine-generated code because they&#8217;re too overwhelmed to properly review it?</span></p></li><li><p><strong><span>Time-to-10th-PR:</span></strong><span> This is currently one of my favorite onboarding metrics because it is highly predictive of a new hire&#8217;s long-term success and speed-to-productivity. But its utility faces an unresolved question: if an AI onboarding assistant can help an engineer generate and ship 10 PRs by their second afternoon, does that metric still capture true structural onboarding health? Or does it just track how quickly someone learned to use AI tools?</span></p></li></ul><p><span>This is the core challenge. The data points themselves have not changed, but the behaviors that generate them have. AI activity metrics, such as tool adoption or token counts, provide critical context for understanding why traditional engineering metrics move the way they do, but they do not replace those metrics.</span></p><p><span>Triangulating between diagnostic metrics, engineering system metrics, and high-level outcome metrics is what lets us translate how teams work into whether they&#8217;re achieving what they set out to. Building a map of these new patterns&#8212;how to interpret them, and what outcomes they predict&#8212;will be critical work for engineering teams and researchers moving forward.</span></p><h2><span>Final thoughts</span></h2><p><span>The instinct to reach for entirely new metrics in this age of AI is understandable. AI is genuinely reshaping how software gets built, and it is reasonable to question whether existing measurement frameworks can keep pace.</span></p><p><span>But our research and data show that the core dimensions of productivity have held up, not because they anticipated AI specifically, but because they were designed around enduring organizational outcomes rather than any particular workflow or technology. Speed, effectiveness, quality, and business impact remain the right questions to ask, whether code is written by a developer at a terminal or generated by an autonomous agent.</span></p><p><span>What has changed is not what we should measure, but how we should interpret it. AI-specific telemetry provides valuable diagnostic context for understanding how work is evolving, but it does not replace outcome-oriented measurement. Likewise, familiar engineering metrics such as PR throughput, merge rates, or onboarding velocity continue to provide meaningful signals, even as the behaviors that generate those signals shift.</span></p><p><span>The priority for engineering leaders is not to rebuild their measurement architecture from scratch. It is to learn to interpret existing frameworks through a new lens, one that recognizes the growing role of AI while remaining anchored to the outcomes that ultimately matter.</span></p><p><span>The framework is stable. The interpretation is where the real work begins.</span></p><div><hr></div><p>That&#8217;s it for this week. Thanks for reading.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.getdx.com/p/revisiting-the-dx-core-4?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.getdx.com/p/revisiting-the-dx-core-4?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p>]]></content:encoded></item><item><title><![CDATA[Beyond the CLI: Agentic AI for async workloads and non-developers ]]></title><description><![CDATA[How Airbnb scaled AI adoption without mandates, why agentic AI is reshaping product development, and the infrastructure powering its vision for AI-first engineering.]]></description><link>https://newsletter.getdx.com/p/beyond-the-cli-agentic-ai-for-async</link><guid isPermaLink="false">https://newsletter.getdx.com/p/beyond-the-cli-agentic-ai-for-async</guid><dc:creator><![CDATA[Justin Reock]]></dc:creator><pubDate>Mon, 22 Jun 2026 13:46:55 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/202175518/3f6d406e67f5579715913bd8049e8875.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>Listen and watch now on <strong><a href="https://youtu.be/lL9-yATNAo0">YouTube</a>, <a href="https://podcasts.apple.com/us/podcast/engineering-enablement-by-abi-noda/id1619140476">Apple</a>, and <a href="https://open.spotify.com/show/3NxjyIsuxeDMQtisDqBy7D">Spotify</a></strong>.</p><p>In this session from DX Annual, Christopher Sanson, Product Lead, AI Developer Experience, and Madison Capps, Engineering Manager, Infrastructure at Airbnb, challenge some of the most common assumptions about AI. Is AI primarily about replacing humans? Do organizations need mandates to drive adoption? And are the productivity gains really as small as some studies suggest?</p><p>Using examples from Airbnb&#8217;s own AI journey, they share how the company achieved widespread adoption of agentic AI through AirChat, community enablement, and internal tooling rather than top-down mandates. They also discuss the impact AI is having on developer productivity, how non-developers are increasingly using coding tools, and how teams are rethinking product development in an AI-first world.</p><p>Finally, Madison takes a deeper look at the infrastructure powering Airbnb&#8217;s AI strategy, including AirChat CLI, the AirChat SDK, and AirChat Remote, along with the company&#8217;s vision for asynchronous agent workflows and the next generation of AI-powered development.</p><div id="youtube2-lL9-yATNAo0" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;lL9-yATNAo0&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/lL9-yATNAo0?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h2><strong>Some takeaways: </strong></h2><p><strong>AI adoption at scale</strong></p><ul><li><p><strong>Successful AI adoption does not require mandates.</strong> Airbnb achieved 97% weekly usage and 90% daily usage of agentic AI tools among engineers without tying adoption to performance reviews or quotas. Christopher argued that the best adoption comes when developers choose to use AI because it genuinely helps them work faster and better.</p></li><li><p><strong>Treat internal AI tools like products, not internal infrastructure.</strong> Airbnb built a recognizable brand around AirChat, invested in onboarding and workshops, created internal marketing materials, and focused heavily on user experience. That product mindset helped turn AirChat into a company-wide platform rather than just another engineering tool.</p></li><li><p><strong>Community-driven learning scales better than centralized training.</strong> AI champions, train-the-trainer programs, hackathons, workshops, and active peer-to-peer learning channels allowed knowledge to spread organically across the company. Over time, the AI community became larger and more active than the team managing the platform itself.</p></li></ul><p><strong>Productivity gains are accelerating</strong></p><ul><li><p><strong>Developers are spending more time actively coding.</strong> Christopher challenged the idea that engineers only spend a small percentage of their time writing code. As coding becomes faster and easier with agentic AI, developers can spend more of their week building software rather than working around implementation bottlenecks.</p></li><li><p><strong>The most active AI users see the largest productivity gains.</strong> Airbnb found that developers who spent four or more hours per day working with agentic AI dramatically increased their output. The relationship between AI usage and productivity became stronger as engineers learned how to incorporate agents into their daily workflows.</p></li><li><p><strong>PR throughput increased by 65% after the introduction of agentic AI.</strong> Airbnb&#8217;s data suggests that productivity gains extend well beyond the single-digit improvements often cited in industry studies. Developers who heavily embraced agentic AI moved from industry-average output to some of the highest throughput levels measured internally.</p></li><li><p><strong>AI-authored code is becoming mainstream.</strong> Roughly 59% of Airbnb&#8217;s code is now primarily authored by AI, and more than half of developers report that AI generates the majority of the code they work with. Christopher argued that this shift is happening far faster than most organizations realize.</p></li></ul><p><strong>AI is spreading beyond engineering</strong></p><ul><li><p><strong>The addressable market for AI is much larger than developers alone.</strong> Airbnb initially expected adoption to level off around its engineering population. Instead, usage continued growing as product managers, designers, finance teams, and operations teams began integrating agentic AI into their work.</p></li><li><p><strong>People will learn new workflows when the value is obvious.</strong> Some non-engineering teams adopted VS Code and terminal-based tools simply because they provided the best access to agentic AI capabilities. Rather than resisting technical tools, employees were willing to learn them in exchange for meaningful productivity gains.</p></li><li><p><strong>Domain experts are increasingly building their own AI-powered solutions.</strong> Airbnb&#8217;s internal platforms allow teams to create specialized applications tailored to their own workflows. This shifts more problem-solving into the hands of the people closest to the business problem.</p></li></ul><p><strong>Rethinking how work gets done</strong></p><ul><li><p><strong>Many existing processes were designed around expensive software development.</strong> Product reviews, lengthy requirements documents, and sequential handoffs evolved in a world where implementation was slow and costly. AI changes those economics and creates opportunities to redesign workflows from first principles.</p></li><li><p><strong>AI enables faster movement from ideas to prototypes.</strong> Rather than spending weeks refining specifications before building anything, teams can generate multiple prototypes quickly, test ideas earlier, and iterate before committing significant resources.</p></li><li><p><strong>Smaller teams can collaborate earlier and move faster.</strong> Airbnb sees opportunities to reduce handoffs between product managers, designers, and engineers by bringing teams together earlier in the process and using AI to accelerate exploration and execution.</p></li></ul><p><strong>Building for asynchronous AI</strong></p><ul><li><p><strong>Current agentic AI tooling still creates friction.</strong> Managing multiple sessions, handling long-running tasks, maintaining context, and switching between workflows remain cumbersome despite major advances in model capabilities.</p></li><li><p><strong>The next frontier is asynchronous agent workflows.</strong> Rather than interacting with a single agent in real time, developers are increasingly orchestrating multiple agents working in parallel, often across long-running tasks that continue without constant supervision.</p></li><li><p><strong>Airbnb is investing in infrastructure, not just models.</strong> AirChat CLI, migration tooling, the AirChat SDK, and AirChat Remote were all built around the belief that future gains will come from workflow orchestration, platform capabilities, and developer experience as much as from improvements in foundation models.</p></li></ul><p><strong>Preparing for an AI-first future</strong></p><ul><li><p><strong>Organizations should build for where developer workflows are heading.</strong> Madison described Airbnb&#8217;s approach as continuously forecasting how engineers are likely to work in the near future and investing in the infrastructure required to support those workflows before they become mainstream.</p></li><li><p><strong>AI-first architecture will become increasingly important.</strong> As throughput rises and more work is delegated to agents, teams will need stronger guardrails, scalable platforms, and systems designed specifically to support AI-assisted development.</p></li><li><p><strong>The biggest bottlenecks are shifting away from code generation.</strong> As AI reduces implementation costs, constraints move elsewhere in the system. Coordination, validation, infrastructure, and workflow management are becoming the new challenges organizations must solve.</p></li></ul><h2><strong>In this episode, we cover:</strong></h2><p>(<a href="https://www.youtube.com/watch?v=lL9-yATNAo0">00:00</a>) Intro</p><p>(<a href="https://www.youtube.com/watch?v=lL9-yATNAo0&amp;t=97s">01:37</a>) Myth #1: AI is about replacing humans</p><p>(<a href="https://www.youtube.com/watch?v=lL9-yATNAo0&amp;t=202s">03:22</a>) Myth #2: You need mandates to drive AI adoption</p><p>(<a href="https://www.youtube.com/watch?v=lL9-yATNAo0&amp;t=321s">05:21</a>) AirChat, agentic AI, and Airbnb&#8217;s adoption strategy</p><p>(<a href="https://www.youtube.com/watch?v=lL9-yATNAo0&amp;t=487s">08:07</a>) Myth #3: AI has little impact on productivity</p><p>(<a href="https://www.youtube.com/watch?v=lL9-yATNAo0&amp;t=573s">09:33</a>) Airbnb&#8217;s increase in coding time and PR throughput</p><p>(<a href="https://www.youtube.com/watch?v=lL9-yATNAo0&amp;t=860s">14:20</a>) Myth #4: AI coding tools are just for coders</p><p>(<a href="https://www.youtube.com/watch?v=lL9-yATNAo0&amp;t=939s">15:39</a>) How non-developers are using coding tools</p><p>(<a href="https://www.youtube.com/watch?v=lL9-yATNAo0&amp;t=1044s">17:24</a>) Rethinking product development in an AI-first world</p><p>(<a href="https://www.youtube.com/watch?v=lL9-yATNAo0&amp;t=1230s">20:30</a>) Myth #5: Vibe coding isn&#8217;t coding</p><p>(<a href="https://www.youtube.com/watch?v=lL9-yATNAo0&amp;t=1336s">22:16</a>) Unsolved problems in agentic AI tooling and how Airbnb is addressing them</p><p>(<a href="https://www.youtube.com/watch?v=lL9-yATNAo0&amp;t=1590s">26:30</a>) Airbnb&#8217;s overall AI philosophy in practice</p><p>(<a href="https://www.youtube.com/watch?v=lL9-yATNAo0&amp;t=1755s">29:15</a>) Using agentic AI to accelerate code migrations</p><p>(<a href="https://www.youtube.com/watch?v=lL9-yATNAo0&amp;t=1818s">30:18</a>) AirChat SDK: How Airbnb enables teams to build AI-powered applications</p><p>(<a href="https://www.youtube.com/watch?v=lL9-yATNAo0&amp;t=1997s">33:17</a>) AirChat Remote and asynchronous agent workflows</p><p>(<a href="https://www.youtube.com/watch?v=lL9-yATNAo0&amp;t=2167s">36:07</a>) Predictions for what&#8217;s next</p><p><strong>Where to find Christopher Sanson:</strong></p><p>&#8226; LinkedIn: <a href="https://www.linkedin.com/in/christophersanson">https://www.linkedin.com/in/christophersanson</a></p><p><strong>Where to find Madison Capps:</strong></p><p>&#8226; LinkedIn: <a href="https://www.linkedin.com/in/madison-capps-66950625">https://www.linkedin.com/in/madison-capps-66950625</a></p><h2><strong>Referenced:</strong></h2><p>&#8226; <a href="https://www.airbnb.com/">Airbnb</a></p><p>&#8226; <a href="https://hbr.org/2011/10/steve-jobss-bicycles-for-the-m">Steve Jobs&#8217;s Bicycles for the Mind</a></p><p>&#8226; <a href="https://www.linkedin.com/in/jennifer-st-pierre-4935a81">Jennifer St Pierre</a></p><p>&#8226; <a href="http://linkedin.com/in/justinreock">Justin Reock</a></p><p>&#8226; <a href="https://getdx.com/blog/ai-generated-merged-code-holds-steady-at-30/">AI-generated merged code holds steady at ~30%</a></p><p>&#8226; <a href="https://x.com/karpathy/status/2015883857489522876">Andrej Karpathy&#8217;s post on X</a></p>]]></content:encoded></item><item><title><![CDATA[The future of engineering at Nationwide, Comcast, TD, and HPE]]></title><description><![CDATA[Leaders from Nationwide, Comcast, TD Bank, and HPE share how large enterprises are building AI-first engineering organizations and preparing for the future of software development.]]></description><link>https://newsletter.getdx.com/p/the-future-of-engineering-at-nationwide</link><guid isPermaLink="false">https://newsletter.getdx.com/p/the-future-of-engineering-at-nationwide</guid><pubDate>Mon, 22 Jun 2026 13:43:30 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/202178587/4892229a6aa383b2a148995c7f3b58d1.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>Listen and watch now on <strong><a href="https://youtu.be/cblHlTvfNFc">YouTube</a>, <a href="https://podcasts.apple.com/us/podcast/engineering-enablement-by-abi-noda/id1619140476">Apple</a>, and <a href="https://open.spotify.com/show/3NxjyIsuxeDMQtisDqBy7D">Spotify</a></strong>.</p><p>In this session from DX Annual, Rebecca Fitzhugh, Lead Principal Engineer at Atlassian, moderates a panel featuring Nidhi Allipuram, Vice President, Enterprise Developer Experience and Platform at Nationwide, Jai Schniepp, Senior Director, DevX Product Management at Comcast, Brent Foster, Vice President and Head of Architecture and Strategy at TD Bank, and Praveena Patchipulusu, Vice President of Engineering at HPE.</p><p>Together, they discuss how large enterprises are approaching AI adoption, what it takes to build an AI-first software development lifecycle, and how engineering leaders are balancing speed, security, governance, and developer experience. They also share their perspectives on the changing role of engineers, human accountability, and how organizations can prepare for the future of software engineering.</p><div id="youtube2-cblHlTvfNFc" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;cblHlTvfNFc&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/cblHlTvfNFc?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h2><strong>Some takeaways: </strong></h2><p><strong>Building an AI-first software development lifecycle</strong></p><ul><li><p><strong>AI adoption is becoming a redesign effort, not a tooling effort.</strong> Several panelists argued that the biggest opportunity is not simply adding AI assistants to existing workflows but rethinking the software development lifecycle itself. Rather than treating AI as a coding tool, organizations are beginning to integrate it into requirements gathering, design, testing, code reviews, and deployment.</p></li><li><p><strong>Training and organizational support matter more than tool selection.</strong> Nationwide found that productivity gains came less from introducing new tools and more from providing engineers with training, coaching, playbooks, and time to learn. Teams consistently reported that air cover, psychological safety, and opportunities to experiment were more valuable than access to additional AI products.</p></li><li><p><strong>Successful adoption requires systems, not mandates.</strong> Organizations cannot simply tell teams to &#8220;go use AI.&#8221; Several panelists described building AI champion programs, governance models, embedded coaching, and structured learning opportunities that help teams develop new habits and scale adoption across large enterprises.</p></li></ul><p><strong>Keeping humans accountable</strong></p><ul><li><p><strong>Humans remain responsible for outcomes regardless of who writes the code.</strong> Every panelist emphasized that accountability does not shift to AI. Whether code is generated by an engineer, a copilot, or an agent, humans remain responsible for validating outputs, making decisions, and owning the results delivered to customers.</p></li><li><p><strong>Validation is becoming more important than approval.</strong> Traditional approval processes may matter less than ensuring the right people validate assumptions, outcomes, and risks. Teams are increasingly focused on creating workflows where humans review and challenge AI-generated work rather than simply acting as signoff gates.</p></li><li><p><strong>Decision-making is becoming a core engineering skill.</strong> As AI takes over more implementation work, engineers are spending more time evaluating tradeoffs, validating outputs, and making judgment calls. The ability to make good decisions quickly may become a larger differentiator than the ability to manually write code.</p></li></ul><p><strong>Security and governance in an AI-powered world</strong></p><ul><li><p><strong>Shift-left practices become even more important with AI.</strong> Security, compliance, and quality checks are being pushed earlier into the development process. Rather than relying on reviews at the end of the pipeline, organizations are embedding guardrails directly into workflows and development platforms.</p></li><li><p><strong>AI-generated infrastructure introduces new challenges.</strong> The conversation extended beyond application code to infrastructure. As AI increasingly generates Terraform, YAML, and cloud configuration files, organizations must build policy-driven validation and security controls to prevent vulnerabilities from entering production environments.</p></li><li><p><strong>Context is both a powerful asset and a potential risk.</strong> One of AI&#8217;s greatest strengths is its ability to use organizational knowledge and historical context. At the same time, exposing that information to AI systems creates new security concerns, making governance and access controls increasingly important.</p></li></ul><p><strong>The changing role of the engineer</strong></p><ul><li><p><strong>Engineers are becoming orchestrators rather than implementers.</strong> As AI takes over more boilerplate work, engineers are expected to focus more on system design, architecture, critical thinking, and coordinating work across humans, agents, and platforms. Success increasingly depends on defining intent and evaluating outcomes rather than writing every line of code manually.</p></li><li><p><strong>Role boundaries are becoming less rigid.</strong> The panel described a future where engineers, product managers, designers, and other builders work more closely together. AI is making it easier for individuals to contribute across traditional functional boundaries, creating smaller teams with broader responsibilities.</p></li><li><p><strong>Critical thinking and creativity become more valuable.</strong> While AI can accelerate execution, it cannot replace human judgment and problem framing. Several panelists argued that creativity, curiosity, and the ability to think differently about problems will become increasingly important as AI capabilities continue to improve.</p></li></ul><p><strong>Rethinking developer experience</strong></p><ul><li><p><strong>Developer experience is becoming workflow experience.</strong> The focus is shifting from individual tools toward creating trusted workflows that help teams move from idea to production more quickly. Organizations are increasingly measuring success by how effectively teams can deliver outcomes rather than by how efficiently they write code.</p></li><li><p><strong>Developer experience now includes agent experience.</strong> As AI agents become active participants in software delivery, organizations must consider how agents consume context, operate within guardrails, and interact with development platforms. Designing effective systems now means thinking about both human and AI users.</p></li><li><p><strong>Breaking down silos creates better outcomes.</strong> Several panelists argued that AI provides an opportunity to reduce friction between product managers, designers, developers, and security teams. The organizations that benefit most may be those that remove barriers between disciplines and enable more collaborative ways of working.</p></li></ul><p><strong>Preparing for the future</strong></p><ul><li><p><strong>The time to experiment is now.</strong> Every panelist encouraged organizations to begin learning through direct experience rather than waiting for the technology to mature. Teams that develop AI skills, workflows, and governance practices today will be better positioned as the technology continues to evolve.</p></li><li><p><strong>Institutional knowledge may become a competitive advantage.</strong> Large enterprises possess decades of documentation, decisions, diagrams, and expertise that often remain difficult to access. Several speakers highlighted the opportunity to unlock that knowledge and make it useful through AI-powered systems.</p></li><li><p><strong>Fundamentals still matter.</strong> Despite rapid technological change, the panel repeatedly returned to the same conclusion: strong engineering fundamentals, sound judgment, accountability, security practices, and critical thinking remain essential regardless of how much AI enters the software development process.</p></li></ul><h2><strong>In this episode, we cover:</strong></h2><p>(<a href="https://www.youtube.com/watch?v=cblHlTvfNFc">00:00</a>) Intro</p><p>(<a href="https://www.youtube.com/watch?v=cblHlTvfNFc&amp;t=148s">02:28</a>) The AI journey across TD Bank, Comcast, and HPE</p><p>(<a href="https://www.youtube.com/watch?v=cblHlTvfNFc&amp;t=359s">05:59</a>) Inside Nationwide&#8217;s AI-assisted development lifecycle</p><p>(<a href="https://www.youtube.com/watch?v=cblHlTvfNFc&amp;t=604s">10:04</a>) Reimagining the software development lifecycle with AI</p><p>(<a href="https://www.youtube.com/watch?v=cblHlTvfNFc&amp;t=692s">11:32</a>) Security, governance, and human accountability</p><p>(<a href="https://www.youtube.com/watch?v=cblHlTvfNFc&amp;t=927s">15:27</a>) Embedding security and guardrails into AI workflows</p><p>(<a href="https://www.youtube.com/watch?v=cblHlTvfNFc&amp;t=1075s">17:55</a>) How AI is changing the role of an engineer</p><p>(<a href="https://www.youtube.com/watch?v=cblHlTvfNFc&amp;t=1312s">21:52</a>) What developer experience looks like in the AI era</p><p>(<a href="https://www.youtube.com/watch?v=cblHlTvfNFc&amp;t=1615s">26:55</a>) What software engineering may look like in 2030</p><p>(<a href="https://www.youtube.com/watch?v=cblHlTvfNFc&amp;t=1967s">32:47</a>) How to prepare for the AI-driven future</p><p><strong>Where to find Rebecca Fitzhugh:</strong></p><p>&#8226; LinkedIn: <a href="https://www.linkedin.com/in/rmfitzhugh">https://www.linkedin.com/in/rmfitzhugh</a></p><p>&#8226; X: <a href="https://x.com/RebeccaFitzhugh">https://x.com/RebeccaFitzhugh</a></p><p><strong>Where to find Jai Schniepp:</strong></p><p>&#8226; LinkedIn: <a href="https://www.linkedin.com/in/jessicaschniepp">https://www.linkedin.com/in/jessicaschniepp</a></p><p><strong>Where to find Nidhi Allipuram:</strong></p><p>&#8226; LinkedIn: <a href="https://www.linkedin.com/in/nidhi-allipuram">https://www.linkedin.com/in/nidhi-allipuram</a></p><p><strong>Where to find Brent Foster:</strong></p><p>&#8226; LinkedIn: <a href="https://www.linkedin.com/in/engineeringthefuture">https://www.linkedin.com/in/engineeringthefuture</a></p><p>&#8226; Website: <a href="https://brentfoster.me">https://brentfoster.me</a></p><p><strong>Where to find Praveena Patchipulusu:</strong></p><p>&#8226; LinkedIn: <a href="https://www.linkedin.com/in/praveena-patchipulusu-158741">https://www.linkedin.com/in/praveena-patchipulusu-158741</a></p><h2><strong>Referenced:</strong></h2><p>&#8226; <a href="https://www.atlassian.com/">Atlassian</a></p><p>&#8226; <a href="https://www.td.com/">TD Bank</a></p><p>&#8226; <a href="https://corporate.comcast.com/">Comcast Corporation</a></p><p>&#8226; <a href="https://www.hpe.com/us/en/home.html">Hewlett Packard Enterprise (HPE)</a></p><p>&#8226; <a href="https://www.nationwide.com/">Nationwide</a></p><p>&#8226; <a href="https://github.com/github/spec-kit">GitHub Spec Kit</a></p><p>&#8226; <a href="https://www.linkedin.com/in/abinoda/">Abi Noda</a></p>]]></content:encoded></item><item><title><![CDATA[Uber’s journey of measuring AI impact on developer productivity]]></title><description><![CDATA[How Uber evolved its approach to measuring AI&#8217;s impact on engineering, why traditional productivity metrics are breaking down, and what new frameworks may be needed in an agent-driven future.]]></description><link>https://newsletter.getdx.com/p/ubers-journey-of-measuring-ai-impact</link><guid isPermaLink="false">https://newsletter.getdx.com/p/ubers-journey-of-measuring-ai-impact</guid><pubDate>Mon, 22 Jun 2026 13:42:42 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/202179543/75f0c689e4474be76ee387b4b684a152.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>Listen and watch now on <strong><a href="https://youtu.be/aQmnolXCH_M">YouTube</a>, <a href="https://podcasts.apple.com/us/podcast/engineering-enablement-by-abi-noda/id1619140476">Apple</a>, and <a href="https://open.spotify.com/show/3NxjyIsuxeDMQtisDqBy7D">Spotify</a></strong>.</p><p>As AI becomes embedded in software development, many of the metrics that engineering organizations have relied on for years are starting to break down.</p><p>In this session from DX Annual, Uber&#8217;s Ty Smith and Abhishek Tibrewal share how their approach to measuring AI&#8217;s impact on developer productivity has evolved over time. They walk through the different phases of their measurement journey, from adoption and engagement to measuring impact, ROI, and agentic value, explaining what they chose to measure at each stage, what worked, what failed, and how their thinking changed along the way.</p><p>They also discuss the role of qualitative feedback before telemetry existed, the challenge of identifying meaningful engagement signals, why &#8220;developer years saved&#8221; failed as an ROI metric, and how AI agents forced them to rethink traditional productivity measurements. Finally, they introduce Uber&#8217;s emerging framework built around feature velocity and explore the unanswered questions that remain as software development becomes increasingly agent-driven.</p><div id="youtube2-aQmnolXCH_M" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;aQmnolXCH_M&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/aQmnolXCH_M?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h2><strong>Some takeaways: </strong></h2><p><strong>Why AI breaks traditional productivity metrics</strong></p><ul><li><p><strong>Many measurement frameworks were built for a world where humans wrote most of the code.</strong> As AI agents become more capable, metrics that once provided useful signals can quickly become misleading.</p></li><li><p><strong>Teams should expect their metrics to break.</strong> Uber&#8217;s measurement journey required repeatedly revisiting assumptions as AI-assisted development evolved into agentic workflows.</p></li></ul><p><strong>Start with stakeholder questions</strong></p><ul><li><p><strong>The best metrics answer real business questions.</strong> Uber worked backward from questions about productivity, ROI, investment priorities, and business value instead of collecting data for its own sake.</p></li><li><p><strong>Measurement should support decision-making.</strong> Metrics influence budgets, tooling investments, enablement efforts, and long-term strategy.</p></li><li><p><strong>Use qualitative signals before telemetry exists</strong></p></li><li><p><strong>Qualitative feedback can be the fastest path to insight.</strong> Before AI tooling generated reliable telemetry, Uber relied on surveys, interviews, and experience sampling to understand adoption and guide investments.</p></li><li><p><strong>Behavioral questions are more useful than perception questions.</strong> Asking developers what they actually did produced stronger signals than asking whether they found AI helpful.</p></li><li><p><strong>Measure engagement through behavior, not demographics</strong></p></li><li><p><strong>Behavioral patterns revealed insights that demographics could not.</strong> Role, tenure, and organization offered limited signal compared to how engineers actually used AI tools.</p></li><li><p><strong>A small group of AI power users emerged early.</strong> Studying usage patterns helped Uber identify engineers who were engaging deeply with AI and generating outsized results.</p></li></ul><p><strong>Correlation is not causation</strong></p><ul><li><p><strong>High AI usage does not automatically prove AI caused higher productivity.</strong> The most productive engineers are often the first to adopt new tools.</p></li><li><p><strong>Rigorous analysis matters when making investment decisions.</strong> Uber used causal methods to better understand the true impact of AI-assisted development.</p></li></ul><p><strong>Why measuring AI ROI is difficult</strong></p><ul><li><p><strong>Developer years saved sounded compelling but failed as an ROI metric.</strong> The approach created anxiety around replacement, required constant recalibration, and did not answer the business questions leadership cared about most.</p></li><li><p><strong>Business leaders ultimately care about outcomes.</strong> Time saved is useful context, but value creation, customer impact, and business results matter more.</p></li><li><p><strong>PRs measure activity, features measure value</strong></p></li><li><p><strong>Agentic AI exposes the limitations of activity-based metrics.</strong> A single agent task can generate many pull requests without creating meaningful customer value.</p></li><li><p><strong>Feature velocity became Uber&#8217;s new North Star.</strong> The goal shifted from measuring engineering output to measuring whether valuable capabilities were actually being delivered.</p></li></ul><p><strong>Building an AI-native measurement framework</strong></p><ul><li><p><strong>Feature velocity works alongside supporting metrics.</strong> Flow efficiency, quality, and capability expansion help create a more complete picture of AI&#8217;s impact.</p></li><li><p><strong>PR classification provides important context.</strong> Understanding the type and complexity of work helps distinguish meaningful progress from routine maintenance and toil.</p></li><li><p><strong>The future belongs to outcome-based metrics</strong></p></li><li><p><strong>The most durable metrics are tied to business outcomes rather than engineering activity.</strong> As AI becomes more autonomous, output alone becomes a less reliable signal.</p></li><li><p><strong>Many important questions remain unanswered.</strong> Organizations still need better ways to measure judgment, autonomy, technical debt, and the value created by increasingly agent-driven software development.</p></li></ul><h2><strong>In this episode, we cover:</strong></h2><p>(<a href="https://www.youtube.com/watch?v=aQmnolXCH_M">00:00</a>) Intro</p><p>(<a href="https://www.youtube.com/watch?v=aQmnolXCH_M&amp;t=90s">01:30</a>) Steve Yegge&#8217;s 8 stages of AI-assisted development</p><p>(<a href="https://www.youtube.com/watch?v=aQmnolXCH_M&amp;t=202s">03:22</a>) Uber&#8217;s shift to a generative AI-powered company</p><p>(<a href="https://www.youtube.com/watch?v=aQmnolXCH_M&amp;t=260s">04:20</a>) Uber&#8217;s pre-AI productivity metrics</p><p>(<a href="https://www.youtube.com/watch?v=aQmnolXCH_M&amp;t=415s">06:55</a>) Important questions from stakeholders that previous metrics didn&#8217;t answer</p><p>(<a href="https://www.youtube.com/watch?v=aQmnolXCH_M&amp;t=505s">08:25</a>) How Uber measures AI before telemetry exists</p><p>(<a href="https://www.youtube.com/watch?v=aQmnolXCH_M&amp;t=671s">11:11</a>) Metrics used to measure adoption</p><p>(<a href="https://www.youtube.com/watch?v=aQmnolXCH_M&amp;t=769s">12:49</a>) Measuring engagement</p><p>(<a href="https://www.youtube.com/watch?v=aQmnolXCH_M&amp;t=870s">14:30</a>) Measuring impact</p><p>(<a href="https://www.youtube.com/watch?v=aQmnolXCH_M&amp;t=992s">16:32</a>) The challenge of measuring AI ROI</p><p>(<a href="https://www.youtube.com/watch?v=aQmnolXCH_M&amp;t=1172s">19:32</a>) Rethinking adoption, engagement, and impact for agentic AI</p><p>(<a href="https://www.youtube.com/watch?v=aQmnolXCH_M&amp;t=1561s">26:01</a>) The new north star: Feature velocity</p><p>(<a href="https://www.youtube.com/watch?v=aQmnolXCH_M&amp;t=1721s">28:41</a>) PR classification + feature velocity: the questions it can answer</p><p>(<a href="https://www.youtube.com/watch?v=aQmnolXCH_M&amp;t=1981s">33:01</a>) What comes next and what&#8217;s still unanswered</p><p>(<a href="https://www.youtube.com/watch?v=aQmnolXCH_M&amp;t=2070s">34:30</a>) Lessons learned and what they&#8217;d do differently</p><p>(<a href="https://www.youtube.com/watch?v=aQmnolXCH_M&amp;t=2231s">37:11</a>) Q&amp;A #1: How Uber defines a feature</p><p>(<a href="https://www.youtube.com/watch?v=aQmnolXCH_M&amp;t=2330s">38:50</a>) Q&amp;A #2: Measuring success and AI ROI</p><p><strong>Where to find Abhishek Tibrewal</strong></p><p>&#8226; LinkedIn: <a href="https://www.linkedin.com/in/aabhishektibrewal">https://www.linkedin.com/in/aabhishektibrewal</a></p><p><strong>Where to find Ty Smith:</strong></p><p>&#8226; LinkedIn: <a href="https://www.linkedin.com/in/tyvsmith">https://www.linkedin.com/in/tyvsmith</a></p><h2><strong>Referenced:</strong></h2><p>&#8226; <a href="https://steve-yegge.medium.com/welcome-to-gas-town-4f25ee16dd04">Welcome to Gas Town</a></p><p>&#8226; <a href="https://x.com/dkhos?lang=en">Dara Khosrowshahi (Uber CEO)</a></p>]]></content:encoded></item><item><title><![CDATA[AI-authored code has nearly doubled, but so has PR size]]></title><description><![CDATA[Findings from our analysis of over 400 organizations from the past year.]]></description><link>https://newsletter.getdx.com/p/ai-authored-code-has-nearly-doubled</link><guid isPermaLink="false">https://newsletter.getdx.com/p/ai-authored-code-has-nearly-doubled</guid><dc:creator><![CDATA[Justin Reock]]></dc:creator><pubDate>Wed, 17 Jun 2026 10:06:27 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/fd573283-a40b-4f73-97dd-a0223e7e2c1c_2400x1260.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em><strong>Welcome to the latest issue of Engineering Enablement,</strong> a weekly newsletter sharing research and perspectives on developer productivity.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.getdx.com/p/ai-authored-code-has-nearly-doubled?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.getdx.com/p/ai-authored-code-has-nearly-doubled?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p>&#128467; Join Brian Houck and me, this Thursday June 18th for a research briefing on measuring AI agents, revisiting the Core 4, and more. <a href="https://getdx.com/webinar/research-briefing-with-brian-houck-measuring-ai-agents-revisiting-the-core4/?utm_source=newsletter">Register here.</a></p><div><hr></div><p>In our <a href="https://getdx.com/report/ai-assisted-engineering-Q1-impact-report/?utm_source=newsletter">Q1 AI impact analysis</a>, we found that 27.4% of code was AI-authored. Because the space is changing quickly, the DX Research team reports on this metric quarterly to track changes in AI&#8217;s impact on organizations&#8217; ability to create code.</p><p>To measure the change in AI-authored code, and the impact on quality, we conducted two analyses:</p><ol><li><p>First we measured the <em>percentage of AI-authored code, </em>using self-reported data from developers. We define the metric as code generated by AI without major human rewrites.</p><ol><li><p><em>As with any self-reported metric, there is potential for bias in both directions&#8212;undercounting from fully autonomous workflows and overcounting when developers treat AI use as a performance signal. In the future we&#8217;ll share what we&#8217;re seeing from <a href="https://getdx.com/blog/introducing-ai-code-insights/">DX&#8217;s AI Code Insights</a>, which automatically measures the percentage of AI-generated code.</em></p></li><li><p>Our sample included DX data from over 400 companies from Q2 (April 2026-June 2026), reported by the average of user responses within each company. We interpret this data as estimates of the proportion of coding workload delegated to AI tools, rather than literal measures of code output. This reflects the assumption that respondents anchor to how often they ask AI to do work rather than measuring how much code AI actually produces.</p></li></ol></li><li><p>Additionally, to evaluate the downstream impact of AI-authored code, we also looked at PR size using telemetry data from the same cohort over the last year (July 2025-June 2026).</p><ol><li><p><em>In the future we&#8217;ll share further investigations on the impact of increased AI-authored code, as well as the impact of increased PR size.</em></p></li></ol></li></ol><p>Here&#8217;s what we&#8217;re seeing.</p><h2>AI-authored code is consistent across organization sizes</h2><p>Our preliminary Q2 findings show that, on average, 51.9% of code is now AI-authored. Newer models, better workflow integration due in large part to usage of CLI tools, AI mandates, and learning curve progress have all contributed to this massive shift. While this indicates that AI is significantly impacting our ability to create code, it says little about the quality of the code being generated.</p><p>When segmented by organization size, our finding still holds. The median percentage of code that is AI-authored holds steady at around 50%. This reflects a broader shift in how code is produced, regardless of team size.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rFvB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F720245ff-f845-4ce4-8569-0d71d7ccb9a2_4200x2728.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rFvB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F720245ff-f845-4ce4-8569-0d71d7ccb9a2_4200x2728.png 424w, https://substackcdn.com/image/fetch/$s_!rFvB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F720245ff-f845-4ce4-8569-0d71d7ccb9a2_4200x2728.png 848w, https://substackcdn.com/image/fetch/$s_!rFvB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F720245ff-f845-4ce4-8569-0d71d7ccb9a2_4200x2728.png 1272w, https://substackcdn.com/image/fetch/$s_!rFvB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F720245ff-f845-4ce4-8569-0d71d7ccb9a2_4200x2728.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rFvB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F720245ff-f845-4ce4-8569-0d71d7ccb9a2_4200x2728.png" width="1456" height="946" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/720245ff-f845-4ce4-8569-0d71d7ccb9a2_4200x2728.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:946,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:256373,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.getdx.com/i/201794900?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F720245ff-f845-4ce4-8569-0d71d7ccb9a2_4200x2728.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!rFvB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F720245ff-f845-4ce4-8569-0d71d7ccb9a2_4200x2728.png 424w, https://substackcdn.com/image/fetch/$s_!rFvB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F720245ff-f845-4ce4-8569-0d71d7ccb9a2_4200x2728.png 848w, https://substackcdn.com/image/fetch/$s_!rFvB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F720245ff-f845-4ce4-8569-0d71d7ccb9a2_4200x2728.png 1272w, https://substackcdn.com/image/fetch/$s_!rFvB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F720245ff-f845-4ce4-8569-0d71d7ccb9a2_4200x2728.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Median pull request size has nearly doubled</h2><p>Because of the dramatic change in AI-authored code, we also looked at whether PR size&#8212;one measure for quality&#8212;has changed for the same cohort of companies over the past year. Interestingly, our data is showing an equally dramatic change: median PR size nearly doubled, growing from 44 lines to 72 lines per pull request between July 2025 and June 2026.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ZKUQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bf39fe7-7fbd-4007-8721-cb9f15146107_2048x1429.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ZKUQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bf39fe7-7fbd-4007-8721-cb9f15146107_2048x1429.png 424w, https://substackcdn.com/image/fetch/$s_!ZKUQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bf39fe7-7fbd-4007-8721-cb9f15146107_2048x1429.png 848w, https://substackcdn.com/image/fetch/$s_!ZKUQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bf39fe7-7fbd-4007-8721-cb9f15146107_2048x1429.png 1272w, https://substackcdn.com/image/fetch/$s_!ZKUQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bf39fe7-7fbd-4007-8721-cb9f15146107_2048x1429.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ZKUQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bf39fe7-7fbd-4007-8721-cb9f15146107_2048x1429.png" width="1456" height="1016" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9bf39fe7-7fbd-4007-8721-cb9f15146107_2048x1429.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1016,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ZKUQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bf39fe7-7fbd-4007-8721-cb9f15146107_2048x1429.png 424w, https://substackcdn.com/image/fetch/$s_!ZKUQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bf39fe7-7fbd-4007-8721-cb9f15146107_2048x1429.png 848w, https://substackcdn.com/image/fetch/$s_!ZKUQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bf39fe7-7fbd-4007-8721-cb9f15146107_2048x1429.png 1272w, https://substackcdn.com/image/fetch/$s_!ZKUQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bf39fe7-7fbd-4007-8721-cb9f15146107_2048x1429.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This finding confirms what many teams would expect: AI tends to <a href="https://arxiv.org/html/2603.27130v2#S4">generate more lines of code than humans</a>. When the majority of code is machine-produced, that verbosity results in larger pull requests.</p><p>More broadly, this metric is becoming one of the most important to watch. Generally, more code can equal more complexity, less portability, and a greater potential for bugs and vulnerabilities. More verbose code <a href="https://static0.smartbear.co/support/media/resources/cc/book/code-review-cisco-case-study.pdf">can also be more difficult to review</a> and maintain. One of the traits of a skilled engineer is the ability to fully implement a use case with exactly as much code as needed to perform the task. When AI undermines that instinct at scale, the result is not just <a href="https://arxiv.org/pdf/2603.22106">technical debt</a>. It is compounding cognitive debt across the team as engineers struggle to understand code they did not write.</p><p>The critical question for leaders: have review and quality processes kept pace with this volume? There&#8217;s been a lot of discussion in engineering leadership communities about how to shift processes to handle code review being the new bottleneck. I also appreciated Camille Fournier&#8217;s recent piece sharing <a href="https://skamille.medium.com/guidelines-for-respectful-use-of-ai-affcc85d7072">guidelines for respectable use of AI</a>, which outlines expectations leaders can set with their teams for using AI. If this is something you&#8217;re actively thinking about, please let me know in the comments&#8212;I&#8217;d love to hear from you.</p><div><hr></div><p>That&#8217;s it for this week. Thanks for reading.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.getdx.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.getdx.com/subscribe?"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[From AI experiments to organizational shift: Lessons from Mercari’s transformation]]></title><description><![CDATA[What Mercari learned after mandating 100% AI adoption&#8212;and why faster code generation didn&#8217;t automatically lead to faster software delivery.]]></description><link>https://newsletter.getdx.com/p/from-ai-experiments-to-organizational</link><guid isPermaLink="false">https://newsletter.getdx.com/p/from-ai-experiments-to-organizational</guid><dc:creator><![CDATA[Justin Reock]]></dc:creator><pubDate>Mon, 15 Jun 2026 13:23:22 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/201672920/47acd4f0792320a06193ca103e5f43a3.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>Listen and watch now on <strong><a href="https://youtu.be/Y8LTIZcv66k">YouTube</a>, <a href="https://podcasts.apple.com/us/podcast/engineering-enablement-by-abi-noda/id1619140476">Apple</a>, and <a href="https://open.spotify.com/show/3NxjyIsuxeDMQtisDqBy7D">Spotify</a></strong>.</p><p>Michael Galloway leads Platform Engineering at Mercari, while Snehal Shinde leads Cost and Performance Engineering. Together, they have been at the center of Mercari&#8217;s effort to become an AI-native company.</p><p>In this session from DX Annual, Michael and Snehal share what happened after Mercari&#8217;s CEO mandated 100% AI adoption across the organization. While AI accelerated code generation and increased engineering output, the team quickly discovered that their existing dashboards could not answer a simple question: was AI actually improving productivity?</p><p>They discuss how Mercari built new visibility into AI usage and software delivery, the bottlenecks they uncovered across the SDLC, why faster coding did not automatically translate into faster delivery, and the lessons they learned rolling out AI at scale. They also share how Mercari is rethinking software development around agents, feedback loops, and new ways of working.</p><div id="youtube2-Y8LTIZcv66k" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;Y8LTIZcv66k&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/Y8LTIZcv66k?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h2><strong>Some takeaways: </strong></h2><p><strong>Measuring AI impact</strong></p><ul><li><p><strong>AI adoption alone does not guarantee business value.</strong> Mercari found that while AI usage increased rapidly across the organization, existing dashboards could not answer the leadership team&#8217;s most important question: whether AI was actually improving productivity.</p></li><li><p><strong>Local optimization does not necessarily improve system-wide performance.</strong> Engineers reported working faster with AI tools, but end-to-end delivery metrics remained largely unchanged because bottlenecks elsewhere in the software delivery process continued to slow teams down.</p></li><li><p><strong>Organizations need visibility into both AI usage and delivery outcomes.</strong> Mercari built new dashboards that combined AI tool data with SDLC metrics to better understand adoption, throughput, quality, and operational performance.</p></li></ul><p><strong>The reality of becoming AI-Native</strong></p><ul><li><p><strong>AI adoption required a cultural transformation, not just a tooling rollout.</strong> Mercari&#8217;s CEO mandated company-wide AI adoption, but success depended on changing workflows, habits, and expectations across engineering, finance, legal, customer support, and other functions.</p></li><li><p><strong>Different teams required different forms of enablement.</strong> Employees varied significantly in their technical backgrounds and familiarity with AI tools, making education, workshops, and support systems essential to driving adoption.</p></li><li><p><strong>The goal was to rethink work itself.</strong> Rather than layering AI onto existing processes, Mercari challenged teams to reconsider what they built, how they built it, and how people worked together.</p></li></ul><p><strong>The bottlenecks AI exposed</strong></p><ul><li><p><strong>AI revealed problems that already existed inside the organization.</strong> Review queues, CI instability, deployment friction, and support requests became more visible as coding accelerated.</p></li><li><p><strong>Code generation was not the primary constraint.</strong> Engineers often spent more time waiting for approvals, navigating organizational boundaries, and dealing with infrastructure limitations than writing code.</p></li><li><p><strong>System complexity amplified AI-related challenges.</strong> As AI-generated changes increased, existing architectural complexity and fragile workflows became harder to ignore.</p></li></ul><p><strong>Finding AI workflow opportunities</strong></p><ul><li><p><strong>Mercari mapped AI opportunities across 33 domains.</strong> The AI task force reviewed functional areas across the company to identify where AI could automate work, where it could not, and where the strongest leverage points existed.</p></li><li><p><strong>The biggest opportunities extended far beyond engineering.</strong> Role-specific workshops helped teams in finance, legal, design, operations, customer support, and other departments find practical AI use cases in their own workflows.</p></li><li><p><strong>Early wins created proof points for broader adoption.</strong> Mercari saw measurable impact from support bots, accounting workflows, platform support automation, and Socrates, an internal BI agent that made company data easier to query and use.</p></li></ul><p><strong>Rethinking software development</strong></p><ul><li><p><strong>Faster coding shifted attention upstream.</strong> As implementation became easier, planning, specification, and decision-making emerged as larger constraints on delivery speed.</p></li><li><p><strong>Agent Spec-Driven Development moves AI earlier in the lifecycle.</strong> Mercari began using agents to analyze documentation, code, and organizational knowledge before implementation work started.</p></li><li><p><strong>Future workflows will focus more on intent than execution.</strong> Teams increasingly define goals, constraints, and success criteria while agents handle larger portions of implementation and validation.</p></li></ul><p><strong>Preparing for an agent-driven future</strong></p><ul><li><p><strong>Feedback loops matter more than ever.</strong> Mercari&#8217;s multi-loop SDLC emphasizes rapid validation, iterative learning, and increasingly autonomous agent workflows.</p></li><li><p><strong>Behavioral change remains harder than technological change.</strong> Organizations must rethink ownership, accountability, and trust before they can fully benefit from agent-based development.</p></li><li><p><strong>The path to AI-native development is iterative.</strong> Mercari expects continued setbacks and learning cycles, applying what they describe as the Stockdale Paradox: maintaining confidence in the destination while remaining honest about current challenges.</p></li></ul><h2><strong>In this episode, we cover:</strong></h2><p>(<a href="https://www.youtube.com/watch?v=Y8LTIZcv66k">00:00</a>) Intro</p><p>(<a href="https://www.youtube.com/watch?v=Y8LTIZcv66k&amp;t=106s">01:46</a>) Mercari&#8217;s scale and engineering culture</p><p>(<a href="https://www.youtube.com/watch?v=Y8LTIZcv66k&amp;t=171s">02:51</a>) DX awards at Mercari</p><p>(<a href="https://www.youtube.com/watch?v=Y8LTIZcv66k&amp;t=224s">03:44</a>) Mercari&#8217;s push to become AI-native</p><p>(<a href="https://www.youtube.com/watch?v=Y8LTIZcv66k&amp;t=394s">06:34</a>) The mandate to rethink everything</p><p>(<a href="https://www.youtube.com/watch?v=Y8LTIZcv66k&amp;t=482s">08:02</a>) Mercari&#8217;s AI visibility problem and how they solved it</p><p>(<a href="https://www.youtube.com/watch?v=Y8LTIZcv66k&amp;t=690s">11:30</a>) Mercari&#8217;s early findings on AI implementation</p><p>(<a href="https://www.youtube.com/watch?v=Y8LTIZcv66k&amp;t=1127s">18:47</a>) Closing the AI awareness gap at Mercari</p><p>(<a href="https://www.youtube.com/watch?v=Y8LTIZcv66k&amp;t=1271s">21:11</a>) Mapping AI opportunities across Mercari</p><p>(<a href="https://www.youtube.com/watch?v=Y8LTIZcv66k&amp;t=1892s">31:32</a>) Unpacking the results from the second rollout</p><p>(<a href="https://www.youtube.com/watch?v=Y8LTIZcv66k&amp;t=2054s">34:14</a>) Agent spec-driven development and what&#8217;s next</p><p>(<a href="https://www.youtube.com/watch?v=Y8LTIZcv66k&amp;t=2257s">37:37</a>) A multi-loop SDLC</p><p>(<a href="https://www.youtube.com/watch?v=Y8LTIZcv66k&amp;t=2450s">40:50</a>) Some hard lessons</p><p>(<a href="https://www.youtube.com/watch?v=Y8LTIZcv66k&amp;t=2575s">42:55</a>) Closing thoughts</p><p><strong>Where to find Michael Galloway:</strong></p><p>&#8226; LinkedIn: <a href="https://www.linkedin.com/in/michaelroygalloway">https://www.linkedin.com/in/michaelroygalloway</a></p><p>&#8226; X: <a href="https://x.com/michaelgalloway">https://x.com/michaelgalloway</a></p><p><strong>Where to find Snehal Shinde:</strong></p><p>&#8226; LinkedIn: <a href="https://www.linkedin.com/in/snehal-shinde">https://www.linkedin.com/in/snehal-shinde</a></p><h2><strong>Referenced:</strong></h2><p>&#8226; <a href="https://www.mercari.com/">Mercari</a></p><p>&#8226; <a href="https://cursor.com/">Cursor</a></p><p>&#8226; <a href="https://devin.ai/">Devin</a></p><p>&#8226; <a href="https://www.anthropic.com/product/claude-code">Claude Code | Anthropic&#8217;s agentic coding system</a></p><p>&#8226; <a href="https://github.com/">GitHub</a></p><p>&#8226; <a href="https://www.datadoghq.com/">Datadog</a></p><p>&#8226; <a href="https://www.linkedin.com/in/tbozarth">Tim Bozarth - Microsoft | LinkedIn</a></p><p>&#8226; <a href="https://www.airbnb.com/">Airbnb</a></p><p>&#8226; <a href="https://jimcollins.com/concepts/Stockdale-Concept.html">Jim Collins - Concepts - The Stockdale Paradox</a></p>]]></content:encoded></item></channel></rss>