<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The Verification Gap]]></title><description><![CDATA[AI writes your code. Who verifies the intent?]]></description><link>https://blog.reqproof.com</link><image><url>https://substackcdn.com/image/fetch/$s_!Wvco!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa6b982f-a214-466c-b35b-07bdf5b4eff6_512x512.png</url><title>The Verification Gap</title><link>https://blog.reqproof.com</link></image><generator>Substack</generator><lastBuildDate>Sun, 11 Oct 2026 05:34:31 GMT</lastBuildDate><atom:link href="https://blog.reqproof.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Leonid Bugaev]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[verificationgap@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[verificationgap@substack.com]]></itunes:email><itunes:name><![CDATA[Leonid Bugaev]]></itunes:name></itunes:owner><itunes:author><![CDATA[Leonid Bugaev]]></itunes:author><googleplay:owner><![CDATA[verificationgap@substack.com]]></googleplay:owner><googleplay:email><![CDATA[verificationgap@substack.com]]></googleplay:email><googleplay:author><![CDATA[Leonid Bugaev]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The birth of the Software Verification Engineer]]></title><description><![CDATA[What will replace code reviews]]></description><link>https://blog.reqproof.com/p/the-birth-of-the-software-verification</link><guid isPermaLink="false">https://blog.reqproof.com/p/the-birth-of-the-software-verification</guid><dc:creator><![CDATA[Leonid Bugaev]]></dc:creator><pubDate>Wed, 30 Sep 2026 14:26:14 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/16723883-c852-42b0-b087-bcb83b9ca301_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>TL;DR:</strong> A new engineering profession is emerging: the <strong>Software Verification Engineer</strong>. Manual coding is dead. Agents will write the code and the tests; someone needs to build the rules, environments, and verification infrastructure that make those changes trustworthy. Human review does not disappear. But what we review becomes the intent, the impact, the evidence, and what remains unresolved&#8212;not every line of the implementation.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Verification Gap! Subscribe to keep updated on future of engineering</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><p>The agent writes all of the code, all of the tests, and everything needed to run them. It opens a pull request. The tests pass.</p><p>What am I supposed to review?</p><p>Take an installation change in a product with multiple components. The agent says it tested the installation, and I believe that it ran something. But did it use the setup I wanted it to use? Did it install everything from scratch, or upgrade an existing installation? Were the other components running the latest versions too, or the versions our customers actually have?</p><p>Those are different questions. Reading the changed code does not tell me which of those things happened.</p><p>If I am working on a user interface, I would like to see the video where the agent clicks through the changes and narrates what it is doing. For an API change, I want the actual test attached: what it called, what came back, and what environment it was running against. Basically, the test I would otherwise do manually, except I want the agent to do it and show me the result.</p><p>Now somebody has to make all of that possible. Prepare the environments, define what needs to be checked, capture the results, and make them useful to the person accepting the change.</p><p>That is the engineering function I see emerging. It is much deeper than automatically writing tests.</p><h2>We still need to review something</h2><p>Our end goal is essentially prompt to production. But there is still a human review piece here. Putting a human at the end to read everything the agents produced is not going to sustain it.</p><p>We need to build our engineering processes around agent-written implementation, rather than treating it as an exception that a human will carefully inspect afterwards. All of the code, all of the tests, everything. The question is how we make that process good enough to trust.</p><p>This is why I think code review is dying as the main way we accept changes. We cannot keep increasing the amount of implementation and expect the same human review process to absorb it.</p><p>And when I accept a product change, I am accepting much more than the code diff. It is about the global picture: people, processes, dependencies, and the environments where the product needs to work. Code review is just one of the places where we try to understand that picture.</p><p>So we still need to review something. We need to understand what we are accepting and why we should believe it works. The question is what we put around the change so that a human can make that decision without reconstructing the entire verification process.</p><h2>A good QA agent is not the whole answer</h2><p>Let&#8217;s imagine we already have the fully automated QA person. A QA agent that is actually good at its job, follows instructions, asks the uncomfortable questions, and knows how to investigate something that looks wrong.</p><p>That would be useful. But go back to the installation change. How does this agent get the environments it needs? Where does it find the supported combinations of components? How does it know which upgrade paths matter? What should it capture to demonstrate that the installation worked?</p><p>Somebody needs to prepare the infrastructure, the rules, and the evidence types. Otherwise, we are still relying on someone remembering to ask, &#8220;Did you try it with this configuration?&#8221;</p><p>When it comes to evidence, it is about trust. Different people have different levels of trust, and different companies need it done differently. I might be happy with a recorded walkthrough for one change. Another change might need compatibility tests, failure scenarios, and evidence that we can recover if something goes wrong.</p><p>The recording is not a replacement for all the other checks. It answers a particular question I have about the change. The same is true of a benchmark, an API transcript, or an installation run. Each needs to tell me something relevant.</p><p>So even with a very good QA agent, there is a lot of engineering left to do. We need to decide what evidence is required and build the means to produce it reliably.</p><h2>The rules cannot live only in the prompt</h2><p>It can start with something like code coverage. Let&#8217;s say the agent gets it to 100%. Fine, but what have we established? We have not established that the tests check the right behaviour. High coverage can coexist with weak tests; that distinction existed long before coding agents. <a href="https://martinfowler.com/bliki/TestCoverage.html">martinfowler.com</a></p><p>Now imagine the agent misunderstood the requirement and wrote both the implementation and the tests around that misunderstanding. Everything could pass, and we would still have the wrong change.</p><p><strong>How do we validate that the agent sticks to the original intent and does not divert from it?</strong></p><p>I want the intended behaviour to remain something the implementation is checked against. If the agent discovers that a requirement needs to change, that needs to become an explicit decision. It should not quietly change the requirement, update the tests to match, and present the result as finished.</p><p>The same applies to the guardrails. Telling an agent to follow a rule is not the same as checking that it followed it. Wherever we can turn that rule into an executable check, we should.</p><p>For the installation change, that might mean the pipeline requires an upgrade run against a defined configuration. A fresh installation should not satisfy that requirement. A description of how the upgrade would probably work should not satisfy it either.</p><p>We need to add as much determinism to the system as possible. I am not asking for the agent to take exactly the same steps every time. I want the acceptance rules to hold regardless of which steps it takes. A missing result should remain missing. A failed check should not become a passed check because the agent found a convincing explanation for it.</p><p>And the agent doing the implementation should not have unrestricted control over the rules that decide whether its implementation is acceptable.</p><p>Security warnings need the same discipline. If an AI system says something is vulnerable, what supports that claim? Can it reproduce the problem? Can it show the relevant path and the conditions required? If it cannot establish the finding, I want to know that too. I do not want a suspicion presented as a verified problem, or a confident dismissal presented as proof that everything is fine.</p><p>Every assumption we can check with code should be checked with code. Where we cannot establish something, it should remain visible as unresolved.</p><h2>The agent needs to close its own evidence gaps</h2><p>The loop should not be a human saying, &#8220;You missed the test here. Please add it,&#8221; followed by another review where the human finds the next missing thing.</p><p>That is still a process where the human has to reconstruct the verification plan, one comment at a time.</p><p>I want the agent to be able to ask: what is currently missing from my evidence? Which requirement have I not demonstrated? Which configuration has not been exercised? Did I actually resolve that warning, or just explain it away?</p><p>Then it should be able to do the missing work. Start another environment, run the upgrade, capture the result, investigate the failure, and try again. The pipeline should check whether the required evidence is actually present and whether it satisfies the acceptance rules.</p><p>There has to be a way to stop and say that it cannot finish, too. If a requirement is ambiguous or an environment is unavailable, that should come back as a specific unresolved issue. The agent should not keep going until it finds a way to call the change complete.</p><p>This is where we need an almost inhuman amount of quality work. Not just more generated tests, but much more verification than a human could reasonably carry out for every change.</p><p>There is a useful precedent here. In his essay on testing, Dan Luu described hardware teams investing in custom test generators and running them across thousands of machines. The engineering effort went into systems that could investigate far more cases than people could write individually. <a href="https://danluu.com/testing/">Dan Luu</a></p><p>That is the kind of ambition I want for agentic software engineering. The agents should not just produce more implementation. They should do the enormous amount of work required to demonstrate that the implementation is acceptable.</p><p>I want the evidence to be thorough enough that even the most stubborn QA person can say, &#8220;Yes, I&#8217;m happy.&#8221; But that needs to mean their questions were answered, not that we produced so much output they gave up reading it.</p><h2>Evidence has to make review easier</h2><p>There is an obvious problem here: if we replace a huge code diff with a huge collection of logs, recordings, and reports, we have just moved the review problem somewhere else.</p><p>The evidence needs to be grouped around what the human is accepting. I should be able to understand the conclusion, inspect the relevant result, and see what was not covered. For each result, I want to know which build and configuration it came from, rather than having to guess whether it still applies to the change in front of me.</p><p>Essentially, I need to answer four questions:</p><ul><li><p><strong>What am I accepting?</strong></p></li><li><p><strong>What could it affect?</strong></p></li><li><p><strong>Why should I believe it works?</strong></p></li><li><p><strong>What still remains unresolved?</strong></p></li></ul><p>Some of those questions extend into production. Pre-production evidence does not remove the need to observe a rollout, limit its exposure, and recover from failures. Charity Majors made this point about gradually earning confidence in releases long before this discussion about agents. <a href="https://charity.wtf/p/shipping-software-should-not-be-scary">charity.wtf</a></p><p>So the verification system needs to be clear about what has been established before release, what can only be checked during the rollout, and what should stop that rollout. &#8220;Ready to deploy&#8221; should not quietly become &#8220;nothing can go wrong.&#8221;</p><h2>This is the job of the Software Verification Engineer</h2><p>The work includes the custom harnesses, the environments, the acceptance rules, the evidence formats, and the way those results reach a human. Different industries will need different systems. Different changes within the same product will need different evidence.</p><p>This is not a claim that verification itself is new. It is a change in what the function owns.</p><p>The Software Verification Engineer is responsible for the system that makes agent-produced changes acceptable. What needs to be demonstrated? Which rules must be enforced? What infrastructure does the agent need? How do we know that the evidence is valid? What happens when it is incomplete?</p><p>The agents can help build that infrastructure too. This is not about preserving a special category of code that humans must still write. The responsibility is deciding what the system must enforce and establishing that it actually does.</p><p>We have been talking about better automation, testing infrastructure, and reliable delivery for years. Agentic development is going to force us to connect those things properly. If agents do the implementation, the way forward is to dramatically increase the trust and engineering quality of the pipelines around them.</p><p>For the next installation change, I do not want to remember to ask whether the upgrade was tested. I want the system to know that the upgrade evidence is required, give the agent the means to produce it, and refuse to call the change ready when it is missing.</p><p>Building that system is the job of the <strong>Software Verification Engineer.</strong></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Verification Gap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/p/the-birth-of-the-software-verification?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">If you have read till this moment, you probably have liked it! Pls share with yours colleagues and friends to support my work!</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/p/the-birth-of-the-software-verification?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://blog.reqproof.com/p/the-birth-of-the-software-verification?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p></div>]]></content:encoded></item><item><title><![CDATA[Stop Choosing Between Features, Bugs, and Technical Debt]]></title><description><![CDATA[AI gives us enough engineering capacity to stop cutting corners and start building software that stays correct.]]></description><link>https://blog.reqproof.com/p/stop-choosing-between-features-bugs</link><guid isPermaLink="false">https://blog.reqproof.com/p/stop-choosing-between-features-bugs</guid><dc:creator><![CDATA[Leonid Bugaev]]></dc:creator><pubDate>Tue, 08 Sep 2026 16:06:06 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/9978e7cf-abf6-41eb-a6a3-8d202711f150_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>So much energy in my life was spent making compromises between quality and speed, or pushing technical debt priorities. Constantly growing backlog, and guilt for the choices I had to do. What if we could fix the bug problem with AI? I mean, completely.</p><p>Image the program always follows the designed spec, and stays that way over time. And the only bugs left are human judgment issues, where you had some hypothesis and it was wrong.</p><p>You probably thinking, Leo, it is just not possible, stop dreaming, and talking about some theoretical stuff. Maybe it just hype farming. But trust me, it is not, and pretty much real, and it happening right when we talk.</p><p>If you think about an automobile or MRT, you wouldn&#8217;t expect some stupid bug that will kill a person there. It just works. So why can they do it and we can&#8217;t?</p><p>This was the question I asked myself while looking for the answer to this problem.</p><p>And the answer is essentially that they have a completely different view on quality and on the price of a mistake &#8212; human life. All the issues need to be found before they even reach development, and during development before they reach production.</p><p>And it&#8217;s totally possible. It&#8217;s just very slow, expensive, and requires a lot of processes.</p><p>All the techniques we invented in consumer engineering were about how to cut corners so that you can deliver faster. Everything is about speed and money. Which is fair, but it doesn&#8217;t have to be like that. There are also other options.</p><p>You must say: We are not NASA. We don&#8217;t need the same level of processes.</p><p><strong>But why not?</strong></p><p>If it&#8217;s because of slow, complex, pricey processes, it&#8217;s solvable with AI.</p><p>These days it can manage your requirements management system. It can do the hazard analysis. It can do 100% MC/DC testing for you.</p><p><strong>Why cut corners? </strong>Just throw the tokens at this problem.</p><p>In my view, this is exactly the problem that deserves to be solved by throwing more tokens at it. It will make so many people happier, and it will make all our lives better.</p><p>The whole point with all of this AI evolution is that you can do an inhuman amount of work.</p><p>Just giving people Claude Code or Codex, etc., is still doing human work, just faster and multiplied by 10x, maybe, or something similar.</p><p>We can do so much more with this technology.</p><p>Back to the roots: no longer cutting corners, no longer picking whether I&#8217;m shipping the feature or fixing the bug, and then regretting it while fighting regressions I introduced or some security issues.</p><p>And it will actually solve another big issue.</p><p>We can&#8217;t scale autonomous AI software factories because we don&#8217;t trust their output.</p><p>But trust is something that should be earned.</p><p>No one trusts the engineer in a regulated software engineering company. No one trusts anyone. You have to earn the trust.</p><p>So if we build our workflows around trust and evidence, we can scale agentic workflows as well.</p><p>Let&#8217;s start thinking differently.</p><p>We are at the point of reshaping what software engineering is supposed to be, lets not screw it up.</p><div><hr></div><p>This is exactly what I&#8217;m building at <a href="https://reqproof.com">Proof</a>. I&#8217;m on a mission to bring the engineering standards of regulated industries to everyday software - without the cost and process that made them impractical before. If any of this resonates with you, just reach out. I love talk about software quality, engineering craft, and what AI changes about it.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Verification Gap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Your AI can fix the bug. But can it find it?]]></title><description><![CDATA[My experiments with GLM-5.3-flash, and the difference between following instructions and knowing what to question.]]></description><link>https://blog.reqproof.com/p/your-ai-can-fix-the-bug-can-it-find</link><guid isPermaLink="false">https://blog.reqproof.com/p/your-ai-can-fix-the-bug-can-it-find</guid><dc:creator><![CDATA[Leonid Bugaev]]></dc:creator><pubDate>Mon, 07 Sep 2026 17:42:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!5rTQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787075b3-dffe-462a-9cf2-28703e67f1e7_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5rTQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787075b3-dffe-462a-9cf2-28703e67f1e7_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5rTQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787075b3-dffe-462a-9cf2-28703e67f1e7_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!5rTQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787075b3-dffe-462a-9cf2-28703e67f1e7_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!5rTQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787075b3-dffe-462a-9cf2-28703e67f1e7_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!5rTQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787075b3-dffe-462a-9cf2-28703e67f1e7_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5rTQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787075b3-dffe-462a-9cf2-28703e67f1e7_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/787075b3-dffe-462a-9cf2-28703e67f1e7_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1473425,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.reqproof.com/i/214598297?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787075b3-dffe-462a-9cf2-28703e67f1e7_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!5rTQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787075b3-dffe-462a-9cf2-28703e67f1e7_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!5rTQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787075b3-dffe-462a-9cf2-28703e67f1e7_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!5rTQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787075b3-dffe-462a-9cf2-28703e67f1e7_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!5rTQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787075b3-dffe-462a-9cf2-28703e67f1e7_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I had this crazy idea - what if I can find such algorithm which can find all the bugs in any software, literally ALL. And while going though it (quite successfully actually!), it helped me understand the difference in intelligence of small vs large models. </p><blockquote><p>If TLDR - if you have clear task definition, or at least enough clues, cheaper models can do quite serious work now. If the task require to invent questions themselves - thats what require true intelligence, and cheap models fall short. Also majority of public benchmarks is a form of cheating. </p></blockquote><p>I was working on bisecting on how exactly benchmarks like SWE bench work, how the eval looks like, what is in the golden data. And what I found is that it does not really ask a question to find the bug. Instead it gives a lot of clues in advance, like a user report, sometimes with logs, and ask to fix the bug. Additionally all the benchmarks use popular OSS projects, like Django, and all the answers in the public internet already - all new models have it as part of both pre and post training - which essentially cheating. And you do not even need to enable internet search for that - it is part of training data.</p><p>Compare it with the work on the real project, where there is no one to help. Genuine new problems. Solving such problem require completely different thinking. And thats why even latest frontier models can look stupid on private codebases. </p><p>While everyone was hyped about mystical Ox Alpha model, which ended up to be GLM-5.3-flash, I wasn&#8217;t getting good results with it on my use cases. And specifically trying to find as much bugs as possible, just by looking at the code, without any clues. Flash model was VERY bad for this case, so I started digging. At the same time Grok 4.6 were perfect for this use-case (~5x larger compared to Flash)</p><p>What bothered me was that it could do apparently complicated work. In one experiment, it wrote an 82-line working reproduction of a Django bug. The checks passed. But I had supplied the function, the triggering input and the symptom - I had already told it what was wrong.</p><p><strong>How much of the intelligence was in my prompt?</strong></p><p>For code review, <strong>I need a model to come up with a useful question before I do</strong>. To notice an assumption worth challenging, choose an input that challenges it, and work out whether the failure actually happens. Otherwise, I&#8217;m still doing the difficult part and delegating the typing.</p><p>Here are three examples from Django 3.0. I&#8217;m showing the failing inputs upfront; discovering them was part of the review. I was using Grok alongside GLM because Claude and Codex had refused the security-related work I was doing. Agreeing to participate was already a competitive advantage.</p><h3>1. The guard that handles nothing</h3><p>Here is the filename construction, with the surrounding path handling removed:</p><pre><code><code>if file_hash is not None:
    file_hash = ".%s" % file_hash

hashed_name = "%s%s%s" % (root, file_hash, ext)</code></code></pre><p>What happens when a custom hash implementation returns <code>None</code>?</p><p>For <code>styles.css</code>, you get <code>stylesNone.css</code>. The branch leaves <code>None</code> untouched, and the next line turns it into text. No exception. A valid string containing the wrong filename. </p><p>The check is sitting right there, mentioning the exact value that causes the problem, without handling it correctly. To see the failure, you have to follow the variable past the reassuring-looking condition.</p><p>The GLM reviews discussed directory errors, exception propagation and a possible race. They reached the right function and missed the corrupted filename. Grok found it. </p><p>This is the difference between identifying something worth investigating and completing the investigation. &#8220;What about <code>None</code>?&#8221; is a useful question. &#8220;It might raise a <code>TypeError</code>&#8221; would still be a wrong answer. The actual failure is quieter.</p><h3>2. A comma inside the exponent</h3><p>Django&#8217;s number formatter converts a number into a string, splits around the decimal point, and groups the integer part into thousands. The relevant steps are:</p><pre><code><code>str_number = str(number)

if "." in str_number:
    int_part, dec_part = str_number.split(".")
else:
    int_part, dec_part = str_number, ""</code></code></pre><p>Except that <code>1e25</code> becomes <code>"1e+25"</code>.</p><p>There is no decimal point. The whole string becomes the &#8220;integer part,&#8221; including the exponent. The grouping loop then works backwards through those characters, inserting a separator every three positions. In the audited Django environment, this produced:</p><pre><code><code>from django.utils.numberformat import format

format(
    1e25, ".",
    grouping=3,
    thousand_sep=",",
    force_grouping=True,
    use_l10n=True,
)
# '1e,+25'</code></code></pre><p>A thousands separator inside the exponent. The loop processes characters; naming its input <code>int_part</code> doesn&#8217;t make those characters digits. </p><p>The GLM findings went elsewhere: empty strings, empty grouping sequences, special Decimal values. Whether those are useful findings depends on which inputs the function is supposed to support. They did not identify this transformation. </p><p>To find this bug, knowing what each line does separately is not enough. You have to carry <code>"1e+25"</code> through all of them without silently replacing it with the ordinary decimal number you expected to see.</p><p>That is a small amount of code, but a real demand on the model. The relevant fact is something that <em>isn&#8217;t</em> there: a dot.</p><h3>3. &#8220;Just use self&#8221;</h3><p>The third example comes with a comment that practically invites you to stop thinking:</p><pre><code><code># If the other Q() is empty, ignore it and just use `self`.
if not other:
    return copy.deepcopy(self)</code></code></pre><p>&#8220;Just use self&#8221; sounds harmless. The implementation asks Python to deep-copy the object and its contents.</p><p>My recorded test put a thread lock inside the query. Combining it with an empty <code>Q()</code> reached this copying operation and raised <code>TypeError: cannot pickle '_thread.lock' object</code>. </p><p>A lock is a deliberately awkward input. Django&#8217;s upstream report shows the same problem with something much more ordinary: a dictionary&#8217;s keys. With the empty query on the left, this failed too:</p><pre><code><code>Q() | Q(x__in={}.keys())
# TypeError: cannot pickle 'dict_keys' object</code></code></pre><p>An empty query, which was supposed to contribute nothing, introduced a new requirement: the contents of the other query had to survive copying. </p><p>And you cannot replace that input with anything vaguely &#8220;unpickleable&#8221; and expect the same result. Ordinary functions, including lambdas, are returned unchanged by <code>deepcopy</code>. The exact object matters. </p><p>This case caught Grok too. It reached the <code>deepcopy</code> line but proposed the wrong trigger. A cheaper-class model found the case in one differently framed session, where the comment had been presented as a claim to challenge. That success did not reliably carry over to another experiment with related guidance. </p><p>That is why &#8220;use the stronger model&#8221; is not the whole answer. Finding the suspicious line is progress. Knowing the phrase &#8220;deep copying can fail&#8221; is progress. The review is unfinished until the specific input produces the specific failure.</p><h2>I tried better prompts</h2><p>I changed the instructions, supplied more clues, split the review into phases, and required more explicit checks. Across three versions of the review instructions, GLM-5.3-flash found none of the seven selected hard defects. Zero out of seven each time. These were deliberately difficult cases, not a random sample of everything it can do. </p><p>One experiment required line-by-line execution traces with intermediate values. Thirty-eight of the forty-eight outputs followed that format. The requested structure appeared; the missing findings did not. The record contains confident, well-formed traces that were wrong. </p><p>At that point, &#8220;make it explain its reasoning&#8221; was not much of an answer. I was already reading the explanations. They were the problem.</p><p>The copying example shows that framing can help. But adding more instructions did not reliably buy the judgment I needed. Better compliance was measurable. It was also insufficient.</p><h2>So what are the leaderboards measuring?</h2><p>SWE-bench gives a model a repository and an existing issue, then asks for a patch. Somebody has already noticed a problem and described it. There may still be difficult engineering ahead, but discovering that something is wrong is not the same assignment. </p><p>Why should I read success at repairing described problems as evidence that a model will independently find problems in my code?</p><p>Then there is the possibility that it has already seen the answer. On February 23, 2026, OpenAI explained why it had stopped reporting SWE-bench Verified scores, citing contamination and flawed evaluations. All three model families it tested could reproduce original fixes or problem-specific details for some tasks. </p><p>I don&#8217;t need to prove that every provider deliberately cheats to question the measurement. When the issue, the patch and the regression test are public, a correct answer alone cannot tell me whether the model worked it out.</p><p>That applies to my Django examples too. Their public histories mean Grok&#8217;s successful answers are not proof of discovery on unseen code. These are inspectable, repeatable examples, not my attempt to sell you another intelligence ranking. </p><p>What interests me is the difference between these tasks. A model can implement a precise bug description while struggling to produce that description itself. A single coding score hides exactly the distinction I need when deciding what work to delegate.</p><h2>Make it ask the question</h2><p>Here is an exercise to try with your own model. Give it the original filename code and enough context to review it, without announcing the bug. Then, in a separate session, give it this version:</p><pre><code><code>if file_hash is None:
    file_hash = ""
else:
    file_hash = ".%s" % file_hash</code></code></pre><p>The original <code>None</code> failure is gone. Does the model find it in the first version and stop reporting it in the second? Can it supply an input and predict the actual result? Does its explanation follow the code when you change the behavior?</p><p>Then look at how much help you had to give it. Asking &#8220;what happens when the hash is <code>None</code>?&#8221; has already supplied an important part of the review. Asking it to investigate the function leaves that decision with the model.</p><p>I don&#8217;t care whether it answers from the code or runs experiments first. Give it a runtime. But it still has to choose an experiment that answers a useful question, rather than execute something irrelevant and report success.</p><p>I hope this article gave you more clarify on how intelligence differ in small vs big models and why it is still too early to claim that local models are on par with frontier ones. It is true only if you performed the &#8220;intelligence&#8221; by yourself, and give the model clues - e.g. bug fixing based on well shaped report. </p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Verification Gap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Help your open-source neighbour first]]></title><description><![CDATA[The biggest projects are closing their contribution queues. In smaller communities, there is still someone on the other side.]]></description><link>https://blog.reqproof.com/p/help-your-open-source-neighbour-first</link><guid isPermaLink="false">https://blog.reqproof.com/p/help-your-open-source-neighbour-first</guid><dc:creator><![CDATA[Leonid Bugaev]]></dc:creator><pubDate>Wed, 02 Sep 2026 13:21:10 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!lfLa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F563cea06-ffb9-496c-86de-16459a9e37eb_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!lfLa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F563cea06-ffb9-496c-86de-16459a9e37eb_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!lfLa!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F563cea06-ffb9-496c-86de-16459a9e37eb_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!lfLa!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F563cea06-ffb9-496c-86de-16459a9e37eb_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!lfLa!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F563cea06-ffb9-496c-86de-16459a9e37eb_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!lfLa!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F563cea06-ffb9-496c-86de-16459a9e37eb_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!lfLa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F563cea06-ffb9-496c-86de-16459a9e37eb_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/563cea06-ffb9-496c-86de-16459a9e37eb_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1369731,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.reqproof.com/i/213846782?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F563cea06-ffb9-496c-86de-16459a9e37eb_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!lfLa!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F563cea06-ffb9-496c-86de-16459a9e37eb_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!lfLa!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F563cea06-ffb9-496c-86de-16459a9e37eb_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!lfLa!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F563cea06-ffb9-496c-86de-16459a9e37eb_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!lfLa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F563cea06-ffb9-496c-86de-16459a9e37eb_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Current open source contribution experience leads to guaranteed frustration on both sides</strong> - project and contributor - if you will follow trap of trying contribute to famous projects like Go, React, Node.js, Rust, Kubernetes, or another project everybody knows. All with a dream to find a good first issue, submit a patch, and put the project name on your CV. In the best case scenario you&#8217;ll just get some bot response with rejection. </p><p>The large projects are not waiting for one more unknown person to send them code. Some are saying this openly. This year tldraw started closing external pull requests automatically. Godot introduced stricter contribution rules, especially for large changes and AI-generated code. GitHub added settings that let maintainers restrict who can open pull requests and limit how many an outside contributor can have open.</p><p>And I totally get it, especially while <a href="https://blog.reqproof.com/p/a-few-months-inside-rsync">helping on the latest rsync release</a>. A maintainer can receive a patch in seconds and spend hours working out whether it should exist. The code may look perfectly reasonable. You still have to check the original problem, compatibility, tests, old decisions, and whether the contributor will be around when the next release breaks something. AI made it easier to produce the patch. It did not give maintainers more time to review it.</p><p><strong>So let the large projects close the queue. There is a lot more open source out there.</strong></p><p>That was the reason I built <a href="https://helpwanted.dev/">Help Wanted</a> in the first place, few years ago. At the time of writing, it has 650,456 issues from 96,576 projects in its index, from the projects who actually <strong>WANT</strong> help. Most developers will never hear the names of nearly all these projects.</p><p>During the last Hacktoberfest fest I used Help Wanted to find smaller projects and opened three pull requests. All three were r<strong>eviewed the same day</strong>. Someone was actually on the other side.</p><p>Meeting a stranger this way can be extremely satisfying. You work on the same problem, explain things to each other, and eventually something works that did not work before. You have been useful to a real person. They understood what you were trying to do. After a few rounds, neither of you feels quite like a stranger anymore. <strong>This is what I want from open-source contribution. </strong></p><p>We have so many new people being enabled to fulfil their own ideas and it is renaissance of the OSS not the end of it! Bigger projects will find a way how to fix contribution inflow. <strong>But we should look at it as opportunity to find the people with the same interests.</strong> </p><p>This is roughly what I mean by helping your neighbour. It could be a dependency you already use, a small game you like, or a tool maintained by one person in a community you know. Maybe it has fifty stars. Who cares? If the maintainer has asked for help and the problem makes sense to you, your time can have a visible effect there.</p><p><strong>We are humans.</strong> We need to speak to each other and organize around the things we care about. A pile of unrelated pull requests does not create a community. People create one by returning and remembering each other. At some point they start deciding together what to do next. Small projects still have room for this.</p><p>I think most of us are looking for fairly ordinary things in the end. We want to have some fun. We want to be useful. We want the people we work with to understand why we care. This is much easier to find in a small community than in a queue where your name is one of thousands.</p><p>And if you are open source maintainer struggling with inflow of contributions, or code quality, <a href="https://reqproof.com/prove">feel free to send me a note</a> - I can help.</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Verification Gap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/p/help-your-open-source-neighbour-first?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/p/help-your-open-source-neighbour-first?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://blog.reqproof.com/p/help-your-open-source-neighbour-first?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p></div>]]></content:encoded></item><item><title><![CDATA[Good Engineering Doesn’t Trust Engineers]]></title><description><![CDATA[AI is taking over coding. To survive the identity crisis, software has to remember what engineering actually is.]]></description><link>https://blog.reqproof.com/p/good-engineering-doesnt-trust-engineers</link><guid isPermaLink="false">https://blog.reqproof.com/p/good-engineering-doesnt-trust-engineers</guid><dc:creator><![CDATA[Leonid Bugaev]]></dc:creator><pubDate>Fri, 28 Aug 2026 13:37:31 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!rRKA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0da29db5-097c-4b6d-8e93-0654d18d7771_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rRKA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0da29db5-097c-4b6d-8e93-0654d18d7771_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rRKA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0da29db5-097c-4b6d-8e93-0654d18d7771_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!rRKA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0da29db5-097c-4b6d-8e93-0654d18d7771_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!rRKA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0da29db5-097c-4b6d-8e93-0654d18d7771_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!rRKA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0da29db5-097c-4b6d-8e93-0654d18d7771_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rRKA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0da29db5-097c-4b6d-8e93-0654d18d7771_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0da29db5-097c-4b6d-8e93-0654d18d7771_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1723176,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.reqproof.com/i/213132414?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0da29db5-097c-4b6d-8e93-0654d18d7771_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!rRKA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0da29db5-097c-4b6d-8e93-0654d18d7771_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!rRKA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0da29db5-097c-4b6d-8e93-0654d18d7771_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!rRKA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0da29db5-097c-4b6d-8e93-0654d18d7771_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!rRKA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0da29db5-097c-4b6d-8e93-0654d18d7771_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>TLDR:</p><blockquote><p><strong>AI may not be the end of software engineering.</strong><br><strong>It may be the first thing forcing us to finally practise it.</strong></p></blockquote><p>I was talking with people who work in factories. Not people who write software for factories. People who spend every day next to machines, materials, operators, maintenance crews, and all the strange things that happen after a clean design meets the physical world.</p><p>They told me they do not trust software engineers.</p><p>They were not making an abstract point about education or job titles. They had seen engineers arrive with a model that worked perfectly on paper and failed to account for what actually happens on the floor. Dust gathers on a sensor. A replacement part comes from a slightly different batch. A guard is bypassed because it slows production. Heat, vibration, wear, and tired people change the conditions. A machine behaves differently on a cold morning than it did during acceptance testing.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Verification Gap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The drawing can be correct. The code can be correct. The system can still be dangerous.</p><p>My first instinct was to defend engineers. I am called one, after all.</p><p>Then I realized that good engineering agrees with the factory workers.</p><p>NASA does not solve difficult problems by finding brilliant engineers and trusting them. Aviation does not accept software because its author is experienced. Nuclear plants do not rely on the developer to remember every way a system might fail.</p><p>They build processes around the assumption that engineers can be wrong.</p><p>Requirements can be wrong. Designs can be wrong. Tests can prove the wrong thing. Entire teams can share the same blind spot. This is why regulated engineering has hazard analysis, requirements traceability, operational validation, independent verification, safety reviews, and evidence that survives after the original engineer has left. NASA even defines independence for critical software across technical, managerial, and financial dimensions. (<a href="https://swehb.nasa.gov/plugins/viewsource/viewpagesrc.action?pageId=129991590&amp;utm_source=chatgpt.com">SWEHB</a>)</p><p>Meanwhile, software engineers are having an identity crisis.</p><p>AI can write code. It can debug it, explain it, refactor it, generate tests, and review the resulting pull request. It may already be better than many of us at some bounded coding tasks, and it will not stop improving.</p><p>So software engineers are asking a frightening question:</p><blockquote><p>If AI writes the code, what exactly is my job?</p></blockquote><p>The question reveals more than we want it to.</p><p>If removing the keyboard also removes our engineering identity, perhaps coding was the identity all along.</p><h2>We copied the title</h2><p>The 1968 NATO conference that made the term <em>software engineering</em> famous was not organised to give programmers a more impressive title. It was a response to a crisis. Software systems were becoming larger, more expensive, less predictable, and harder to maintain. The term was an attempt to push software toward the discipline expected from engineering. (<a href="https://homepages.cs.ncl.ac.uk/brian.randell/NATO/?utm_source=chatgpt.com">Newcastle Uni Comp Sci Homepages</a>)</p><p>More than half a century later, we kept the word <em>engineer</em>, but much of the industry reduced the work to this:</p><pre><code><code>ticket
  &#8595;
code
  &#8595;
tests
  &#8595;
pull request
  &#8595;
green CI
  &#8595;
production</code></code></pre><p>The ticket contains a few sentences. The software engineer fills in the missing requirements while writing the implementation. The same engineer writes the tests and decides which cases matter. Someone from the same team reads the diff, usually with the same product context, the same architecture, the same deadline, and many of the same assumptions.</p><p>Then the pipeline turns green.</p><p>One small process has defined the truth, implemented the truth, and proved the truth.</p><p>We use words borrowed from engineering: architecture, design, reliability, infrastructure, incident, post-mortem. <strong>But often it is a shadow of engineering.</strong> The requirement is a Jira ticket. The safety argument is a pull-request comment. The design rationale lives in someone&#8217;s head. Validation means watching production metrics after release.</p><p>The code is treated as the only reality that matters.</p><p>This worked better when writing code was expensive. Implementation moved slowly enough that experienced people could hold a surprising amount of the system in their heads. A good developer could compensate for a weak process with memory, care, and judgment.</p><p>AI removes that protection. It can produce more changes than a human can understand line by line. It can turn one vague sentence into thousands of lines of plausible code before anyone has asked whether the sentence was true.</p><p>AI did not create software engineering&#8217;s identity crisis.</p><p>It exposed it.</p><h2>The physical world always gets a vote</h2><p>NASA&#8217;s systems-engineering process does not begin with implementation. It begins with stakeholder expectations: who needs the system, how they intend to use it, where it will operate, what constraints exist, and what success means.</p><p>Only much later is a component realised by buying it, building it, reusing it, or coding it. In NASA&#8217;s own description, coding is one implementation method inside a larger engineering process. It is not the process itself. (<a href="https://www.nasa.gov/reference/2-0-fundamentals-of-systems-engineering/?utm_source=chatgpt.com">NASA</a>)</p><p>This is also why engineering separates verification from validation.</p><p>Verification asks whether we built the product according to its requirements.</p><p>Validation asks whether we built the right product for its intended purpose and environment.</p><p>Those questions sound almost identical until a perfectly verified system fails in operation. NASA explicitly distinguishes compliance with written requirements from proving that a product accomplishes its intended purpose in its intended environment. (<a href="https://www.nasa.gov/reference/5-3-product-verification/?utm_source=chatgpt.com">NASA</a>)</p><p>You can implement every requirement correctly and still build the wrong system.</p><p>You can also write the wrong requirement, implement it perfectly, achieve complete test coverage, and congratulate yourself when every test passes.</p><p>This is where the factory worker matters. Their experience is not an annoying opinion to collect after the design is finished. It is part of the engineering input.</p><p>When an operator says, &#8220;That valve sometimes sticks after the machine has been cold all night,&#8221; the engineering response should not be, &#8220;The specification says it does not.&#8221;</p><p>The response should be:</p><blockquote><p>Our model is missing something.</p></blockquote><p>The model does not get the final vote. The physical world does.</p><h2>A hazard can write a requirement</h2><p>Most software requirements begin with desired behaviour:</p><blockquote><p>The user should be able to open the valve.</p></blockquote><p>Safety engineering begins with a different question:</p><blockquote><p>What happens if the valve opens at the wrong time?</p></blockquote><p>Then it keeps going.</p><p>What if it never opens? What if it opens twice? What if the sensor reading is stale? What if the command arrives three seconds late? What if the valve reports that it is closed when it is still open? What if the software works exactly as designed, but the operator misunderstands the display?</p><p>That creates a different flow:</p><pre><code><code>operational reality
        &#8595;
possible hazards
        &#8595;
safety constraints
        &#8595;
requirements
        &#8595;
system design
        &#8595;
implementation
        &#8595;
verification
        &#8595;
evidence</code></code></pre><p>NASA&#8217;s software-safety guidance says that preliminary hazard analysis identifies hazard causes and possible controls, which then become inputs to safety requirements. Its requirements go further: system hazard analyses and software safety analyses must create or identify the software requirements needed when software may cause, contribute to, mitigate, or control a hazard. (<a href="https://swehb.nasa.gov/plugins/viewsource/viewpagesrc.action?pageId=111575600&amp;utm_source=chatgpt.com">SWEHB</a>)</p><p>This starts before implementation, but it is not a one-time document exercise. New information from design, testing, operations, and failures changes the hazard model. The requirements and evidence have to change with it.</p><p>Now consider a function like this:</p><pre><code><code>func OpenValve() error</code></code></pre><p>Nothing inside that signature tells you how much evidence it deserves.</p><p>You need to know what the valve controls. You need to understand the pressure, material, temperature, timing, failure modes, operator response, maintenance history, and whether an independent mechanism can stop the flow.</p><p>Failure might mean a delayed batch.</p><p>It might mean a destroyed machine.</p><p>It might mean a dead person.</p><p>The criticality of code lives outside the code.</p><p>This is another place where ordinary software often reverses engineering. We look at the diff and decide how risky the change feels. We count files, lines, dependencies, and services touched.</p><p>Engineering starts with the consequence.</p><h2>Code is not the source of truth</h2><p>In an engineering system, the requirement is the obligation.</p><p>The design is an argument for satisfying that obligation. The code is one realisation of the design. Tests, analysis, simulations, reviews, and operational observations are different forms of evidence.</p><p>The implementation can change while the obligation remains.</p><p>You can rewrite the system in another language. You can replace one algorithm with another. You can move from custom hardware to an off-the-shelf component. You can throw away the current test suite and build a better one.</p><p>The requirement should survive all of that.</p><p>Code and tests matter, but they are replaceable artifacts. They are not the reason the system exists, and they are not the final authority on what the system should do.</p><p>Mainstream software inverted this relationship.</p><p>The code became the source of truth. Tests became an explanation of the current code. Requirements became temporary prose that started rotting as soon as the ticket was closed.</p><p>Six months later, nobody knows whether a strange condition is intentional, defensive, obsolete, or accidental. We read the implementation and try to reconstruct the decision that produced it.</p><p>That is not traceability.</p><p>It is archaeology.</p><p>This matters even more with AI. Generated code can be internally consistent and still be based on the wrong intent. Generated tests can confirm the same misunderstanding. A generated explanation can make the whole mistake sound reasonable.</p><p>No model can prove a system against an intent that was never made explicit.</p><h2>The builder is not the proof</h2><p>A normal software team often asks the developer to do all of these things:</p><ul><li><p>interpret the requirement;</p></li><li><p>decide the design;</p></li><li><p>write the implementation;</p></li><li><p>select the tests;</p></li><li><p>write those tests;</p></li><li><p>explain why the change is safe.</p></li></ul><p>A colleague then checks whether it all looks plausible.</p><p>This can be good work. It is not independent evidence.</p><p>NASA uses Independent Verification and Validation for critical software specifically to introduce a different perspective. Technical independence means the people doing the analysis were not involved in developing the system. Managerial independence lets them choose what to analyse and how. Financial independence protects the work from pressure by the development organisation. NASA&#8217;s rationale is direct: a genuinely different perspective can find subtle errors that the development team overlooks. (<a href="https://swehb.nasa.gov/plugins/viewsource/viewpagesrc.action?pageId=129991590&amp;utm_source=chatgpt.com">SWEHB</a>)</p><p>Not every billing page needs an independent verification organisation. Rigor should follow risk.</p><p>But the principle matters:</p><blockquote><p>The assumptions that created the system should not be the only assumptions used to prove it.</p></blockquote><p>This does not change merely because agents are involved.</p><p>An agent can write the code, generate the tests, review the diff, and produce a confident safety summary. A second agent can review it. A third can vote on the result.</p><p>But if all three receive the same incomplete requirement, share the same context, and optimise for the same target, their agreement may not mean much.</p><p>Three agents agreeing can be one assumption repeated three times.</p><p>Even strong coverage does not fix an upstream mistake. NASA requires 100 percent MC/DC coverage for identified safety-critical software components, meaning each condition in a decision must be shown to affect the outcome independently. That is serious evidence about the implementation. It still cannot tell us that the requirement was correct or that the system is safe in its real environment. (<a href="https://swehb.nasa.gov/plugins/viewsource/viewpagesrc.action?pageId=211386381&amp;utm_source=chatgpt.com">SWEHB</a>)</p><p>A green pipeline tells us that the checks we selected passed.</p><p>It does not tell us that we selected the right checks.</p><h2>So what is left when AI writes the code?</h2><p>This is where the software-engineering identity crisis becomes useful.</p><p>If your idea of engineering is turning tickets into code, AI is coming directly for the centre of your identity. Coding faster will not solve that. Learning one more framework will not solve it. Becoming better at prompting a model may extend the same identity for a while, but it does not answer the question.</p><p>The answer is not to prove that humans will always write smarter code.</p><p>Some code is genuinely difficult. Some parts require deep performance work, hardware knowledge, novel algorithms, or careful human judgment. But AI is getting good at many of those tasks too. Building our professional identity around the remaining areas where humans currently outperform it is a shrinking defence.</p><p>The way out is to stop treating coding as the definition of engineering.</p><p>Engineering is understanding the system before choosing the implementation. It is making intent explicit. It is finding the constraints hidden in the environment. It is asking what can go wrong before somebody discovers the answer in production.</p><p>It is deciding what must always be true, what must never happen, how severe failure would be, which uncertainty remains acceptable, and what evidence is strong enough for the risk involved.</p><p>It is also deciding where an agent can act alone and where human judgment is required.</p><p>AI can help with all of this. It can propose hazards, formalise requirements, analyse designs, generate test cases, search for counterexamples, and inspect evidence. This is not an argument that humans own reasoning and machines should only type code.</p><p>The point is that engineering is not a task owned by one kind of worker. It is a system for turning uncertain intent into explicit obligations, and explicit obligations into evidence.</p><p>The code is part of that system.</p><p>It is not the system.</p><p>If the only thing separating us from a code generator was that we personally typed the code, we were not defending engineering. We were defending a temporary monopoly on construction.</p><h2>We need to become engineers again</h2><p>The answer is not to make every software company imitate NASA.</p><p>Most software is not flight control. A change to button text does not need a hazard review, formal verification, and an independent assurance organisation. Copying every ceremony from a regulated industry would make ordinary development slower without making it meaningfully safer.</p><p>The important principle is proportionality.</p><p>A visual change may need a preview and a reviewer. An authentication change deserves stronger evidence. A destructive database operation should prove its safety conditions. Software controlling a medical device or physical machine belongs in another category entirely.</p><p>But every change should have some clear relationship between intent, implementation, and evidence.</p><p>That is the direction behind ReqProof.</p><p>A requirement should exist above the current code. It should say what must be true, under which conditions, and why. It should connect to hazards, constraints, implementation, tests, analysis, and operational evidence.</p><p>When something changes, we should be able to answer simple questions:</p><p>What obligation changed? Why did it change? Was the change made by a human or an agent? What evidence was produced? Which assumptions were challenged? What remains uncertain? Where was human approval required?</p><p>NASA&#8217;s FRET project demonstrates one part of this model. It lets engineers express requirements in structured natural language, gives those requirements precise semantics, and translates them into temporal logic for analysis. The important idea is not the syntax. It is that the requirement becomes something we can reason about, not prose that disappears after implementation. (<a href="https://ntrs.nasa.gov/citations/20220019339?utm_source=chatgpt.com">NASA Technical Reports Server</a>)</p><p>ReqProof takes that idea into an agent-driven software lifecycle.</p><p>The code may be generated. The tests may be generated. Parts of the analysis may be generated. But the obligation remains visible, and the evidence stays attached to the change.</p><p>We should not need to trust that an agent understood the requirement.</p><p>We should not need to trust that the developer remembered every constraint.</p><p>We should be able to inspect the argument.</p><h2>Engineering starts where trust ends</h2><p>So, are software engineers real engineers?</p><p>Some are. Some are programmers with a more expensive title. The same person may be doing engineering on one project and simply implementing features on another.</p><p>The language does not decide it. The material does not decide it. The complexity of the code does not decide it.</p><p>A better test is whether the organisation can explain what must be true, what can go wrong, why the design should work, and what evidence supports that claim&#8212;without asking us to trust the person who built it.</p><p>The factory workers were right. They should not have to trust an engineer&#8217;s clean model over what they see every day. Their knowledge should shape the requirements. Their experience should change the hazard analysis. Their objections should remain visible until someone produces evidence that they have been addressed.</p><p>Good engineering does not ask the factory floor to trust the engineer.</p><p>It gives the factory floor a way to prove the engineer wrong.</p><p>AI may not be the end of software engineering.</p><p><strong>It may be the first thing forcing us to finally practise it.</strong></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Verification Gap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[A Few Months Inside Rsync]]></title><description><![CDATA[Thirty years of code, hundreds of findings, and the pressure on the people who keep it alive]]></description><link>https://blog.reqproof.com/p/a-few-months-inside-rsync</link><guid isPermaLink="false">https://blog.reqproof.com/p/a-few-months-inside-rsync</guid><dc:creator><![CDATA[Leonid Bugaev]]></dc:creator><pubDate>Thu, 27 Aug 2026 19:24:28 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/643f2f37-d07c-43f6-a339-9fd2955b11e0_1731x909.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I was part of the team behind latest Rsync release, using <a href="https://reqproof.com/rsync">Proof</a> to perform the full continuous audit. And I wanted to share how it feels to be behind popular non commercial OSS project, and what kind of pressure and effort it takes to maintain it. <br>There is also an official <a href="https://reqproof.com/rsync">case study page</a>, if you onto it.</p><p>By our internal count, a few hundred security reports and candidate findings passed through the rsync 3.5.0 cycle. Out of them Proof formally filed 99 findings; 28 were later withdrawn. The release shipped with 33 security fixes. These are three different numbers, and the distance between them is where most of the work happened.</p><p>I have spent about ten years working in open source, so I thought I understood pressure on a mature project: the backlog that never reaches zero, behaviour nobody wants to touch, users depending on things that were never properly documented, and a very small number of people who understand why the code has the shape it does. </p><p><strong>rsync still changed my sense of scale.</strong></p><h2>Thirty years of promises</h2><p>From the outside, rsync is a command that moves files. Inside, it is thirty years of operating systems, filesystems, protocols, security assumptions and compatibility decisions compressed into one codebase.</p><p>Every flag is a promise to somebody. An option added in 2004 may still be part of a backup system that has run quietly for twenty years. A strange exit code may decide whether a production job retries. A path-handling rule may look inconsistent until you find the old platform or deployment that made it necessary.</p><p>People often use &#8220;legacy&#8221; as a synonym for bad engineering. I used to see old code that way more than I do now. <strong>Some of it is debt, obviously, but some of it is simply the price of remaining useful for thirty years.</strong> Most software disappears before it has to carry this much history.</p><p>The repository does not contain the whole project. You can read every line and still miss why a behaviour exists, which compatibility promise it protects, or which apparently cleaner implementation was already tried and rejected. A great deal of that knowledge lives with <strong>Andrew Tridgell,</strong> and other project maintainers. I do not think rsync could be maintained safely today without Andrew, or someone with an equivalent amount of project memory, and there is no quick way to manufacture such a person. rsync has been evolving from his original 1996 release ever since.</p><p>This is a different kind of engineering from building the first clean version. <strong>Keeping software alive while the world changes around it depends as much on judgment and memory as on implementation</strong>. <strong>At exactly the point where that human context has become most valuable, generating work for the people who hold it has become almost free.</strong></p><h2>There is always another report</h2><p>Maintaining a project like rsync now means knowing that somewhere in the queue there may be a serious vulnerability&#8212;and that it may look almost identical to the last ten reports that went nowhere.</p><p>A report arrives describing remote code execution, path traversal or an authentication bypass. It contains real function names and enough technical detail to sound convincing. It may have taken five minutes to generate. Deciding whether it is real can take hours.</p><p>You have to reproduce the configuration, understand who controls the input, inspect the surrounding code, and often recover the history behind the behaviour. Sometimes the report is nonsense. Sometimes it is wrong but points towards a different problem. Sometimes it is real. You cannot know without doing the work.</p><p>By our internal count, a few hundred candidate reports and potential findings passed through the rsync 3.5.0 cycle. They came from our audit, fuzzing, Trail of Bits, independent researchers, and people running models against the repository. They were not a few hundred vulnerabilities. They were a few hundred claims that experienced people had to understand well enough to classify.</p><p>This lands on top of everything else: reviewing patches, supporting old branches, fixing regressions, preparing releases and answering users. Each report pulls the maintainer into another part of a codebase where a careless decision may affect backup systems all over the world.</p><p>The asymmetry is difficult to ignore. Reports can be generated much faster than a project can investigate them. There is almost no cost to being wrong, while the cost of checking the claim falls on the people with the deepest context and the least replaceable time.</p><p>Then LLM use inside our own process became another reason to attack the project. On Discord and elsewhere, months of triage, reproductions, withdrawn findings, patches and tests were reduced to: you used LLMs, screw you, we are done with rsync.</p><p>I found that infuriating. We were not publishing raw model output as security research. The models produced claims and made mistakes. Maintainers and contributors still had to reproduce them, reject them, correct the requirements, review the fixes and take responsibility for the release.</p><p>We have made it almost free to create more work for open-source maintainers. We have not made judgment, historical knowledge or responsibility any cheaper.</p><p>And there is always another report.</p><h2>Deciding what was real</h2><p>Most of the difficult work began after a report looked plausible. Who controls the input? Which process owns the path? Does the behaviour require a non-default configuration? Is it documented? Would changing it break a workflow that has existed for fifteen years? Is it a vulnerability, a robustness problem, or behaviour that looks dangerous only when removed from its history?</p><p>One of our requirements appeared almost embarrassingly simple:</p><blockquote><p>A host matching <code>hosts deny</code> must not be admitted.</p></blockquote><p>The check failed. When a hostname in a deny rule could not be resolved, rsync skipped the rule. Under the affected configuration, the control failed open and could admit the host it was intended to block. That became CVE-2026-70452, rated HIGH.</p><p>We also found that transferred filenames could place raw control characters into rsync logs. A filename entered as data, but when an administrator later opened the log in a terminal, those bytes could become terminal escape sequences. rsync now escapes them.</p><p>Both problems look obvious after they have a name, a patch and a regression test. Before that, they sit among thousands of unusual behaviours, many of which exist for legitimate reasons.</p><p><a href="https://reqproof.com">Proof</a> formally filed 99 findings and later withdrew 28 of them. We kept the withdrawals visible because an audit has to be capable of discovering that it was wrong. Large finding counts create pressure to preserve claims after the evidence has weakened, but a claim had to survive a working reproduction, a coherent threat model, maintainer review and the history of the code. Sometimes it did not.</p><p>My role grew beyond filing reports. I joined the rsync admins group and worked on triage, tests, pull-request review, and the arguments around where a security issue ends and expected behaviour begins. </p><p>I will not pretend I was casual about seeing it there. rsync was already part of the engineering landscape when I started in open source. Being trusted inside the release mattered to me, and it also changed the feeling of the work. Once a patch lands, you think about the path you did not test. When a requirement changes, you wonder who depended on the previous behaviour.</p><p><a href="https://reqproof.com">Proof</a> helped because it gave that uncertainty somewhere to live. A model-generated claim did not become the conclusion. It entered a requirements graph connected to code, evidence and a concrete reproduction. Maintainers and contributors could challenge the assumption, correct the threat model, or show that the supposed boundary had never existed. When our understanding changed, the graph changed with it.</p><p>For each concrete requirement, it built an executable tripwire against the real rsync source tree. While a defect still reproduced, the tripwire remained live. When a fix landed, it stopped reproducing, and the accepted behaviour moved into an upstream regression test. The models still made mistakes; the important part was that those mistakes had a visible place to be corrected, and the corrected requirement could run against the next version.</p><p>Trail of Bits worked on the same release through its Patch the Planet program. They found issues we did not find, and we found issues they did not find, including some that had survived in rsync for decades and, in ten cases, since the first release in 1996. A companion fuzzing pass and other researchers found more. The 33 fixes in rsync 3.5.0 came from that combined work, not from one tool or one team.</p><h2>What the project kept</h2><p>The line in the rsync 3.5.0 release notes that matters most to me is:</p><blockquote><p>&#8220;Every fix ships with a regression test in the test suite that fails on the unfixed tree.&#8221;</p></blockquote><p>A report can disappear into an archive, and an argument about a threat model becomes difficult to reconstruct. Even Andrew cannot be expected to remember every detail forever. A test keeps part of that reasoning inside the project.</p><p>Most audits end with a document whose accuracy begins decaying as soon as the code changes. This one kept running throughout the development of 3.5.0. When upstream changed, we merged the changes, reran the tripwires, retired fixed findings, added new ones, and corrected the record.</p><p>rsync 3.5.0 shipped on August 13. The reports did not stop, and neither did the audit. The tripwires remain in place and will run against the next release candidate. But our long term plans is to configure real-time continuous audit, watch for the updates.</p><p>The project now retains more of what we learned instead of asking the same few people to remember all of it again.</p><div><hr></div><p>If you are non-commercial open-source project, reach us at <a href="https://reqproof.com">Proof</a> and we can help you with your backlog, issue triage, and ease maintainer pains. </p><p>If you are a commercial project looking to move your software quality scale in pace with modern agentic flows you are even more welcome :) </p><p>Huge kudos to all the OSS maintainers making our day to day life possible. Keep this people safe.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Verification Gap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/p/a-few-months-inside-rsync?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://blog.reqproof.com/p/a-few-months-inside-rsync?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Celebrating ten years of jsonparser by taking back the “fastest” title]]></title><description><![CDATA[From 'fastest' to 'one of the fastest' and back, with a formal proof along the way.]]></description><link>https://blog.reqproof.com/p/celebrating-ten-years-of-jsonparser</link><guid isPermaLink="false">https://blog.reqproof.com/p/celebrating-ten-years-of-jsonparser</guid><dc:creator><![CDATA[Leonid Bugaev]]></dc:creator><pubDate>Thu, 30 Jul 2026 19:19:37 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ras6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7edddec0-17ab-4fa5-85e6-167520295876_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ras6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7edddec0-17ab-4fa5-85e6-167520295876_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ras6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7edddec0-17ab-4fa5-85e6-167520295876_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!ras6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7edddec0-17ab-4fa5-85e6-167520295876_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!ras6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7edddec0-17ab-4fa5-85e6-167520295876_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!ras6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7edddec0-17ab-4fa5-85e6-167520295876_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ras6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7edddec0-17ab-4fa5-85e6-167520295876_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7edddec0-17ab-4fa5-85e6-167520295876_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2368178,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.reqproof.com/i/209145048?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7edddec0-17ab-4fa5-85e6-167520295876_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ras6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7edddec0-17ab-4fa5-85e6-167520295876_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!ras6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7edddec0-17ab-4fa5-85e6-167520295876_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!ras6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7edddec0-17ab-4fa5-85e6-167520295876_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!ras6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7edddec0-17ab-4fa5-85e6-167520295876_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p><em>Ten years ago I claimed it was the fastest JSON parser for Go. The ecosystem caught up, and &#8220;the fastest&#8221; quietly became &#8220;one of the fastest.&#8221; <strong>This summer I went back to make the claim true again</strong>, and to formally prove the library correct while I was at it. The technical appendix at the bottom has every fix, every failed experiment, and the code.</em></p><p><strong>TL;DR:</strong></p><ul><li><p><a href="https://github.com/buger/jsonparser">jsonparser</a> turned ten: 5,600+ stars, used inside Grafana, Tyk, Keybase, lux, and the Sentry, Solana, and even indirectly in Docker and Istio</p></li><li><p>Six releases in one summer: 50 open issues and 12 open PRs down to zero, 12 real bugs fixed, 25+ new APIs, zero breaking changes</p></li><li><p>It became mine first Go library with a full formal <a href="https://reqproof.com">proof</a> case: 123 formally prooved requirements, 100% MC/DC coverage, zero audit findings</p></li><li><p><strong>It is the fastest parser in the benchmark again, ahead of gjson and sonic, and still the only one that allocates nothing</strong></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Verification Gap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><p>Ten years ago, I was building a company dashboard. It needed to analyze gigabytes of JSON data (server logs, API responses, analytics events), extract insights, and visualize them in real time.</p><p>I had a choice. I could set up a proper database pipeline: ingest, transform, index, query, build aggregation layers and ETL jobs and wait minutes for each dashboard refresh. But I simply wanted real-time analytics from plain files, with no database and no ingestion pipeline to maintain. Just read the JSON, pull out the fields I needed, and show the numbers. The problem was that Go&#8217;s standard <code>encoding/json</code> was too slow for that. Deserialising gigabytes of JSON into structs or <code>map[string]interface{}</code> took too long and allocated too much memory.</p><p>Performance is a feature, which can enable completely new use-cases.</p><p>I didn&#8217;t need a full parser. I needed three fields out of a 1MB JSON document, so there was no reason to deserialize the entire thing or build a tree. If I just walked the bytes and extracted the path I wanted, I could skip 90% of the work.</p><p>So on March 20, 2016, I released <a href="https://github.com/buger/jsonparser">buger/jsonparser</a>, a path-based, zero-allocation JSON parser that worked directly on bytes. The original README made a wonderfully restrained claim:</p><blockquote><p>&#8220;Alternative JSON parser for Go (so far fastest)&#8221;</p></blockquote><p>It was 10x faster than <code>encoding/json</code> and allocated nothing. The real-time dashboard worked, reading gigabytes of JSON from plain files with no database in sight. People like to overengineer. I like plain files and fast code.</p><p>The library caught on. Over the next decade it accumulated 5,600+ stars and 450+ forks. The ecosystem moved, as ecosystems do. Generated-code parsers like easyjson and ffjson became serious competitors. Go&#8217;s own <code>encoding/json</code> improved. The claim gradually softened from &#8220;the fastest&#8221; to &#8220;one of the fastest.&#8221;</p><p>Somewhere along the way, the dashboard itself was retired. I moved on to other projects. The library kept working without me, which is both the best and the strangest fate for code you write: at some point, it stops needing you.</p><p>When I finally sat down to see who actually depends on jsonparser, I expected a few personal projects and some CLI tools. What I found it is now used by Grafana, Lux, Tyk, Keybase, Solana, Sentry, and even indirectly by projects like Docker and Itsio, and much more! A nice thought, right up until you realize it also makes my bugs other people&#8217;s problems. And I have a few public CVEs which freaked out everyones scanners (nothing THAT serious thou! - but read the small text when when you join Google OSS Fuzz program &#128517;)</p><h2>In July 2026, I went back in</h2><p>I received Anthropic&#8217;s <a href="https://claude.com/contact-sales/claude-for-oss">Claude for Open Source</a> program, which gives maintainers of qualifying projects six months of the $200-a-month Max plan for free; jsonparser qualified on stars alone. This comeback is part of what that grant bought. If you maintain something old and widely used, apply. The backlog you have been avoiding for years is suddenly a summer project!</p><p>The repository had 50 open issues and 12 open pull requests. But most important we no longer were the fastest! </p><p>For its tenth birthday, someone did.</p><p>Six releases shipped: v1.1.0 through v1.6.0. The scoreboard:</p><ul><li><p>50 open GitHub issues &#8594; zero</p></li><li><p>12 open pull requests &#8594; zero</p></li><li><p>4 formally tracked known issues &#8594; zero</p></li><li><p>12 real bugs found and fixed</p></li><li><p>25+ new backward-compatible APIs added</p></li><li><p>Zero breaking changes</p></li></ul><p><strong>And after a decade of competition, jsonparser was once again the fastest parser in the benchmark, across all three payload sizes.</strong></p><p>The scoreboard matters less than what we found on the way there, though.</p><h2>Formal verification, and what it did not magically solve</h2><p>The comeback began with putting jsonparser through <a href="https://reqproof.com/">Proof</a>, which made it the public Go library with a full L3 assurance case. Concretely, that means 123 formal requirements, full traceability across the public API, 100% Modified Condition/Decision Coverage, and an audit result of zero errors and zero warnings. The proof artifacts live with the code rather than in a detached report.</p><p>MC/DC is stronger than ordinary line or branch coverage: it demands evidence that each condition in a decision can independently flip the outcome, which is why safety-critical software relies on it.</p><p>But 100% MC/DC does not mean &#8220;there are no bugs.&#8221; We proved that distinction almost immediately.</p><p>The scalar-array corruption path was covered. Both sides of its main decision were exercised. The problem was that the wrong input category entered the wrong branch. MC/DC knew the branches were reachable; it did not know which output was semantically correct. The code ran everywhere, and nobody had ever written down what it was <em>supposed</em> to do. That is the verification gap.</p><p>That led to a <a href="https://github.com/buger/jsonparser/blob/master/docs/proof-gap-root-cause.md">blameless proof-gap postmortem</a>. Every escaped defect forced the assurance model to get stronger. For the aliasing bug, we added an explicit &#8220;must not mutate the input buffer&#8221; obligation. For the <code>EachKey</code> inconsistency, we added a cross-API consistency gate. For the benchmark mistake (more on that shame in a moment), we added a benchmark-honesty lint.</p><h2>Why ordinary fuzzing missed it</h2><p>jsonparser had already been through OSS-Fuzz. It found a real <code>Delete</code> panic. It still missed several of the bug classes above.</p><p>The reason was reachability. The fuzz harness mutated JSON bytes, but used fixed key paths like <code>"test"</code>. No amount of byte mutation can discover a panic that requires an empty path component if the path is never mutated. Likewise, a fuzzer cannot reach an out-of-range array-index mutation when it only generates ordinary object keys.</p><p>So I built <a href="https://github.com/probelabs/json-fuzz">probelabs/json-fuzz</a>, a structure-aware JSON fuzzer. It generates valid JSON from a grammar, then applies mutations at meaningful structural boundaries: after a colon, inside a Unicode escape, before a closing delimiter, or in an adversarial key path. It runs at roughly 250,000 inputs per second, and a crash is not the only failure it looks for: its gates include output validity, round trips, numeric differential tests against <code>encoding/json</code>, offset bounds, determinism, aliasing, and input preservation.</p><p>That fuzzer found bugs years of normal use and blind mutation had missed.</p><h2>The benchmark had been lying since 2017</h2><p>The funniest and most painful discovery was sitting in <a href="https://github.com/buger/jsonparser/issues/126">issue #126</a>, opened in October 2017.</p><p>The benchmark&#8217;s payload types had ffjson-generated <code>MarshalJSON</code> and <code>UnmarshalJSON</code> methods. When the supposed <code>encoding/json</code> benchmark received those types, Go correctly called their custom methods. So the &#8220;encoding/json&#8221; column was not really measuring <code>encoding/json</code>. It was measuring ffjson-generated code through the standard-library interface.</p><p>The issue was right. It stayed open for almost nine years. To the person who filed it: you were correct the entire time, and I am sorry it took a decade and a birthday to say so. <strong>However, benchmark was wrong on the other side - our numbers were even better then we originally thought!</strong> </p><p>For sure comparing with encoding/json is not fully fair, as it does not only parsing. And Go team encoding/json got 2.4x faster over 10 years.</p><p>We fixed the benchmark by introducing plain payload types, adding the competitors people would actually ask about (tidwall/gjson and ByteDance&#8217;s sonic), updating every library to its latest version, recording the hardware and Go version, and taking results as the median of five runs.</p><h2>Path to regain the fastests badge</h2><p>To get back to the &#8220;fastests&#8221; first place, I started with &#8220;<a href="https://github.com/uditakhourii/adhd">ADHD</a>&#8221; skill: a structured process that looks at the same problem through multiple independent cognitive frames (hardware engineer, competitor, speedrunner, assumption-remover, biologist), deliberately producing divergent hypotheses before converging.</p><p>Five frames generated 30 optimization ideas. Fifteen of those became isolated worktree experiments. Most did not win, which is fine. The job is to make many claims cheap to falsify, not to have one clever idea.</p><p>The biggest win came from an almost absurd piece of repeated work. <code>stringEndConfig</code> found the first quote, then searched the entire remaining parent document for a backslash. On a 24KB payload, it could walk tens of kilobytes for every string, looking for a character that wasn&#8217;t there.</p><p>Bounding that scan to the actual string body took the large benchmark from roughly 128&#181;s to 22&#181;s. A 5.8x improvement from a one-line fix.</p><p>Two more wins followed: a single 8-byte SWAR (SIMD-Within-A-Register) pass that fuses the quote and backslash scans (another eight percent), and a fast-skip trick I found by reading gjson&#8217;s source, which skips every byte above 0x5C in one unsigned comparison because all JSON structural characters sit at or below the backslash. That one bought 11 to 19 percent on the small and medium payloads, exactly where gjson had been winning. The final margin owes something to the runner-up&#8217;s own technique. Credit where it is due: tidwall writes fast code.</p><p>The final large-payload results on an Apple M4 Max with Go 1.26.3, all libraries at latest versions:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!dPY_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe14d8221-2ba5-4972-992c-32ea0b344ddf_1262x610.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!dPY_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe14d8221-2ba5-4972-992c-32ea0b344ddf_1262x610.png 424w, https://substackcdn.com/image/fetch/$s_!dPY_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe14d8221-2ba5-4972-992c-32ea0b344ddf_1262x610.png 848w, https://substackcdn.com/image/fetch/$s_!dPY_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe14d8221-2ba5-4972-992c-32ea0b344ddf_1262x610.png 1272w, https://substackcdn.com/image/fetch/$s_!dPY_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe14d8221-2ba5-4972-992c-32ea0b344ddf_1262x610.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!dPY_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe14d8221-2ba5-4972-992c-32ea0b344ddf_1262x610.png" width="1262" height="610" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e14d8221-2ba5-4972-992c-32ea0b344ddf_1262x610.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:610,&quot;width&quot;:1262,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:72827,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.reqproof.com/i/209145048?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe14d8221-2ba5-4972-992c-32ea0b344ddf_1262x610.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!dPY_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe14d8221-2ba5-4972-992c-32ea0b344ddf_1262x610.png 424w, https://substackcdn.com/image/fetch/$s_!dPY_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe14d8221-2ba5-4972-992c-32ea0b344ddf_1262x610.png 848w, https://substackcdn.com/image/fetch/$s_!dPY_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe14d8221-2ba5-4972-992c-32ea0b344ddf_1262x610.png 1272w, https://substackcdn.com/image/fetch/$s_!dPY_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe14d8221-2ba5-4972-992c-32ea0b344ddf_1262x610.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>That is 13% ahead of gjson, the other path-based parser and the obvious head-to-head comparison. It is twice as fast as ByteDance&#8217;s SIMD-accelerated sonic, and 6.5x faster than <code>encoding/json</code>, while remaining the only parser in the table that allocates nothing. The medium and small payload benchmarks land the same way, with jsonparser variants at the top of both tables; the full results are in the README.</p><p>Ten years later, the &#8220;fastest&#8221; claim came back, this time on a corrected benchmark, against the strongest competition available, with a guard against repeating the old mistake.</p><h2>Ten years of feature requests, without breaking ten years of users</h2><p>Closing the backlog did not mean mechanically closing old tickets. Many represented legitimate gaps in the API, each one somebody&#8217;s real workflow.</p><p>The highlights: wildcard paths, a JSONPath compiler, error-returning iteration, a <code>Config</code> type with opt-in lenient parsing, and a streaming <code>ReaderParser</code> for inputs of 10GB and beyond. More than 25 new APIs in total; the release notes have the full list.</p><p>The compatibility policy was non-negotiable. Six releases is aggressive enough; forcing thousands of downstream users through a migration at the same time would have been irresponsible. Docker was not going to migrate for me.</p><h2>What a comeback really means</h2><p>The repository is now at <a href="https://github.com/buger/jsonparser/releases/tag/v1.6.1">v1.6.1</a>. Zero open issues, zero open PRs, zero known issues, 123 formal requirements, zero audit findings.</p><p>&#8220;Zero known issues&#8221; does not mean &#8220;zero bugs forever.&#8221; If this session proved anything, it is that mature software can hide incorrect assumptions behind green tests, full coverage, widespread production use, and confident benchmark labels.</p><p>What changed is that jsonparser now has better ways to turn the next surprise into a permanent improvement. And it settled something for me personally: AI is what made six releases in one summer feasible for a single maintainer. The proof case is what made them safe to ship. Neither would have been enough alone.</p><p>There is a fashionable take right now that <strong>AI is killing open source</strong>. Agents strip-mine repositories for vulnerabilities, cloning a project takes an afternoon, so why maintain anything at all?</p><p>This project is my answer. The same economics that let an agent hunt for bugs in anyone&#8217;s code let me formally prove a ten-year-old library correct, in my free time, at a depth that used to require a safety-critical budget. I would not have believed that sentence a year ago.</p><blockquote><p>The win is not that AI makes the work cheap. The win is that expertise and a clear idea of what should exist are now enough to ship it.</p></blockquote><p>That does not look like open source dying to me. It looks like leverage finally landing on the maintainer&#8217;s side of the table.</p><p>In 2016, the project was an experiment in whether performance alone could replace an entire infrastructure category. Could you skip the database and just read files fast enough?</p><p>Ten years later, the answer is still yes. And the project picked up a second question along the way: how far can one maintainer take an old, widely used library if he is willing to reopen every assumption?</p><p>Including the flattering ones.</p><p>The dashboard I originally built is long gone. But the library outlived its reason to exist and grew into a hundred reasons I never planned. It is now inside more software than I can track, carries a formal proof case, and is faster than it has ever been.</p><p>So, to everyone who filed an issue over these ten years, sent a pull request, argued with me in a comment thread, or shipped jsonparser without ever saying hello: thank you. It turns out a library&#8217;s tenth birthday is less about the code than about the people, the slow and mostly invisible collaboration between one maintainer and thousands of strangers he will never meet.</p><p>That feels like a better anniversary than a cake.</p><div><hr></div><p><em>jsonparser is on <a href="https://github.com/buger/jsonparser">GitHub</a>. The formal verification tooling is <a href="https://reqproof.com/">Proof</a>. The fuzzer is <a href="https://github.com/probelabs/json-fuzz">probelabs/json-fuzz</a>.</em></p><div><hr></div><h2>Appendix: the performance story, fix by fix</h2><p>For readers who want the exact mechanics behind &#8220;the fastest claim came back,&#8221; this is the full account: three fixes that worked, ten experiments that did not, and why.</p><h3>The starting point</h3><p>The main story covers how the benchmark had been mismeasuring <code>encoding/json</code> since 2017. With that fixed, the honest starting point on an Apple M4 Max with Go 1.26.3 was roughly 128&#181;s on the large payload. Not the fastest anymore, by a wide margin.</p><h3>Fix 1: bound the backslash scan (5.8x)</h3><p><code>stringEndConfig</code> found the closing quote via <code>bytes.IndexByte(data, '"')</code>, then checked for escapes via <code>bytes.IndexByte(data, '\\')</code>. But <code>data</code> was the entire remaining parent slice, not the string body, so the escape check could walk 24,000 bytes for a single short string.</p><pre><code><code>// Before (scans entire remaining document):
firstBackslash := bytes.IndexByte(data, '\\')

// After (scans only the string body):
if bytes.IndexByte(data[:firstQuote], '\\') == -1 {
    return firstQuote + 1, false
}
</code></code></pre><p>Result: 128&#181;s to 22&#181;s. One line changed.</p><h3>Fix 2: one SWAR pass instead of two SIMD calls (8% more)</h3><p>After fix 1, the CPU profile still showed 46% of time in <code>bytes.IndexByte</code>. The function made two SIMD calls per string, one for the quote and one for the backslash, and each call carries roughly 5 to 10ns of function-call overhead before any scanning happens. The replacement is a single inline 8-byte SWAR loop that checks for both characters at once:</p><pre><code><code>const swarLsb = 0x0101010101010101
const swarMsb = 0x8080808080808080
broadcastQuote := uint64(quote) * swarLsb
broadcastBackslash := uint64('\\') * swarLsb

for i+8 &lt;= n {
    w := binary.LittleEndian.Uint64(data[i:])
    xq := w ^ broadcastQuote
    xb := w ^ broadcastBackslash
    quoteHit := (xq - swarLsb) &amp; ^xq &amp; swarMsb
    bsHit := (xb - swarLsb) &amp; ^xb &amp; swarMsb
    if quoteHit|bsHit == 0 { i += 8; continue }
    break
}
</code></code></pre><p>The SWAR formula <code>(x - 0x0101...01) &amp; ^x &amp; 0x8080...80</code> detects zero bytes (matching bytes after the XOR broadcast) in a 64-bit word using pure ALU operations. Go&#8217;s NEON <code>bytes.IndexByte</code> is faster per byte, but for the &#8220;find quote, then find backslash&#8221; pattern the call overhead dominates, and the fused loop wins. Result: 22&#181;s to 21&#181;s.</p><h3>Fix 3: gjson&#8217;s fast-skip (11 to 19% on small and medium)</h3><p>gjson was still 1.8x faster on the medium payload. Its inner loops carry a deceptively simple trick:</p><pre><code><code>for ; i &lt; len(c.json); i++ {
    if c.json[i] &gt; '\\' { continue }  // skip bytes &gt; 0x5C
    // Only structural characters reach here (all &lt;= 0x5C)
}
</code></code></pre><p>Every JSON structural character sits at or below <code>\</code> (0x5C): the quote is 0x22, the colon 0x3A, the comma 0x2C, the open bracket 0x5B, the backslash itself 0x5C. A single unsigned comparison therefore skips the vast majority of content bytes in one branch. We applied it to three hot loops: the per-byte tail of <code>stringEndConfig</code>, and the main loops of <code>blockEndConfig</code> and <code>searchKeysConfig</code>. For the latter two, <code>{</code>, <code>}</code>, and <code>]</code> are explicitly excluded from the skip, since those structural characters sit above 0x5C.</p><p>Payload Before After Improvement Small (190B) 382 ns 339 ns 11.3% Medium (2.4KB) 3,899 ns 3,141 ns 19.4% Large (24KB) 20,788 ns 20,114 ns 3.2%</p><p>Medium benefits most because it does the most sibling skipping, navigating past large nested objects to reach the target key.</p><h3>What didn&#8217;t work: ten failed experiments</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!LyJD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee8b8f63-6cce-482e-b788-0e9a4b3bd33f_1408x1838.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!LyJD!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee8b8f63-6cce-482e-b788-0e9a4b3bd33f_1408x1838.png 424w, https://substackcdn.com/image/fetch/$s_!LyJD!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee8b8f63-6cce-482e-b788-0e9a4b3bd33f_1408x1838.png 848w, https://substackcdn.com/image/fetch/$s_!LyJD!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee8b8f63-6cce-482e-b788-0e9a4b3bd33f_1408x1838.png 1272w, https://substackcdn.com/image/fetch/$s_!LyJD!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee8b8f63-6cce-482e-b788-0e9a4b3bd33f_1408x1838.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!LyJD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee8b8f63-6cce-482e-b788-0e9a4b3bd33f_1408x1838.png" width="1408" height="1838" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ee8b8f63-6cce-482e-b788-0e9a4b3bd33f_1408x1838.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1838,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:447456,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.reqproof.com/i/209145048?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee8b8f63-6cce-482e-b788-0e9a4b3bd33f_1408x1838.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!LyJD!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee8b8f63-6cce-482e-b788-0e9a4b3bd33f_1408x1838.png 424w, https://substackcdn.com/image/fetch/$s_!LyJD!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee8b8f63-6cce-482e-b788-0e9a4b3bd33f_1408x1838.png 848w, https://substackcdn.com/image/fetch/$s_!LyJD!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee8b8f63-6cce-482e-b788-0e9a4b3bd33f_1408x1838.png 1272w, https://substackcdn.com/image/fetch/$s_!LyJD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee8b8f63-6cce-482e-b788-0e9a4b3bd33f_1408x1838.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The lesson from the failures: on Apple Silicon, Go&#8217;s compiler and runtime are already extremely well optimized. NEON <code>bytes.IndexByte</code> is hand-tuned assembly, and the branch predictor chews through per-byte switches at about a cycle per byte. The wins came only from eliminating work entirely (the bounded scan), fusing operations (one SWAR pass instead of two calls), or reducing per-byte branch cost (the fast-skip). Micro-optimizations to dispatch and depth logic returned under 2% because they targeted the wrong 33%.</p><h3>The cumulative journey</h3><p>Stage Large payload What changed Before any fix ~128&#181;s <code>stringEnd</code> scanned the entire parent document Fix 1: bounded scan ~22&#181;s (5.8x) Scan only the string body Fix 2: SWAR string scan ~21&#181;s (+8%) Fuse quote and backslash into one 8-byte pass Fix 3: gjson fast-skip 20&#181;s (+5%) Skip non-structural bytes in one comparison <strong>Total</strong> <strong>128&#181;s &#8594; 20&#181;s</strong> <strong>6.4x, and the top of the leaderboard</strong></p><div><hr></div><p><em>If you read all the way past the appendix: thank you. You just finished roughly four thousand words about a JSON parser, which statistically makes you one of my people. The issue tracker is at zero right now, so if you find something, you know where to file it. I promise the response time will beat nine years.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Verification Gap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The second 100%]]></title><description><![CDATA[Coverage is the number we invented so we could stop thinking.]]></description><link>https://blog.reqproof.com/p/the-second-100</link><guid isPermaLink="false">https://blog.reqproof.com/p/the-second-100</guid><dc:creator><![CDATA[Leonid Bugaev]]></dc:creator><pubDate>Wed, 08 Jul 2026 09:41:43 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!E7tK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fa59170-77f0-4c0a-9c5c-28c830af36be_1000x640.svg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!E7tK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fa59170-77f0-4c0a-9c5c-28c830af36be_1000x640.svg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!E7tK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fa59170-77f0-4c0a-9c5c-28c830af36be_1000x640.svg 424w, https://substackcdn.com/image/fetch/$s_!E7tK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fa59170-77f0-4c0a-9c5c-28c830af36be_1000x640.svg 848w, https://substackcdn.com/image/fetch/$s_!E7tK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fa59170-77f0-4c0a-9c5c-28c830af36be_1000x640.svg 1272w, https://substackcdn.com/image/fetch/$s_!E7tK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fa59170-77f0-4c0a-9c5c-28c830af36be_1000x640.svg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!E7tK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fa59170-77f0-4c0a-9c5c-28c830af36be_1000x640.svg" width="1456" height="932" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9fa59170-77f0-4c0a-9c5c-28c830af36be_1000x640.svg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:932,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:6789,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/svg+xml&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.reqproof.com/i/206000749?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fa59170-77f0-4c0a-9c5c-28c830af36be_1000x640.svg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!E7tK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fa59170-77f0-4c0a-9c5c-28c830af36be_1000x640.svg 424w, https://substackcdn.com/image/fetch/$s_!E7tK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fa59170-77f0-4c0a-9c5c-28c830af36be_1000x640.svg 848w, https://substackcdn.com/image/fetch/$s_!E7tK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fa59170-77f0-4c0a-9c5c-28c830af36be_1000x640.svg 1272w, https://substackcdn.com/image/fetch/$s_!E7tK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fa59170-77f0-4c0a-9c5c-28c830af36be_1000x640.svg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I spent years being proud of coverage numbers. I want that time back.</p><p>Here&#8217;s a session check, close to code I&#8217;ve actually shipped:</p><pre><code><code>func (s *Store) CanUseSession(t Token, now time.Time) bool {
&#9;return t.SignatureValid() &amp;&amp;
&#9;       now.Before(t.ExpiresAt) &amp;&amp;
&#9;       !s.nonces.Seen(t.Nonce)
}</code></code></pre><p>And here&#8217;s the test I&#8217;ve written for it a hundred times:</p><pre><code><code>func TestCanUseSession(t *testing.T) {
&#9;store := NewStore()
&#9;if !store.CanUseSession(validToken(), time.Now()) {
&#9;&#9;t.Fatal("valid session rejected")
&#9;}
}</code></code></pre><p>100% line coverage. Every condition executed. The report is green.</p><p>Now delete the nonce check. Test passes. Delete the expiration check. Test passes. Replace the whole body with <code>return true</code>. Test passes. This test can&#8217;t tell the difference between my function and no function. All it proves is that good input produces a good answer &#8212; which is also what code with no security checks does.</p><p>Branch coverage is the same lie with more steps. Add the case everyone adds &#8212; a tampered token gets rejected &#8212; and both branches are hit, the report looks even better, and two of the three security checks can still be deleted without a single test noticing. That&#8217;s what our coverage metrics measure: whether code ran. Not whether any of it was ever forced to matter. And we hold serious meetings about whether the threshold should be 80% or 90%.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">I write The Verification Gap about testing, requirements, AI-generated code, and the missing evidence behind software we&#8217;re supposed to trust.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><h2>Avionics solved this twice</h2><p>In the 80s, avionics certification hit exactly this problem, and their answer had two parts. We took neither. I&#8217;ll get to the second part later, because it took watching this kind of failure to understand it.</p><p>Part one is MC/DC &#8212; Modified Condition/Decision Coverage. Ugly name, simple idea: every condition in a decision must independently flip the result, or you haven&#8217;t tested the decision.</p><p>For my session check, that means four cases:</p><pre><code><code>signature valid   expired   nonce seen   result
yes               no        no           allow
no                no        no           deny    // signature matters
yes               yes       no           deny    // expiration matters
yes               no        yes          deny    // nonce matters
</code></code></pre><p>So let&#8217;s do it properly this time:</p><pre><code><code>func TestCanUseSession(t *testing.T) {
&#9;store := NewStore()
&#9;now := time.Now()

&#9;if !store.CanUseSession(validToken(), now) {
&#9;&#9;t.Fatal("valid session rejected")
&#9;}
&#9;if store.CanUseSession(tamperedToken(), now) {
&#9;&#9;t.Fatal("bad signature accepted")
&#9;}
&#9;if store.CanUseSession(expiredToken(), now) {
&#9;&#9;t.Fatal("expired token accepted")
&#9;}

&#9;replayed := validToken()
&#9;store.CanUseSession(replayed, now) // first use: allowed
&#9;if store.CanUseSession(replayed, now) {
&#9;&#9;t.Fatal("replayed token accepted")
&#9;}
}</code></code></pre><p>100% MC/DC. Every condition forced to matter. Delete the expiration check and a test fails. Flip the nonce logic and a test fails. This is a <em>good</em> test &#8212; I&#8217;d approve it in review today. It felt like the ceiling of rigor.</p><p>Remember that feeling.</p><h2>The bug that still ships</h2><p>That test suite is good enough to make the function feel done. The code is simplified, but this is the exact class of bug I&#8217;ve watched ship: in-memory state standing in for a guarantee that needed to survive the process. Here&#8217;s the incident this design is one deploy away from.</p><p><em>s.nonces</em> is an in-memory map. Every deploy restarts the process, the map comes up empty, and every token captured before the deploy becomes replayable after it. Deploy daily, and replay protection &#8212; the thing with a dedicated, passing, well-designed test &#8212; effectively doesn&#8217;t exist for a window after every single release.</p><p>Which test failed? None. Which test <em>could</em> fail? None. Look at the replay test again: it creates a store, uses a token, replays it against the same store, in the same process, in the same millisecond. Within that world, the test is correct. The bug doesn&#8217;t live in that world. It lives in the question the test never asked: <em>what remembers the nonces, and for how long?</em></p><p>That&#8217;s the part that stings. MC/DC did its job perfectly &#8212; every condition in the decision was proven to matter. But it can only interrogate the logic I wrote. It can&#8217;t interrogate the logic I didn&#8217;t write. The bug wasn&#8217;t in a branch. It was in a question:</p><ul><li><p>What happens to seen nonces across a restart?</p></li><li><p>Across two instances behind a load balancer?</p></li><li><p>Whose clock is <code>now</code>, and how far do our servers drift?</p></li><li><p>What does the client learn from a rejection &#8212; does the error message tell an attacker <em>which</em> check failed?</p></li></ul><p>None of these are coverage questions. A suite could hit 100% of anything a tool can measure and never touch one of them. Nobody writes them down, so nobody tests them, so production tests them for us.</p><h2>A missing test is annoying. A missing question is dangerous.</h2><p>I eventually got a word for those questions: obligations.</p><p>An obligation isn&#8217;t a test case. It&#8217;s a class of behavior the change must account for. &#8220;Replayed token rejected&#8221; was a test case, and we had it. &#8220;Replay protection survives restarts and failover&#8221; was an obligation, and it existed nowhere &#8212; not in the ticket, not in the tests, not in review, only in the gap between what the ticket said and what production required.</p><p>This is why I now think of coverage as two different numbers. The first 100% is structural: did the code run, was the logic exercised. Line coverage, branch coverage, MC/DC &#8212; they all live here, and MC/DC is the honest end of it. The second 100% is semantic: did we cover what the code is <em>responsible</em> for &#8212; the requirement, the obligations, the assumptions, the risks. Every tool we have measures the first. Every incident I can remember came from the second.</p><p>Not every obligation becomes a test. Some are handled elsewhere, some are accepted risks, some need fuzzing or a runtime assertion or one honest sentence in a doc. Clock skew might be &#8220;accepted, &#177;30 seconds, here&#8217;s why&#8221; &#8212; that&#8217;s fine. What&#8217;s not fine is the obligation silently not existing, so that the difference between <em>considered and accepted</em> and <em>never thought about</em> is invisible.</p><p>And these aren&#8217;t exotic cases. Restarts, failover, cache staleness, error messages that say too much &#8212; they&#8217;re Tuesday. If nobody writes them down, they live in memory: the author&#8217;s, the reviewer&#8217;s, maybe nobody&#8217;s.</p><h2>Tests are evidence, not truth</h2><p>This is part two of what avionics knew &#8212; the part I skipped earlier. In that world, a test that doesn&#8217;t trace to a requirement is evidence of nothing. Every test exists to support a claim; every requirement must have evidence behind it. They understood forty years ago that executing code proves nothing by itself. The industry kept the percentage and threw away both ideas that gave it meaning.</p><p>For most of my career, this was my whole model of testing:</p><pre><code><code>code -&gt; tests -&gt; CI -&gt; confidence
</code></code></pre><p>Tests were code that exercised other code. More tests, more coverage, more confidence. The replay failure above fits this model perfectly and the model never blinks &#8212; the code is exercised, CI is green, confidence is high, and the bug ships anyway.</p><p>The model I&#8217;d steal from avionics, minus the paperwork:</p><pre><code><code>requirement -&gt; obligations -&gt; evidence -&gt; confidence</code></code></pre><p>Same code, same tests, same CI &#8212; but now they sit inside an argument. The requirement says what must be true. The obligations say what must be considered. The tests are demoted from &#8220;the thing itself&#8221; to what they always actually were: evidence for specific claims.</p><p>Here&#8217;s my session check under that model:</p><pre><code><code>Requirement:
A session token is accepted only if its signature is valid,
it has not expired, and it has never been used before.

Obligations:                              Evidence:
- forged signature rejected               TestCanUseSession (MC/DC)
- expired token rejected                  TestCanUseSession (MC/DC)
- replayed token rejected, same process   TestCanUseSession (MC/DC)
- replay survives restart                 MISSING &#8212; nonce store is in-memory
- replay survives failover                MISSING
- clock skew between servers              accepted risk: &#177;30s, documented
- malformed token rejected safely         fuzz target, nightly
- rejection reveals which check failed    MISSING
</code></code></pre><p>Same function, same tests. But now the failure isn&#8217;t a surprise buried in an architecture diagram &#8212; it&#8217;s a row that says MISSING, visible in review, before the deploy. The test suite didn&#8217;t get better. The *argument* got better, and the argument is what caught it.</p><p>This changes what review means. &#8220;Do we have tests?&#8221; becomes &#8220;which claim does this test prove?&#8221; &#8220;What&#8217;s the coverage number?&#8221; becomes &#8220;which obligations have no evidence?&#8221; It&#8217;s much harder to fake an argument than a number.</p><h2>Code does not remember why</h2><p>Engineers love saying code is the source of truth, because code runs and docs rot. Half right. Code is the source of truth for what the system does. It&#8217;s a terrible source of truth for what the system was supposed to do.</p><p>Code can&#8217;t tell you whether a behaviour is intentional or accidental. Whether the weird branch exists because of a customer, a migration, or a 3am incident nobody wrote down. Whether clock skew was accepted or never considered. So every change starts with archaeology: reconstruct intent from the diff, stale tickets, Slack threads, and whoever hasn&#8217;t left yet.</p><p>This was already failing when humans wrote all the code. AI makes it fail faster, because generating plausible code is now cheap and verifying intent is not. And when the code and the tests come from the same prompt &#8212; the same incomplete understanding &#8212; they agree with each other perfectly. The replay bug above, generated fluently in seconds: in-memory nonce store, matching in-process replay test, green CI, high coverage, missing obligation still missing. That&#8217;s not verification. That&#8217;s a model grading its own homework.</p><p>Humans always did this too. AI just industrialised the polished misunderstanding.</p><h2>What the second 100% looks like</h2><p>Not MC/DC everywhere. Not a requirements document for a button color. Start where being wrong is expensive: auth, billing, deletion, migrations, tenant boundaries, anything that moves money or can&#8217;t be un-shipped.</p><p>Write the requirement in plain English. List the obligations. Attach evidence &#8212; and let MISSING be visible. That&#8217;s the entire practice. The obligations table above is not a big artifact; it&#8217;s twenty lines that would expose the failure before it became an incident, and it answers questions a coverage report can&#8217;t: Which obligations are handled? Which are accepted risks &#8212; accepted by a person, on the record? Which decisions need MC/DC instead of a happy path? What goes stale if the requirement changes?</p><p>The first 100% covers the structure of the code. The second covers its responsibility. One asks whether the code ran. The other asks whether anyone understood what the code had to protect.</p><h2>Why I&#8217;m building Proof</h2><p>This bothered me long enough that <a href="https://reqproof.com">I&#8217;m building a tool around it</a>. Requirements that survive the ticket closing. Obligations that live in the repo instead of someone&#8217;s head. Tests that know which claim they support. Code, tests, and requirements that invalidate each other when they drift.</p><p>Not compliance theatre. Not another number to gamify. An evidence chain that answers the only question that matters in a review: <strong>Why do we trust this change?</strong></p><p>Coverage can&#8217;t answer that. It can only tell you which lines ran.</p><p>The second 100% is everything your coverage tool cannot see &#8212; and after enough incidents, you learn that&#8217;s exactly where the bugs live.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/p/the-second-100?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">If this changed how you think about tests, coverage, or code review, subscribe to The Verification Gap and share this post</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/p/the-second-100?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://blog.reqproof.com/p/the-second-100?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p></div>]]></content:encoded></item><item><title><![CDATA[Bugs Are Misunderstandings Made Executable]]></title><description><![CDATA[Why software fails, how the industry fights back, and whether bug-free code is even possible.]]></description><link>https://blog.reqproof.com/p/bugs-are-misunderstandings-made-executable</link><guid isPermaLink="false">https://blog.reqproof.com/p/bugs-are-misunderstandings-made-executable</guid><dc:creator><![CDATA[Leonid Bugaev]]></dc:creator><pubDate>Tue, 30 Jun 2026 16:40:49 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f47aa4ab-98be-499f-a40d-0742d9ff670c_1200x630.svg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A software bug is often <strong>a misunderstanding that has become executable</strong>.</p><p>At the deepest level, a bug is not merely &#8220;a programmer made a typo.&#8221; It is a mismatch among four things: <strong>what people wanted</strong>, <strong>what was specified</strong>, <strong>what was implemented</strong>, and <strong>the world in which the software actually runs</strong>. Software is hard because humans speak in intentions, but computers execute exact instructions. The gap between intention and exactness is where bugs live.</p><p>Dijkstra captured one side of the problem: &#8220;Program testing can be used to show the presence of bugs, but never to show their absence.&#8221; His point was not that testing is useless; it is that testing samples behavior, while a nontrivial program may have an enormous or effectively unbounded space of possible behaviors. At the theoretical limit, computability theory gives a hard boundary: there is no general algorithm that can decide all interesting behavioral questions about arbitrary programs; Stanford&#8217;s entry on computability summarizes Turing&#8217;s proof that the halting problem is unsolvable.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Verification Gap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>1. The philosophical reason software has bugs</h2><p>Software is an attempt to turn an ambiguous, changing, social world into precise, repeatable machinery.</p><p>That creates several layers of failure.</p><p>First, <strong>the world is underspecified</strong>. A user says, &#8220;make checkout fast,&#8221; &#8220;don&#8217;t lose my data,&#8221; &#8220;show the correct price,&#8221; or &#8220;make login secure.&#8221; Each sounds clear until you ask: correct in which currency, time zone, tax jurisdiction, discount rule, retry scenario, fraud case, browser, network condition, and accessibility context?</p><p>Second, <strong>software is exact but our understanding is approximate</strong>. A tiny missing condition can matter. A human can infer &#8220;obviously don&#8217;t charge the customer twice.&#8221; A payment service needs an explicit idempotency rule, retry behavior, database constraint, observability path, and failure recovery plan.</p><p>Third, <strong>software composes abstractions that leak</strong>. An app depends on libraries, operating systems, networks, CPUs, databases, caches, browsers, cloud services, humans, and business rules. Each layer promises something. Each layer has exceptions.</p><p>Fourth, <strong>software changes faster than most engineered artifacts</strong>. A bridge is not redeployed ten times a day. A web service might be. Features, dependencies, infrastructure, data distributions, regulations, and user behavior all shift underneath it.</p><p>Fifth, <strong>bugs are partly observer-relative</strong>. The same behavior can be correct for engineering, wrong for product, acceptable for one customer, illegal in one jurisdiction, and catastrophic in a safety-critical context. A bug is not only a broken instruction; it is a violated expectation.</p><p>A useful mental model:</p><blockquote><p><strong>Bug risk &#8776; ambiguity &#215; complexity &#215; change &#215; coupling &#215; consequence &#247; feedback speed</strong></p></blockquote><p>That is not a formal equation, but it captures the industry&#8217;s lived reality: bugs grow where systems are unclear, large, changing, interconnected, high-stakes, and slow to learn from.</p><h2>2. A classification of bugs</h2><h4>1. Requirements bugs</h4><p><strong>What goes wrong:</strong> The team builds the wrong thing.</p><p><strong>Why it happens:</strong> The desired behavior was ambiguous, incomplete, misunderstood, or changed after implementation began.</p><p><strong>Example:</strong> &#8220;Cancel subscription&#8221; does not clearly specify whether refunds are prorated, immediate, delayed, or handled differently by country.</p><div><hr></div><h4>2. Domain and modeling bugs</h4><p><strong>What goes wrong:</strong> The software&#8217;s model of reality is too simple or simply wrong.</p><p><strong>Why it happens:</strong> Real-world rules are messier than the abstraction chosen by the developers.</p><p><strong>Example:</strong> Tax rules, calendar rules, currency conversion, medical workflows, legal requirements, logistics, and identity systems often contain exceptions that the original model did not capture.</p><div><hr></div><h4>3. Algorithmic and logic bugs</h4><p><strong>What goes wrong:</strong> The code computes the wrong result.</p><p><strong>Why it happens:</strong> The programmer&#8217;s reasoning was flawed, an edge case was missed, or the algorithm was implemented incorrectly.</p><p><strong>Example:</strong> Off-by-one errors, wrong rounding, incorrect sorting, missing conditions, or a branch that handles most cases but fails on one unusual input.</p><div><hr></div><h4>4. State bugs</h4><p><strong>What goes wrong:</strong> The system remembers the wrong thing, forgets something important, duplicates an action, or corrupts stored information.</p><p><strong>Why it happens:</strong> State is hard. Software must track sessions, caches, retries, database writes, background jobs, user actions, and partially completed operations.</p><p><strong>Example:</strong> A customer clicks &#8220;pay,&#8221; the network times out, the system retries, and the customer is charged twice because the payment operation was not designed to be idempotent.</p><div><hr></div><h4>5. Interface and integration bugs</h4><p><strong>What goes wrong:</strong> Two parts of the system disagree about how they are supposed to communicate.</p><p><strong>Why it happens:</strong> Different components make different assumptions about data formats, units, versions, schemas, encodings, or API behavior.</p><p><strong>Example:</strong> One service sends a timestamp in milliseconds, while another service interprets it as seconds.</p><div><hr></div><h4>6. Data bugs</h4><p><strong>What goes wrong:</strong> The code may be logically correct, but the data it receives is wrong, incomplete, stale, duplicated, malformed, or unexpected.</p><p><strong>Why it happens:</strong> Real production data is often messier than test data. It may contain null values, old records, migration artifacts, inconsistent formats, or historical exceptions.</p><p><strong>Example:</strong> A report works perfectly in testing but crashes in production because one old customer record is missing a field that newer records always have.</p><div><hr></div><h4>7. Concurrency bugs</h4><p><strong>What goes wrong:</strong> The software behaves differently depending on timing.</p><p><strong>Why it happens:</strong> Multiple operations happen at the same time and interact in unexpected ways.</p><p><strong>Example:</strong> Two users try to buy the last ticket at the same moment, and both transactions appear to succeed because the system did not correctly lock or coordinate access to shared state.</p><div><hr></div><h4>8. Distributed-systems bugs</h4><p><strong>What goes wrong:</strong> Different machines, services, or regions develop different views of reality.</p><p><strong>Why it happens:</strong> Networks are unreliable. Messages can be delayed, duplicated, reordered, dropped, or retried. Systems can be temporarily unavailable while other parts continue running.</p><p><strong>Example:</strong> Two services both believe they own the same task because a coordination message was delayed or lost.</p><div><hr></div><h4>9. Resource and performance bugs</h4><p><strong>What goes wrong:</strong> The application works under normal conditions but fails under pressure.</p><p><strong>Why it happens:</strong> The system consumes too much memory, CPU, database capacity, network bandwidth, or time. Small inefficiencies become serious at scale.</p><p><strong>Example:</strong> A page loads quickly with 100 users but times out with 100,000 because every page view triggers dozens of unnecessary database queries.</p><div><hr></div><h4>10. Configuration and deployment bugs</h4><p><strong>What goes wrong:</strong> The right code runs with the wrong settings, in the wrong environment, or against the wrong dependency.</p><p><strong>Why it happens:</strong> Modern software depends heavily on environment variables, secrets, permissions, feature flags, regions, infrastructure settings, and deployment pipelines.</p><p><strong>Example:</strong> A production service accidentally points to a staging database, or a feature flag is enabled for all users instead of a small test group.</p><div><hr></div><h4>11. Security bugs</h4><p><strong>What goes wrong:</strong> The software behaves correctly for ordinary users but becomes exploitable when used maliciously.</p><p><strong>Why it happens:</strong> The system was designed for cooperative use, but attackers search for unexpected paths through input fields, permissions, APIs, dependencies, and trust boundaries.</p><p><strong>Example:</strong> A form accepts user input and passes it directly into a database query, creating an injection vulnerability.</p><div><hr></div><h4>12. UX and human-factor bugs</h4><p><strong>What goes wrong:</strong> The user makes a mistake because the interface encourages, hides, or fails to prevent it.</p><p><strong>Why it happens:</strong> Software is not only code; it is also a conversation with the user. Bad design can make the wrong action look safe, obvious, or reversible.</p><p><strong>Example:</strong> A destructive action uses vague wording, so users click it without realizing they are permanently deleting data.</p><div><hr></div><h4>13. Toolchain and platform bugs</h4><p><strong>What goes wrong:</strong> The application fails because of something beneath or around the code.</p><p><strong>Why it happens:</strong> Software depends on compilers, runtimes, operating systems, browsers, hardware, cloud platforms, libraries, and frameworks. Those layers can also contain bugs or behave differently than expected.</p><p><strong>Example:</strong> An application works in one browser but fails in another because of a subtle difference in how they implement a web standard.</p><div><hr></div><h4>14. Process and organizational bugs</h4><p><strong>What goes wrong:</strong> The organization creates the conditions for defects.</p><p><strong>Why it happens:</strong> Bugs are not always born inside code. They can come from rushed deadlines, unclear ownership, poor communication, weak review practices, bad incentives, or fragmented teams.</p><p><strong>Example:</strong> A risky migration ships without a rollback plan because no single team clearly owns the full production impact.</p><p>Security has its own mature bug vocabulary. MITRE&#8217;s CWE is a community-developed list of software and hardware weaknesses, and the CWE Top 25 highlights common, impactful weaknesses that can guide engineering investment, SDLC changes, and architectural prevention.</p><h2>3. Why bugs keep happening</h2><p>The naive explanation is: &#8220;programmers make mistakes.&#8221;</p><p>The deeper explanation is: <strong>software development is knowledge work under uncertainty</strong>.</p><p>A programmer is not simply typing instructions. They are translating a fuzzy goal into a formal mechanism while negotiating incomplete requirements, legacy constraints, hidden dependencies, deadlines, user expectations, economic tradeoffs, and future change.</p><p>Many bugs are not failures of syntax. They are failures of <strong>shared understanding</strong>.</p><p>For example:</p><p>A requirements bug is a failed conversation.</p><p>An integration bug is a failed contract.</p><p>A concurrency bug is a failed mental model of time.</p><p>A security bug is a failed imagination of adversarial behavior.</p><p>A performance bug is a failed extrapolation from small to large.</p><p>A deployment bug is a failed connection between code and environment.</p><p>A process bug is a failed organization design.</p><p>This is why &#8220;just hire better programmers&#8221; does not solve the problem. Better programmers help, but most serious software failures emerge from the <strong>system around the programmer</strong>: unclear goals, changing context, inadequate feedback, insufficient isolation, weak testing, poor observability, and incentives that reward shipping over understanding.</p><h2>4. The industry response</h2><p>The industry&#8217;s response is not one thing. It is a layered immune system.</p><h4>Prevention: make whole classes of bugs harder or impossible</h4><p>This includes better requirements, design reviews, type systems, memory-safe languages, static analysis, linters, code review, architecture constraints, secure coding standards, threat modeling, and safer frameworks.</p><p>For security, NIST&#8217;s Secure Software Development Framework organizes practices into preparing the organization, protecting software, producing well-secured software, and responding to vulnerabilities; NIST explicitly frames SSDF as a risk-based way to reduce vulnerabilities, mitigate impact, and address root causes rather than as a mere checklist. OWASP SAMM similarly defines business functions and security practices across governance, design, implementation, verification, and operations to help organizations improve software security posture.</p><p>For safety-critical systems, industries use standards and assurance processes. In aviation, RTCA describes DO-178C as the current core document for design assurance and product assurance for airborne software, referenced by FAA guidance. NASA&#8217;s software assurance materials emphasize systematic software assurance, software safety, independent verification and validation, and formal inspections to detect and eliminate defects early in the lifecycle.</p><h4>Detection: find bugs before users do</h4><p>This includes unit tests, integration tests, end-to-end tests, property-based tests, fuzzing, static analysis, dynamic analysis, penetration testing, model checking, simulation, staging environments, chaos experiments, and formal verification.</p><p>The key point: different techniques see different bug classes. Unit tests catch local logic errors. Integration tests catch contract mismatches. Fuzzing finds weird inputs. Static analysis finds certain patterns without running the program. Formal methods can prove specific properties. Production monitoring catches what pre-release methods missed.</p><h4>Containment: assume some bugs will escape</h4><p>Modern systems often assume failure and try to reduce blast radius.</p><p>That means feature flags, canary releases, staged rollouts, circuit breakers, rate limits, graceful degradation, retries with idempotency, backups, rollback plans, isolation boundaries, and observability.</p><p>Site reliability engineering formalizes this with service-level objectives and error budgets. Google&#8217;s SRE material describes SLOs as a way to measure reliability and error budgets as a way to balance reliability work against other engineering work.</p><h4>Measurement: make software delivery visible</h4><p>The DevOps/DORA movement responds to bugs partly by measuring flow and instability. DORA&#8217;s software delivery metrics include change lead time, deployment frequency, failed deployment recovery time, change fail rate, and deployment rework rate; DORA frames these as ways to understand both throughput and instability in software delivery.</p><p>The important cultural shift is that high-performing teams do not merely ask, &#8220;How many bugs did we write?&#8221; They ask:</p><p>How quickly do we detect them?</p><p>How many users are affected?</p><p>Can we roll back safely?</p><p>Do we learn from incidents?</p><p>Did we remove a class of defect, or only patch one instance?</p><h4>Learning: treat incidents as information</h4><p>Good organizations do postmortems, root-cause analysis, vulnerability disclosure, bug bounty programs, incident reviews, and process improvements.</p><p>NASA&#8217;s own lesson summaries illustrate this mindset: its Software Engineering Handbook discusses the Mars Climate Orbiter loss as a mismatch of imperial vs. metric units and ties the lesson to stronger software assurance, requirements validation, interface verification, and data-consistency testing.</p><p>The mature industry response is therefore not &#8220;be perfect.&#8221; It is:</p><blockquote><p>Prevent what you can, detect what you cannot prevent, contain what escapes, recover quickly, and learn so the same class of failure becomes less likely.</p></blockquote><h2>5. Is bug-free software possible?</h2><p>It depends what &#8220;bug-free&#8221; means.</p><p><strong>In the absolute, everyday sense: no, not for large real-world applications.</strong></p><p>A large application cannot be proven &#8220;bug-free&#8221; in the same way a theorem can be proven, because the word &#8220;bug&#8221; depends on user expectations, changing requirements, unstated assumptions, environment behavior, business rules, legal rules, and future situations nobody has imagined yet.</p><p>Also, even mathematically, there are hard limits. General automatic verification of arbitrary program behavior runs into undecidability barriers: not every meaningful question about arbitrary programs can be mechanically decided.</p><p>But <strong>in a narrower, more precise sense: yes, sometimes</strong>.</p><p>You can build software that is bug-free <strong>relative to a formal specification, for specified properties, under stated assumptions</strong>.</p><p>That distinction matters enormously.</p><p>The seL4 project, for example, proves that the seL4 OS kernel implements its specification correctly, with computer-checked mathematical evidence; its documentation also describes proofs of functional correctness, security properties, compilation correctness on certain architectures, and the fact that proof work evolves with hardware and features. CompCert is a formally verified C compiler intended for high-assurance software; its core guarantee is that generated executable code behaves as prescribed by the semantics of the source program, although its own documentation is careful about scope, noting parts such as source-text transformation and assembling/linking that are not fully formally verified.</p><p>So the honest answer is:</p><p><strong>Perfect software is possible only after you shrink the meaning of &#8220;perfect.&#8221;</strong></p><p>You must specify:</p><p>Perfect with respect to which behavior?</p><p>For which inputs?</p><p>Under which hardware assumptions?</p><p>With which compiler/runtime assumptions?</p><p>Against accidental failure, adversarial attack, or user misunderstanding?</p><p>For how long, as the environment changes?</p><p>A formally verified system may still have a bad specification. It may correctly implement the wrong requirement. It may be secure in one threat model and vulnerable in another. It may be logically correct and still unusable. It may satisfy every stated property and still surprise users.</p><h2>6. A note on Proof</h2><p>If bugs are misunderstandings made executable, then the obvious question is: how do we catch misunderstandings before they become executable?</p><p>That question is one of the reasons I started building <strong><a href="https://reqproof.com">Proof</a></strong>.</p><p>By <a href="https://reqproof.com">Proof</a>, I do not mean a magical guarantee that software can never fail. I mean a more practical layer between requirements, code, tests, documentation, and review. Tests tell us whether selected examples worked. Proof asks whether we described the intended behavior clearly enough that a tool, a reviewer, or an AI agent can challenge it before production does.</p><p>This matters even more in the age of AI-generated software. AI can now produce code very quickly. It can also produce plausible tests, plausible documentation, and plausible explanations. But plausibility is not evidence. If the original intent is vague, AI may simply automate the misunderstanding faster.</p><p>That is the gap Proof is meant to address: preserving intent as software changes. A requirement should not be a disposable note that dies once implementation begins. It should remain connected to the code that implements it, the tests that check it, and the evidence that shows what has actually been verified.</p><p><a href="https://reqproof.com">Proof</a> does not eliminate all bugs. Nothing does. But it can move some bugs earlier &#8212; from production incidents into requirements, obligations, inconsistencies, missing edge cases, and untested assumptions. In that sense, Proof is not about perfection. It is about debugging intent before reality does it for us.</p><h2>7. The practical goal is not &#8220;zero bugs&#8221;</h2><p>For most software, &#8220;zero bugs&#8221; is not an engineering goal. It is a slogan.</p><p>A better goal is:</p><blockquote><p><strong>No catastrophic bugs, no repeated bugs, no silent bugs, no unactionable bugs, and fewer entire classes of bugs over time.</strong></p></blockquote><p>The best teams try to make bugs:</p><p>less likely through design,</p><p>less severe through isolation,</p><p>more visible through observability,</p><p>faster to fix through deployment discipline,</p><p>less repeatable through learning,</p><p>and sometimes impossible through stronger languages, constraints, and formal methods.</p><p>The deepest lesson is that software quality is not only a technical property. It is a property of a whole system: people, incentives, tools, architecture, feedback loops, and philosophy of risk.</p><p>A bug is what happens when reality finds a path through your assumptions.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Verification Gap! Subscribe for free to receive new posts and support my work!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Six things I realised after Mythos disappeared]]></title><description><![CDATA[Once I heard that Anthropic releasing Mythos 5, and it will be temporary in subscription layer only for 3 weeks, my first action was to buy one more account to squeeze as much intelligence of it as possible.]]></description><link>https://blog.reqproof.com/p/six-things-i-realised-after-mythos</link><guid isPermaLink="false">https://blog.reqproof.com/p/six-things-i-realised-after-mythos</guid><dc:creator><![CDATA[Leonid Bugaev]]></dc:creator><pubDate>Tue, 16 Jun 2026 19:15:32 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/077d9b73-1922-42b6-ade5-a8206832b1a8_2400x1260.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Once I heard that Anthropic releasing Mythos 5, and it will be temporary in subscription layer only for 3 weeks, my first action was to buy one more account to squeeze as much intelligence of it as possible. Feds removed it 3 days later, the rest is history. </p><p>Can&#8217;t say that my inner world has changed, as I naturally always try to imagine the worst case scenarios, but I defo not expected it so soon. It made me reflect on my current AI usage, both personal and business one to re-evaluate all the risks.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Verification Gap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>TLDR: </p><ul><li><p>Gains I get from Claude Code and Codex are unbelievable (LLM + tight harness), no true OSS alternative yet, and I ready to pay way more for it. </p></li><li><p>I do not need more intelligence, current level is already amazing - I need more throughput.</p></li><li><p>LLM costs for businesses killing margins, geo political risks create so much uncertainty, turbulence and even more fragmentation ahead. Those who own harness and tune products to avoid vendor lock-in will survive.</p></li><li><p>Like any constraint, tuning your products for cheaper and dumber models, makes it better, and you become way more creative. And it scales much better as well.</p></li></ul><h2><strong>1. My productivity is now limited by LLM throughput</strong></h2><p>First of all, I don&#8217;t regret buying the additional account. Actually, I already have at least five now: one on Codex, two on Claude, and also a couple of open-source or hosted models, and I pay like $700/m for it already. But that is not really the point.</p><p>My productivity is now basically measured by how much access I have to the capability of the language model. I am not limited by ideas. I am not limited by tasks. I am limited by how many high-quality model sessions I can run before I hit a limit.</p><p>As an individual person, of course, I cannot spend millions on this like a business. But I try to squeeze as much intelligence as I can from current subsidised AI market before it finish.</p><p>So yes, I got hooked to Codex and Claude Code and my output depends on them.</p><h2><strong>2. I don&#8217;t need more intelligence. I need more bandwidth.</strong></h2><p>The funny thing is that I don&#8217;t even feel I need some dramatically smarter model right now. Of course, everyone wants a better model, and if tomorrow something twice as smart appears, I will obviously try to burn it down with tokens immediately. But honestly, the current level of intelligence from Opus and Codex 5.5 is already enough for the majority of tasks I am doing.</p><p>It already does almost everything I wanted it to do. The problem is not that I am waiting for AGI to become productive. I am productive already. The problem is that I do not have enough throughput.</p><p>I am not really limited by intelligence anymore. I am limited by bandwidth, limits, price, speed, and access. If I had cheap, almost unlimited access to the current level of models, I could do a ridiculous amount of things. Not in the future, not with some theoretical next generation, but with what already exists now.</p><h2><strong>3. We got spoiled, and everything can be cut off</strong></h2><p>The second thing is that the risk is real. Everything can be cut off at any moment.</p><p>We have become spoiled by products like Claude Code and Codex. They are good in a dangerous way, because after using them it becomes painful to go back. You start expecting the model to understand the task, keep the context, follow the direction, and behave like something close to a useful engineering partner.</p><p>Then you try open-source models&#8230; In benchmarks, some of them look close to Opus, in real life, very nuanced. They have bugs, edge cases, harness not optimised for them, hard to steer in the process.</p><p>But this is also useful. They show where your system depends too much on the best model. If your workflow only works when the model is extremely smart, maybe the workflow is not good enough.</p><h2><strong>4. Weak models are annoying, but they make your system better</strong></h2><p>Limitations always enforce creativity. Weak models are annoying because they do not forgive you. They do not magically understand what you meant. They do not compensate for bad context, unclear prompts, or lazy workflow design. But this is exactly why they are useful.</p><p>When I made experiment on weaker models, they forced me to improve my own applications. Better prompting, better context, better task boundaries, better harness. A strong model can hide a lot of bad architecture. A weaker model exposes it immediately.</p><p>Of course I would prefer Opus level models, without rate limit, and very cheap, but we live in the real world. </p><p>Good news is that many business tasks tasks do not need the strongest model. If the task is clear, constrained, and repeatable, you can often use cheaper models, simpler models, or even deterministic tools. This is where the economics start to make sense.</p><p>So one lesson is simple: prepare for weaker models. Not because they are great, but because your system should not collapse without the best one.</p><h2><strong>5. If the model can disappear, the harness can disappear too</strong></h2><p>If it is so easy to cut access to a model, then it is also easy to cut access to the harness around it. Claude Code itself, Codex, OpenCode, or anything similar can change, disappear, become expensive, become limited, or move in a direction that does not work for you.</p><p>This is why depending on someone else&#8217;s harness in serious development is a mistake, especially if this is your business.</p><p>I think I made the right decision when I started building my own <a href="https://github.com/probelabs/probe">harness</a>, <a href="https://github.com/probelabs/visor">my own workflow engine</a>. Not because I want to rebuild everything for fun, although apparently I do enjoy making my life more complicated. But because I want independence.</p><p>I want to own how models are called. I want to own how context is prepared. I want to own routing, evals, constraints, tools, and the workflow itself. Even open-source harnesses can become someone else&#8217;s roadmap. That is fine for experiments. For business, I do not want my core workflow to be someone else&#8217;s pricing experiment.</p><p>Building your own harness is realistic, especially if you have a lot of real material to test on.</p><h2><strong>6. For individuals, this is a subscription. For business, it is unit economics.</strong></h2><p>For individual use, I am ready to pay even more for products like Claude Code or Codex. They are so good that even if they cost twice as much, I would probably still use them. As an top tier engineer, the economics work very well for me. If it increases my output, improves my loops, and lets me keep working, it is worth it.</p><p>But business economics are completely different. You do not want to build a business where your whole margin is eaten by the LLM. That simply does not work. Price and speed become key business components. If you can optimize your application for a model that is 10 or 100 times cheaper, and it still works well, that changes the game.</p><p>Some businesses may only become possible when this pricing vs intelligence problem is solved. Plan-based pricing is dying, and API-based pricing already changing how management looks to LLM costs.</p><p>There is also the regulatory part. Mythos was one signal. Manus and Meta was another. Governments will intervene more. Companies will want privacy and protection. Serious enterprise players will not want to depend on a model, a provider, a country, or a harness they do not control.</p><p>So the conclusion for me is very simple.</p><p>Be independent from the model. Optimize your application for weaker, not-so-smart models. Stop thinking about this as open-ended magic and start thinking about workflows.</p><p>And most importantly: own the harness.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Verification Gap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Death of Security by Obscurity]]></title><description><![CDATA[I should feel very scared right now.]]></description><link>https://blog.reqproof.com/p/death-of-security-by-obscurity</link><guid isPermaLink="false">https://blog.reqproof.com/p/death-of-security-by-obscurity</guid><dc:creator><![CDATA[Leonid Bugaev]]></dc:creator><pubDate>Thu, 28 May 2026 18:21:35 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/16ad71ee-35f2-47fb-bf83-0c69a01c7cfa_1584x672.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I should feel very scared right now. Everyone should be freaking scared. But I think we have had so many emotional events happening in the world recently that people have stopped feeling much of anything, and that is the most dangerous part of where we are. The line we used to tell ourselves &#8212; &#8220;we are not a bank, we are not NASA, we don&#8217;t need that level of security&#8221; &#8212; is no longer an option. We just haven&#8217;t felt it yet.</p><p>Try this thought experiment. Imagine your company&#8217;s source code is made public tomorrow. All of it. How would you feel? I bet most of you would be freaking scared. Not because of IP. Because of the quality. Because of the spaghetti conditions in some files. Because of the strange customer-specific branch nobody touched in three years. Because of the comment that says &#8220;// TODO: fix this before prod&#8221; still sitting in prod. Because of the auth path that &#8220;almost&#8221; works. Because of the secret that probably should have been rotated. </p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Verification Gap! Subscribe to read my journey on re-discovering software engineering craft</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>For years, a lot of teams treated security as something between a badge, a process, and a hope. Of course everyone said the right thing. &#8220;We treat security as a first-class citizen,&#8221; and so on. But in practice, unless you were in a regulated industry, security was often optional in the only sense that matters &#8212; optional in priority. You followed best practices. You ran dependency scanners. Maybe you had a penetration test every six months because customers asked for it. Maybe you had a badge in your sales deck. But you weren&#8217;t really thinking about security as part of how the product works.</p><p>Banks, Automotive, Aerospace was different. In those industries, security is existential &#8212; if a bank loses trust, the bank is dead, and if an automotive system fails the wrong way, people die. So they built heavy processes around it: requirements, reviews, evidence, traceability, release gates. All the painful stuff. For a long time it was easy for the rest of us to look at that and say: yes, but we are not a bank.</p><p>I used to think this way too. At Tyk, I work with banks, governments, and large enterprises, and I was often annoyed by how slow some of their security processes were. Every release needed another check. Every patch had to go through another team. Every dependency update could become a discussion. Sometimes it took weeks or months. From the outside it looked like bureaucracy, and a lot of it was bureaucracy. But I changed my mind on the core idea. Those industries understood something the rest of us could safely ignore for a while: security is not something you add at the end. It is part of what the system is.</p><h2>Security is a market</h2><p>This is the part most engineers do not internalise. There is a real economy of people who make money by finding vulnerabilities in software. Some sell to bug bounty programmes. Some sell to brokers. Some sell to whoever is buying. And some just use what they find directly &#8212; exfiltrate data, sell the data, blackmail companies, take systems hostage.</p><p>That market used to be expensive to enter. Finding bugs took time. Understanding a custom system took time. Building an exploit took time. So attackers focused where the return was high &#8212; WordPress, Drupal, popular CMS plugins, well-known SaaS &#8212; anything they could exploit a million times after building it once. If you were niche, you were maybe scanned, but rarely understood. That asymmetry was your moat. Nobody admitted it out loud, but it was the moat.</p><p>A few years ago at Tyk we had a slightly crazy idea: let&#8217;s find open-source Tyk users across the world and see if some of them could become paid users. The idea wasn&#8217;t crazy. The crazy part was how easy the technical side was. I wrote a scanner that could scan the public IPv4 internet in a matter of hours. <strong>The whole internet</strong>. Once you do something like that yourself, &#8220;nobody will find us&#8221; stops sounding like a serious argument.</p><p>You can reproduce a small version of this at home. Run a basic HTTP application on a fresh public IP, expose a port, and watch the logs. Within minutes you start seeing requests: WordPress paths, admin URLs, old plugin routes, random probes, exploit attempts for software you are not even running. Most of it is dumb traffic. That is exactly the point. The internet does not need to know who you are before it starts touching your system.</p><p>So scanning was already cheap. The thing that just changed is understanding.</p><h2>What AI actually changed</h2><p>AI did not invent insecure software. We were already very good at writing insecure software. What AI changed is the cost of finding the insecurity, and the cost of understanding an unfamiliar system. A model can read a codebase and ask the questions a tired team will never ask, because everyone is busy shipping the next thing. It can build a personalised exploit for an unfamiliar service in hours, sometimes minutes. It can chain three small bugs that nobody would have chained manually because the manual cost was too high.</p><p>The reason this is genuinely scary is the floor, not the ceiling. The economics have flipped to the point where a kid in a basement with a decent model can build a personalised exploit for your specific codebase, scan the whole public internet in an afternoon, find every instance of your software, and run that exploit against all of them. Nothing about that sentence requires a state actor.</p><p>And this is not theoretical. Anthropic&#8217;s Project Glasswing reports that Anthropic and around 50 partners used Claude Mythos Preview to find more than ten thousand high- or critical-severity vulnerabilities across important software in the first month, with the bottleneck shifting from finding vulnerabilities to verifying and patching them. &#65532; Cloudflare pointed the same model at more than fifty of their own repositories and found that real vulnerability research needs a harness &#8212; architecture context, narrow tasks, validation &#8212; but with that harness, it works. &#65532; Anthropic also says Mythos-class capabilities will soon exist in many AI labs. &#65532; Open source historically catches up in around six months. The gap closes; it does not stay open.</p><p>So the relevant question is not whether this exact model is public today. The question is whether your security model assumes this capability stays rare. I would not bet a company on that.</p><h2>Assume your source code is already out</h2><p>Here is the uncomfortable truth: you should stop wondering whether your source code is going to leak. You should assume it already has.</p><p>I know how that sounds. But look at what just happened to GitHub. On May 20, 2026, GitHub said it had detected and contained a compromise of an employee device involving a poisoned third-party VS Code extension. GitHub&#8217;s assessment was that GitHub-internal repositories were exfiltrated, with the attacker&#8217;s claim of around 3,800 repositories being directionally consistent with the investigation. GitHub said it had no evidence of impact to customer repositories outside its own internal ones, but it still had to rotate critical secrets and continue analysing logs. &#65532;</p><p>Think about who this happened to. This is GitHub. Owned by Microsoft. With MDM, EDR, hardened endpoints, mature processes, more security engineering than almost any company on earth. One developer. One poisoned VS Code extension. Source code gone.</p><p>And the exposure window was tiny. The Nx Console advisory says the malicious version was live in the Visual Studio Marketplace for about 18 minutes and in OpenVSX for about 36 minutes. Eighteen minutes was enough.</p><p>Now compare that to your company. If this happened to GitHub, with everything GitHub has, do you really believe nobody has done the equivalent to your team? Be honest. How many extensions did your team install last week? Do you know what any of them do? When did each one last update? Who reviewed it?</p><p>So assume it. Assume some snapshot of your source code is already out there. It does not even need to be the latest one &#8212; an older snapshot is enough to know where to look. And once an attacker has that, they can use AI to do exactly what defenders are starting to do: read everything, model the system, and build the exploit shaped specifically for you.</p><p>That changes the threat model fundamentally. Black-box testing &#8212; sending requests, observing responses, fuzzing endpoints, inferring behaviour &#8212; is already dangerous. White-box is a different category. With your code in hand, the attacker does not need to guess. They can follow your authentication logic. They can inspect your authorisation paths. They can see your tenant isolation, find the one resolver that does not validate ownership, see the timeout path nobody tested, find the internal endpoint that was &#8220;safe&#8221; because nobody knew it existed, find the retry that is not idempotent, read the comment that says &#8220;this should never happen,&#8221; and find the strange customer-specific branch that everyone forgot about.</p><p>If your answer to that scenario is &#8220;well, they probably do not know how the system works,&#8221; then the source code itself was part of your security boundary. And that boundary is gone.</p><h2>CI/CD is not plumbing anymore</h2><p>The other place this hits hard is the build system. We used to think of CI/CD as plumbing. From an attacker&#8217;s point of view, CI/CD is one of the most interesting machines in the company &#8212; it usually has source code, deployment credentials, package publishing tokens, cloud access, GitHub tokens, and secrets for half of the internal systems.</p><p>The Trivy incident in March 2026 is the most uncomfortable example because Trivy is a security tool. A trusted security scanner became the attack vector &#8212; version tags in aquasecurity/trivy-action were force-pushed to credential-stealing malware, and the action stole everything CI runners had access to. &#65532; CanisterWorm followed a similar pattern: attackers stole npm tokens from compromised pipelines and used them to publish backdoored versions across every namespace they could reach. The malicious packages ran on postinstall &#8212; install was enough. &#65532;</p><p>So zero trust now has to include code. But notice the ladder. First, don&#8217;t trust your dependencies &#8212; pin versions, quarantine updates. Then, don&#8217;t trust the actions running in your pipelines &#8212; pin those to commit SHAs, not tags. Then ask the harder question: what if the platform itself is compromised, the way GitHub just was? At that point pinning helps but is no longer a complete answer. Each step up the ladder gives the attacker less leverage, but no step makes the problem zero.</p><p>This sounds exhausting. It is. But the alternative is worse.</p><h2>The CVE model is dead</h2><p>The CVE model still matters, but it cannot be the centre of your security process anymore. For a lot of teams the hidden workflow is still: wait for the CVE, check the severity, patch by priority, hope customers update before something bad happens. That workflow was already fragile. The new world breaks it.</p><p>VulnCheck found that in the first half of 2025, 32.1% of known exploited vulnerabilities had exploitation evidence on or before the day the CVE was issued. For a large share of exploited vulnerabilities, the CVE was not the start of danger. It was already late. &#65532;</p><p>The system producing CVEs is also overloaded. NIST said CVE submissions increased 263% between 2020 and 2025. NIST enriched nearly 42,000 CVEs in 2025, more than any prior year, and still said it was not enough to keep up. In April 2026, NIST moved to a risk-based enrichment model where some CVEs are listed but not immediately enriched. &#65532;</p><p>And none of that machinery will know the weird things inside your own system. A CVE will not know that your GraphQL resolver crashes a customer&#8217;s system on one malformed input. It will not know that your retry path is unsafe when the downstream service writes but your load balancer times out first. It will not know that your PII is in the same database as everything else, and one forgotten SQL injection path could expose more than anyone expected. You need a model of your own system, not just a feed of public bugs.</p><p>The intuitive response is &#8220;patch faster.&#8221; That is not enough either. No matter how fast you patch, if your architecture and development flow do not enforce security boundaries by default &#8212; if they do not force you to think about security as you write the code &#8212; patching will not save you. Every new release becomes a new attack surface, and a personalised exploit can be ready in five minutes. I am not exaggerating.</p><p>Cloudflare made the more important point. They described teams talking about a two-hour SLA from CVE release to patch in production, but if regression testing takes a day, getting to two hours means skipping something. Their conclusion is architectural: make exploitation harder even when a bug exists, put defences in front of the application, design the system so one flaw does not give access to everything else, and make fixes deployable everywhere at once. &#65532;</p><p>That is the difference between reactive security and designed security. Reactive security tries to outrun the attacker. Designed security assumes bugs exist and limits what one bug can do.</p><h2>So how do you live in this world?</h2><p>The first thing is to actually accept where you are. You have to go through the stages &#8212; angry, scared, in denial, eventually acceptance. Most teams stop at denial. They tell themselves the old story: we are not big enough, we are not interesting enough, who would target us. That story is over.</p><p>Once you accept it, the next move is not &#8220;make everything secure.&#8221; That is too vague and usually becomes theatre. Start with the worst case. What can kill your company?  How would I leak all customer data? How would I bypass authentication? How would I bypass tenant isolation? How would I poison a release? How would I get production credentials out of CI? How would I make one customer&#8217;s system go down with a single malformed request? How would I turn a slow downstream service into a cascading outage?</p><p>This is uncomfortable, but it gives you priorities. The worst case is not the same for every company. For a bank, anything touching the money is existential. If a bank loses money, the bank is dead. For a SaaS company automating LinkedIn outreach, leaking a list of customer emails is bad but survivable. The same company being used to impersonate its users and send messages on their behalf is not survivable. That is the death of the company.</p><p>If everything is critical, nothing is critical. But some things really are: authentication, authorisation, tenant isolation, secrets, CI/CD, package publishing, admin APIs, customer data, billing data, and anything that turns one bug into many customers affected. Be honest about which of these are existential for you, and put real process around those.</p><p>By real process I do not mean bureaucracy for its own sake. I mean what banks and government suppliers actually do, and I am saying this from experience &#8212; I spent years building software that went through that kind of pen-testing, the long kind, with people who do this for a living. Pin everything that runs in your build. Make secrets short-lived. Quarantine dependency updates before they reach your main branch. Treat anything that executes code in your dev or build environment as part of the product, because it is.</p><p>And then there is the boring part, the part that works much better than people want to admit: <strong>checklists</strong>.</p><p>Checklists are not exciting. They are one of the main reasons software engineering has ever shipped anything reliable. They are why planes fly. They are why a CT scanner does not lie to a radiologist. The reason they work is not because they are clever. They work because they force you to think about things you would otherwise skip.</p><p>If you have a login system, you should be forced to think about password reset, previous password reuse, account enumeration, timing attacks, brute force, lockout, and what happens when the email provider is down. If you make an outbound HTTP call, you should be forced to think about timeouts, DNS hangs, retries, idempotency, downstream slowness, partial success, and what happens if the service receives your request but your load balancer times out before you get the response. If your Go code starts a goroutine, you should be forced to think about cancellation, ownership, leaks, blocked channels, and behaviour under load.</p><p>None of this is advanced security research. It is basic engineering. But it only happens when the process forces it to happen. The hardest bugs to find are not the ones you wrote wrong; they are the ones you never wrote at all, because you never thought about that case.</p><p>And if you decide not to handle something, fine &#8212; say it. Log it as a known issue. Put an expiration on it. Hidden gaps are the dangerous ones.</p><h2>Putting security on autopilot</h2><p>After years of pen-testing work with banks and governments, I noticed the same kinds of gaps kept coming back. Login systems that didn&#8217;t think about timing attacks. HTTP calls that didn&#8217;t handle partial success. Goroutines that leaked under load. Not exotic bugs &#8212; basic ones, in different codebases, over and over.</p><p>So I started writing them down. Open-source projects I had investigated, real incidents I had seen up close, every checklist I had built across years of audits &#8212; all of it into a catalogue. That catalogue is what Proof runs on.</p><p><a href="https://reqproof.com">Then I built automation around it</a>, because checklists in someone&#8217;s head do not scale. This is also why I care about MC/DC. Line coverage tells you a line ran. Modified Condition/Decision Coverage &#8212; required by the FAA for Level A software, where failure could be catastrophic &#8212; asks whether every logical part of a decision independently affected the outcome. &#65532; The bug is rarely &#8220;this function was never tested.&#8221; It is &#8220;this function was tested, but not when auth is false, tenant is different, feature flag is enabled, downstream state is stale, and the retry path is active.&#8221; Not the happy path. The combination nobody specified.</p><p><a href="https://reqproof.com">Proof</a> works in both directions. From spec to code: if your requirement says &#8220;user can log in,&#8221; Proof attaches sub-requirements for password reset, timing attacks, enumeration, lockout, dependency failures &#8212; and the CI check literally fails if a required item has no test attached. If you decide not to handle one of them, fine &#8212; mark it as a known issue, attach an expiry, and it stays visible instead of disappearing. From code to spec: static patterns scan for signals (HTTP client, goroutine, database call, queue), and when one is found, Proof asks the spec whether you have described what happens when the service is slow, down, or partially successful. If the spec is missing, the code is shouting at the spec.</p><p>Over time the spec stops being a static document and becomes a living source of truth &#8212; in my case, a graph of small interconnected requirements that I treat as more authoritative than the code itself, because the code can drift and the spec is the contract. Spec links to code, code to tests, tests back to spec. If one of them changes, the link becomes suspect, and you have to look again.</p><p>This does not replace human security work. It just makes the questions impossible to skip.</p><h2>Security is everyone&#8217;s problem now</h2><p>Not every company needs to copy bank-level process. That would kill a lot of teams. But every software company needs to copy the posture. Security is not something you bolt on at the end. Your moat is no longer the software you build, or the market expertise, or being first. Your moat is being something people can trust to depend on for a long time.</p><p>So assume the internet will find you. Assume your dependencies are not safe by default. Assume one developer tool can become the entry point. Assume your source code has already leaked. Assume attackers can use AI to understand your system faster than you can.</p><p>Then ask what still protects you. Do you know which parts of the system can kill the company? Can you rotate secrets quickly? Do you have tests for the behaviours that matter, not just the lines that execute? Do you know when code, tests, and specs drift apart?</p><p>Security by obscurity is what dies when attention gets cheap. That world is going away. The practical question now is simple. When someone looks closely at your system &#8212; with automation, with your source code in front of them, with AI, and with patience &#8212; what will they find?</p><p>And what will be your answer?</p><div><hr></div><p>If any of this is hitting close to home and you want to see what putting security on autopilot looks like in practice, get in touch. I'd love to hear what your team is dealing with &#8212; and show you how Proof works on real code. <a href="https://reqproof.com">reqproof.com</a></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Verification Gap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Source of truth: Code, Spec, or Requirement?]]></title><description><![CDATA[When code becomes easy to produce, the hard part is remembering what we meant.]]></description><link>https://blog.reqproof.com/p/code-spec-or-requirement</link><guid isPermaLink="false">https://blog.reqproof.com/p/code-spec-or-requirement</guid><dc:creator><![CDATA[Leonid Bugaev]]></dc:creator><pubDate>Thu, 14 May 2026 14:50:13 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!nDMH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf0a4b03-1908-4c6b-b09b-66bd6f1809b9_1376x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nDMH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf0a4b03-1908-4c6b-b09b-66bd6f1809b9_1376x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nDMH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf0a4b03-1908-4c6b-b09b-66bd6f1809b9_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!nDMH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf0a4b03-1908-4c6b-b09b-66bd6f1809b9_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!nDMH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf0a4b03-1908-4c6b-b09b-66bd6f1809b9_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!nDMH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf0a4b03-1908-4c6b-b09b-66bd6f1809b9_1376x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nDMH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf0a4b03-1908-4c6b-b09b-66bd6f1809b9_1376x768.png" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bf0a4b03-1908-4c6b-b09b-66bd6f1809b9_1376x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1473621,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.reqproof.com/i/197232948?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf0a4b03-1908-4c6b-b09b-66bd6f1809b9_1376x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nDMH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf0a4b03-1908-4c6b-b09b-66bd6f1809b9_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!nDMH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf0a4b03-1908-4c6b-b09b-66bd6f1809b9_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!nDMH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf0a4b03-1908-4c6b-b09b-66bd6f1809b9_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!nDMH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf0a4b03-1908-4c6b-b09b-66bd6f1809b9_1376x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>The code runs. The code breaks. The code is what production uses.<br><br>Specs and docs can help, but they often become stale. So we learned not to trust them too much.<br><br>Code is honest in a way documents are not. It may be wrong, but it does exactly what it does.<br><br>But I think there was another reason we trusted code so much.<br><br>For a long time, code was manual work. We wrote it ourselves. We spent time with it. We were not only typing syntax; we were thinking through the system while writing it. The thinking and the implementation were almost the same activity.<br><br>That is why code deserved this level of trust. Not because code was perfect. It was not. But because the code carried a lot of the human judgement that produced it.</p><p>With agentic coding, this becomes more complicated.</p><p>As an individual contributor, I can move incredibly fast now. I can open Claude Code or a similar tool, give it direction, and shape the project while it is being built. I may start with a rough spec or just an idea, but the real decisions often happen during implementation. I try something. I see the code. I realise the original idea was not quite right. I adjust. The agent tries another version.</p><p>This is not bad. Actually, it is one of the best parts of working this way. Implementation gives feedback. Sometimes the code teaches you what the spec should have been.</p><p>So code-first is very tempting. It keeps speed. It keeps flow. It works especially well when the person steering the agent has the whole picture in their head.</p><p>But that is also the problem.</p><p>In that case, the real source of truth is not the code and not the spec. It is the experienced person.</p><p>They know the edge cases. They know the dependencies. They know which interface is fragile. They know why some strange behaviour exists. They know when the generated code is technically correct but still wrong. They are continuously filling the gaps.</p><p>The agent is not really working from a complete spec. It is working with a human who carries the missing context.</p><p>This works while the system fits inside one brain.</p><p>But real systems do not stay like this. They grow. They get delegated. Teams split. People leave. New people join. Some engineers understand the product but not the architecture. Some understand the architecture but not the domain. Some are junior. Some are moving fast. And the agent only knows what we gave it, plus whatever it can infer.</p><p>At that point, <strong>memory becomes a bad source of truth</strong>.</p><p>We forget dependencies. We forget edge cases. We forget why something was built in a strange way. We forget which downstream system depends on a behaviour. Not because people are bad, but because we are humans.</p><p>Agents do not magically solve this. If something is not described, they will improvise. They will choose something plausible. And plausible is often enough to pass the first review.</p><p>The same is true for humans. If something is not specified, we should not expect it to behave exactly as we imagined.</p><p>This is where I think the word &#8220;spec&#8221; is not enough.</p><p>In many software teams, a spec is a temporary artifact. You write it to start the task. It helps the engineer or the agent. Then implementation happens, some things change, and the spec is effectively dead. Maybe it still exists in Notion, Linear, Jira, GitHub, or a markdown file. But nobody really trusts it six months later.</p><p>If the spec is temporary, of course code wins.</p><p>But maybe the better word is requirement.</p><p>And this is not a new idea. This is basically how regulated industries already work. In aerospace, medical, automotive, and similar places, requirements are normal. They can produce multi-hundred-page documents explaining the whole system, but usually it is not really one big document in the simple sense. It is a set of requirements, sub-requirements, interfaces, tests, evidence, and links. A graph that can be turned into a document when needed.</p><p>That difference matters.</p><p>A spec often says: here is what we want to build.</p><p>A requirement says: here is what the system must do, what it must not do, how we know it was implemented, and where the evidence is.</p><p>The requirement does not die when the task is done. The code implements it. The test proves it. The evidence is attached to it. If something changes, the requirement becomes part of the review again.</p><p>This is the part I think normal software engineering may need to borrow, but without copying all the heaviness.</p><p>Not because we suddenly want bureaucracy. I don&#8217;t. But because agentic engineering makes implementation faster, and the faster we implement, the easier it becomes to lose intent.</p><p>The dangerous drift is not always a bug. The code can work. The tests can pass. CI can be green. But the product intent may have moved a little. The domain behaviour may not be exactly right. A security assumption may be weaker. An interface may have changed in a way nobody noticed.</p><p>Everything is green, but it is green around slightly wrong intent.</p><p>I also do not think the answer is just &#8220;write a huge spec first.&#8221; That can fail too. A detailed spec can make wrong assumptions before reality pushes back. Then implementation starts, the spec is challenged, and now you have spec drift instead of code drift.</p><p>So for me the real question is not code-first or spec-first.</p><p>The real question is how we manage the drift between intent and implementation.</p><p>If the code does something different from the requirement, it should not silently become the new truth. But if the requirement is wrong, it should not block reality forever either. There should be a stop. The team should ask: is the code wrong, is the requirement wrong, or did we learn something?</p><p>Until that is resolved, something is broken in the process.</p><p>This is where traceability matters. Not as documentation for documentation&#8217;s sake, but as an invalidation mechanism.</p><p>If code changes, the related requirement and tests should become suspicious. If a test changes, the requirement should be checked. If a requirement changes, the code and tests should definitely be reviewed. The system should say: this thing changed, so these other things are no longer fully trusted.</p><p>This is also why evidence matters.</p><p>I do not trust humans. I do not trust AI either. I want the evidence.</p><p>A green CI run is useful, but evidence of what? That the code passes current tests? Or that the system still matches the original intent?</p><p>Those are not the same thing.</p><p>In larger projects, the original intent becomes a spec, then tasks, then subtasks, then test cases. Every step narrows the focus. Everyone implements their piece. Everyone tests their piece. Locally, everything can look correct. But who validates the final system against the original intent?</p><p>Usually not systematically.</p><p>This is the part I want to think more about. Maybe requirement management, in some lighter and more modern form, becomes much more important for normal software teams. Not because requirements are new, but because agentic development changes the cost of not having them.</p><p>Code is still the runtime truth. It tells us what the system does.</p><p>But requirement is the intent truth. It tells us what the system is supposed to mean.</p><p>And evidence is what connects them.</p><p>I do not have the <a href="https://reqproof.com">full answer yet</a>. I still do not know how detailed requirements should be before they become harmful. I do not know which parts should be formal and which should stay flexible. I do not know how to make traceability work inside Git, PRs, CI, and agentic coding tools without making everyone hate it.</p><p>But I think the direction is becoming clearer.</p><blockquote><p>In the old world, code was expensive to produce, so code naturally became the main asset. In the agentic world, code may become cheaper to produce, but intent becomes easier to lose.</p></blockquote><p>And if intent is the thing we can lose, maybe intent is the thing we need to manage much more seriously.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe to receive new posts and updates on my journey on building <a href="https://reqproof.com">Proof</a></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Trust Is the Bottleneck]]></title><description><![CDATA[AI can write specs, code, tests, and docs. If all of them agree on the wrong intent, green CI isn&#8217;t enough.]]></description><link>https://blog.reqproof.com/p/engineerings-ai-bottleneck-is-trust</link><guid isPermaLink="false">https://blog.reqproof.com/p/engineerings-ai-bottleneck-is-trust</guid><dc:creator><![CDATA[Leonid Bugaev]]></dc:creator><pubDate>Tue, 05 May 2026 16:01:15 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/2d387e0a-8965-470d-87b9-c8743c912f31_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Everyone is asking the same question now: if AI can help us create much more code, why aren&#8217;t engineering teams suddenly moving much faster?</p><p>I think the question is right, but the answer usually stops too early.</p><p>AI does make some things dramatically faster. MVPs are faster. Prototypes are faster. The time to validate an idea is reduced a lot. You can explore directions that previously weren&#8217;t worth the effort. This is real, and I don&#8217;t want to pretend otherwise. But creating the first version of something isn&#8217;t the same as maintaining a product, and creating more pull requests isn&#8217;t the same as creating more trusted change.</p><p>This is where the economics breaks. If your team can create ten times more pull requests, your product doesn&#8217;t automatically move ten times faster. Your company doesn&#8217;t become ten times faster. The economy doesn&#8217;t double or triple. Because the expensive part of mature engineering was never only the typing of code.</p><p><strong>The expensive part is trust.</strong></p><p>Can I trust this change? Does it match the intent? Does it break a hidden customer flow? Does it affect backwards compatibility? Are the docs updated? Are the tests proving the right thing? Did we think about security, performance, malformed input, error states, release notes, migration, support?</p><blockquote><p>A pull request doesn&#8217;t answer all of this. <br>A pull request is just something asking to be trusted.</p></blockquote><p>So I don&#8217;t think the interesting question is &#8220;can AI create more code?&#8221; It can. <em>The interesting question is: what needs to exist around the code so we can safely absorb more change?</em></p><p>If we can scale trust, we can unlock the real scaling of AI. But not by sending maintainers ten times more PRs. That only moves the bottleneck. What I want is a pull request that comes with enough context that I can actually believe it: why this change exists, what it affects, which tests prove it, which docs changed, what can break, and what still needs a human decision.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">AI made implementation cheaper, but trust is still expensive. Join me on journey to find what to trust in new, post AI world.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>  If that is the kind of engineering problem you care about, subscribe.</p><h2>If your trust model is green CI, you are in trouble</h2><p>AI isn&#8217;t going back in the box. Even if you personally didn&#8217;t join the hype train, people around you probably did. Engineers use it to write code. PMs use it to write specs. Someone asks it to validate the plan, then write the code, then write the tests, then check everything against the same plan again.</p><p>I do it too. I ask AI to help me write the spec. Then I ask it to validate the plan. Then I ask it to write the code. Then validate its own code. Then write the tests. Then check everything against the original plan. It&#8217;s tempting because it works surprisingly well. In a lot of cases, it feels almost magical.</p><p><strong>But this is exactly why it becomes dangerous.</strong></p><p>For a long time, our basic engineering trust model was something like this: write the code, write the tests, pass CI/CD, review the pull request, ship. It was never perfect, but it was much better than nothing. Green CI never meant the product was correct. It meant the code passed the checks we had.</p><p>The problem is that those checks don&#8217;t prove intent. They don&#8217;t prove the requirement was correct. They don&#8217;t prove the tests were testing the right thing. They don&#8217;t prove the documentation was complete. They don&#8217;t prove the change matched the real product behavior we needed.</p><blockquote><p>They prove that the current artifacts agreed with the current checks.</p></blockquote><p>With AI, the whole chain can be generated. The spec can be wrong, the code can follow the wrong spec, the tests can validate the wrong code, the docs can describe the wrong behavior, and CI can still be green. Everything agrees with everything, but the intent is wrong.</p><blockquote><p>That&#8217;s not trust. That&#8217;s a consistent mistake.</p></blockquote><p>This is why high coverage isn&#8217;t enough either. In my <a href="https://blog.reqproof.com/p/i-had-near-100-test-coverage-it-didnt">previous article about jsonparser</a>, the painful part wasn&#8217;t that I had no tests. I had near-100% coverage in the area that mattered. The problem was that malformed input behavior was never properly described. So the tests proved what existed, not what should have existed.</p><p><strong>You cannot test what you never described.</strong></p><p>Security makes this even less optional. For years, many teams survived with some quiet version of security by obscurity. Not officially, of course. Everyone says security matters. But in practice, a lot of software depended on nobody looking too closely, or on attackers moving slowly enough that maintainers had time to react.</p><p>That assumption is breaking. <a href="https://www.vulncheck.com/blog/state-of-exploitation-1h-2025">VulnCheck reported</a> that in the first half of 2025, 32.1% of known exploited vulnerabilities had exploitation evidence on or before the day the CVE was issued. This doesn&#8217;t mean every vulnerability becomes an exploit in hours, but it does mean the old time cushion isn&#8217;t something you can build your product around anymore.</p><p>So things that felt optional before become normal engineering requirements: malformed input, authorization boundaries, resource limits, timeout behavior, error states, data exposure, public API behavior. These aren&#8217;t enterprise extras. They&#8217;re product requirements.</p><p>This is the uncomfortable part: the trust problem is now everyone&#8217;s problem. Even if your company hasn&#8217;t &#8220;adopted AI,&#8221; your people probably have. Even if your CI is green, it may be green against the wrong intent. Even if your coverage is high, it may cover the behavior you remembered to describe, not the behavior the product actually needs.</p><p>So we need a different source of truth. Not instead of CI/CD, not instead of tests, not instead of code review. Above them. Something that says what the system is supposed to do, which obligations apply, what evidence proves them, and what becomes suspicious when something changes.</p><blockquote><p>Otherwise AI won&#8217;t only help us move faster. <br>It will help us move faster with a false feeling of safety.</p></blockquote><h2>The outside structure is not the product</h2><p>I know this problem from open source. For the last 12 years at least, I worked a lot in open source. I had my own popular open source projects, and today at Tyk we build an open source API Gateway.</p><p>Open source is hard. Not because people are bad. Usually it&#8217;s the opposite. Someone from the outside sends you a pull request. Maybe it&#8217;s a bug fix. Maybe a new feature. Maybe it&#8217;s useful. Maybe it&#8217;s technically correct. Maybe they spent their evening on it.</p><p>But as a maintainer, you still need to get inside the context. You need to understand what&#8217;s happening and why this person is doing it. You can be fast and accept too much, or stay picky and make people unhappy. Neither option really solves the trust problem.</p><p>The real issue isn&#8217;t that contributors are bad. The issue is that they see the outside structure. They see the code. Maybe they see the tests. Maybe they see the docs. But they don&#8217;t see the intent in the same way the owner of the project sees it. They don&#8217;t know all the small product promises made over the years. They don&#8217;t know which ugly thing is accidental and which ugly thing is load-bearing. They don&#8217;t know which customer flow depends on some behavior that looks strange from the outside.</p><p><em>They are not inside of this bubble.</em></p><p>And this isn&#8217;t only open source. The same thing happens inside a company. Someone from support knows the product very well. They see customer pain every day. They may even be technical enough to raise a pull request. Someone from solutions architecture can do the same. Another team can contribute to your service. AI makes all of this easier.</p><p><strong>But internal doesn&#8217;t automatically mean trusted.</strong></p><p>A support engineer may understand the product from the customer side, but not the architecture. Another team may understand code, but not the local history. AI may generate something that looks clean, but it has no real ownership unless someone gives it context and checks it.</p><p>These contributions can become shallow. Not useless. Shallow. They touch the visible layer of the system, but they aren&#8217;t backed by the deep intent of the people who own this part of the product.</p><p>We tried relaxing quality gates a few times. More people contributing sounds obviously good, especially when every company has more backlog than humans. But we had cases where a simple line, a simple fix, broke everything. We had other cases where the fix was so big in scope that it was too dangerous to move there.</p><blockquote><p>The conclusion wasn&#8217;t &#8220;only engineers can write code.&#8221;<br>The conclusion was: if you want to scale engineering, it&#8217;s always about trust.</p></blockquote><p>This is also why &#8220;move fast&#8221; changes meaning when you have customers. When you&#8217;re still searching for an MVP, you can break things and call it learning. But when customers put your product inside their infrastructure, the product is no longer fully yours. They pay you for stability, security, and predictable behavior. In a way, you give away part of the ownership.</p><p>At Tyk, this is very real. We build software used by banks, governments, and large enterprises. Quality assurance isn&#8217;t some internal slogan. It&#8217;s part of the relationship with customers. Every software has bugs; I don&#8217;t want to pretend otherwise. But the price of a bug isn&#8217;t the same everywhere. Sometimes it&#8217;s legal. Sometimes it&#8217;s regulatory. Sometimes it&#8217;s very big money. Forget even the money for a second: what if the bank goes down? It can become a national-level issue.</p><blockquote><p>Speed isn&#8217;t how quickly you can make a change. <br>Speed is how quickly you can safely absorb change.</p></blockquote><p>Lehman&#8217;s software evolution work has a phrase that fits here: &#8220;The safe rate of change per release is constrained by the process dynamics.&#8221; In the same passage, he says that as the number, size, and architectural distance of changes increase, complexity and fault rate grow more than linearly.</p><p>This sounds academic, but it matches product reality. You can move only as fast as your safety norms allow. Your current team, your architecture, your customer base, your process, your quality gates, your review culture &#8212; all of this defines your real speed.</p><p>If AI gives you more change than your trust system can absorb, you aren&#8217;t scaling engineering. You&#8217;re scaling incoming work.</p><h2>Temporary specs become archaeology</h2><p>One of the deeper problems is how we treat specifications in consumer engineering.</p><p>Most of what we call a specification is a temporary artifact. You start with all the best practices. Maybe an RFC. Then it becomes a detailed Jira ticket. Maybe later there is an ADR. There are comments in GitHub. A Slack thread. A Confluence page. A few decisions made during review because reality was different from the original assumption.</p><p>At the moment, this feels normal. This is how software gets built. But after some time, all these artifacts pile up. If you want to understand how a component works, you need to dig through history. You need to understand why it ended up in this final state. Why this was done and not that. You may be lucky and find the exact explanation. In most cases it&#8217;s lost in someone&#8217;s head.</p><blockquote><p>This is archaeology, not development.</p></blockquote><p>The bigger problem is that these artifacts are independent. The RFC isn&#8217;t connected to all the code. The Jira ticket isn&#8217;t connected to all the tests. The docs are scattered across ten pages. The final implementation isn&#8217;t connected back to the original assumptions. <strong>It&#8217;s not a graph.</strong></p><p>So we trust a person &#8212; engineer, architect, lead, PM &#8212; to hold the high-level picture in their head. We trust them to find all dependencies. We trust them to notice backwards compatibility issues. We trust them to know which docs need updating. We trust them to remember which customer flow can break.</p><p>And of course people forget. Not because they&#8217;re careless. Because this is too much context for one person to carry.</p><p>You fix a bug and break some other flow. You build a feature and forget a dependency with another service. You update two documentation pages and miss the other eight. The feature exists, but it&#8217;s unusable for one group of users. The implementation works, but not in the real production shape.</p><blockquote><p>The spec was supposed to create clarity. <br>But because it was temporary, it becomes one more historical artifact.</p></blockquote><p>This is also why spec-first isn&#8217;t enough. Spec-driven development is better than no spec. Planning before coding is obviously better than jumping into implementation. But if the spec is still treated as a temporary artifact, after a few iterations you end up in the same position, with intent chaos.</p><p>During development, the spec always changes. You start with assumptions. Researched assumptions, but still assumptions. Then implementation begins and reality appears. The architecture doesn&#8217;t work. A limitation appears. A reviewer notices a security issue. QA finds a case. A customer dependency changes the direction. And in many teams, the spec isn&#8217;t updated.</p><p>The real knowledge moves into GitHub comments, Slack messages, review threads, and people&#8217;s heads. The Jira ticket becomes stale. The implementation says one thing. The ticket says another.</p><p>Imagine someone from QA comes back from vacation and needs to test the feature. They see the ticket. They see the implementation. They have no idea what is happening. Why this? Why not that? Is this intended?</p><p>It&#8217;s so common. I bet a lot of you feel the same.</p><h2>How I can know what I don&#8217;t know?</h2><p><strong>A lot of bugs are not in the code first. They are in the missing specification.</strong></p><p>Have you actually described what should happen if the input is malformed? Have you described that this functionality must not allow SQL injection? What happens if the third-party service times out? What is the error state? Have you described authorization boundaries? Resource limits? Performance boundaries? What happens when something goes from ten requests per second in a test environment to thousands in production?</p><p>There are also more subtle cases. Concurrency. Non-deterministic behavior. Map iteration. Merge order. I&#8217;m looking at you, Go.</p><p>Have you described that the behavior should be deterministic? Did you write a test for it? Did the test prove the requirement, or did it just execute the code?</p><p>This is where checklists, obligations, processes, discipline, and all the boring stuff come in. I know people hate boring process. I hate fake process too. Documents nobody reads. Boxes people tick after the fact. Quality theatre.</p><p>But the useful version is different. An obligation is not a test case. It&#8217;s a category of behavior you are required to describe: malformed input, boundary behavior, error handling, access denied, determinism, idempotency, atomicity, nil safety, overflow safety, encoding safety.</p><blockquote><p>The obligation doesn&#8217;t tell you the answer. <br>It forces you to ask the question.</p></blockquote><p>That is why I like it. It turns &#8220;maybe someone remembers&#8221; into a deterministic process. The checklist itself is human judgment. But checking whether the spec covered the checklist can be mechanical.</p><p>This is also where AI can help without pretending to own the product. If the code uses goroutines, the system can ask where cancellation, lifecycle, and error propagation are described. If code depends on map iteration or merge logic, it can ask whether determinism or commutativity matters. If code reads time directly, it can ask whether time is part of the behavior and how this is tested. If code changes a public API, it can ask where compatibility and documentation obligations are.</p><p><strong>This isn&#8217;t AI judging architecture taste. It&#8217;s tooling surfacing missed questions.</strong></p><p>That is the &#8220;how I know what I don&#8217;t know&#8221; loop. Spec obligations force code and test evidence. Code shape can reveal missing spec questions.</p><h2>What regulated industries got right</h2><p>Consumer engineering and regulated engineering live in different worlds. Different tools. Different conferences. Different language. Some of it is archaic. Some of it is bureaucracy. I don&#8217;t want every SaaS team to become an avionics certification team.</p><p><strong>But we shouldn&#8217;t ignore what they learned.</strong></p><p>I expected to find paperwork. Annoyingly, I found a lot of things our world forgot to learn.</p><p>In aviation, automotive, medical devices, space systems, the spec isn&#8217;t treated as a temporary note. It&#8217;s a source of truth that lives together with the software. Requirements have IDs. They have layers. They are linked to documentation, tests, implementation, verification evidence. You can see blast radius. You can see what a change affects. During review, if implementation differs from the spec, the spec must be updated.</p><blockquote><p>The useful idea is not the paperwork. <br>The useful idea is that intent is durable, traceable, and connected to evidence.</p></blockquote><p><a href="https://software.nasa.gov/software/ARC-18066-1">NASA&#8217;s FRET</a> is one example of this direction. It lets users enter hierarchical system requirements in structured natural language, gives those requirements unambiguous semantics, and can show them as natural language, formal logic, diagrams, and interactive simulation.</p><p>That doesn&#8217;t mean every product team needs FRET or formal methods everywhere. It means the requirement is not just a document. It&#8217;s something you can analyze, link, verify, and keep alive.</p><p>This is where requirement management becomes interesting again for consumer engineering. Not the old heavy version copied blindly from regulated industries. Not paperwork for paperwork. But the useful part: a source of truth, cross-links, invalidation, traceability, and evidence.</p><p>Combined with everything consumer engineering learned over the years: CI/CD, pull requests, fast feedback, developer experience, automated tests, observability, docs, release automation.</p><blockquote><p>We should not throw away modern engineering. <br>We should add the missing trust layer.</p></blockquote><h2>From pull request to evidence pack</h2><p>Today a pull request usually gives me code, maybe tests, maybe a description. But it doesn&#8217;t give me the whole chain.</p><p>It doesn&#8217;t tell me the original intent. It doesn&#8217;t tell me which obligations apply. It doesn&#8217;t show the blast radius. It doesn&#8217;t show which docs changed or should have changed. It doesn&#8217;t show which specs this conflicts with. It doesn&#8217;t show what changed during implementation compared to the plan.</p><p>So the reviewer has to reconstruct all of that.</p><p>Again, archaeology.</p><p><strong>What I want instead is an evidence pack.</strong> Not enterprise theatre. Not documents for the sake of documents. A practical package that makes the change reviewable.</p><p>Here is the intent. Here are the requirements. Here are the obligations. Here are the tests that witness them. Here are the docs. Here is the blast radius. Here is how it aligns with existing specs and where we checked for conflicts. Here is what changed during implementation. Here is what still needs human judgment.</p><p>Then the pull request isn&#8217;t only code. It&#8217;s the full chain of development.</p><p>This matters for open source. It matters for support engineers contributing fixes. It matters for other internal teams. It matters for AI agents. You don&#8217;t trust the contributor blindly. You don&#8217;t trust AI blindly. You trust the evidence chain, and then you still apply human judgment where judgment is needed.</p><p>This will feel slower at first. Writing obligations is slower than writing a vague ticket. Linking tests to requirements is slower than writing random tests. Updating docs through the graph is slower than pushing a change and hoping someone remembers. But not all friction is bad.</p><p>The question is whether the friction creates trust. </p><blockquote><p>Bureaucracy gives you friction without trust. <br>Evidence gives you friction that lets more people move safely.</p></blockquote><p>If this trust exists, then AI can actually help us scale. Not by dumping more pull requests into the same review bottleneck, but by making more changes reviewable, traceable, and safe to absorb in parallel.</p><p>Without trust, maintainers become managers of incoming things. Instead of thinking about architecture, future, and vision, they review an endless stream of pull requests, fixes, and generated artifacts.</p><p><strong>That is not the scaling I want.</strong></p><h2>Why I am building Proof</h2><p>This is why I am building <a href="https://reqproof.com">Proof</a>.</p><p>I don&#8217;t want another tool whose main purpose is to create more code. We already have many of those. <strong>The problem isn&#8217;t that we can&#8217;t produce enough artifacts. The problem is that the artifacts don&#8217;t preserve intent.</strong></p><p>I want specs to stop being temporary. I want requirements to live with the software. I want obligations to force the boring questions before they become production bugs. I want code, tests, docs, and requirements to invalidate each other when they drift. I want a reviewer to see the evidence chain instead of rebuilding it from memory.</p><p>AI will make engineering faster. That part is already happening. But faster without trust is not enough.</p><p>For me, the real question is this: how can I end up in the position where it&#8217;s not just a pull request coming from someone from the outside, but a well-thought evidence pack that makes me believe I can merge it as soon as possible?</p><p>That is the scaling I care about.</p><p>Not just more code.</p><p>More trusted change.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Verification Gap! Subscribe to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[I Had Near 100% Test Coverage. It Didn’t Matter.]]></title><description><![CDATA[You cannot test for what you never described.]]></description><link>https://blog.reqproof.com/p/i-had-near-100-test-coverage-it-didnt</link><guid isPermaLink="false">https://blog.reqproof.com/p/i-had-near-100-test-coverage-it-didnt</guid><dc:creator><![CDATA[Leonid Bugaev]]></dc:creator><pubDate>Wed, 29 Apr 2026 17:06:16 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/7de09747-0619-4bc2-aca5-b1c0702247c4_1451x720.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I woke up and saw a wall of emails in my personal account. Then logged into my corporate Slack, and it was filled with Zendesk messages from customers. Everyone was looking for me.</p><p>The library I wrote, <a href="https://github.com/buger/jsonparser">jsonparser</a>, which got used by a lot of projects, got its very own public CVE. So everyone started freaking out looking at their scanners.</p><p>&#8220;That&#8217;s what the fame is,&#8221; was my first thought.</p><p>Now I remember some notifications I kept ignoring from the Google OSS Fuzz project, I signed up multiple years ago.</p><p>This lib was written in the pre-AI-agents era (so weird to say that now!). Every piece was handcrafted manually, using best practices, with full test coverage. </p><p>I checked the function which had the issue, and it literally had near 100% test coverage. But it did not matter, because the issue was in handling of malformed input data. One of the edge cases which was missed. In other words, the issue was in the specification of what this function should do and how it should behave in edge cases.</p><p>But it opened one more can of worms. I wrote this library like 6 years ago. I don&#8217;t remember anything. And my only source of truth is the code and the tests, which is rather cryptic and looks more like archaeology.</p><p>The issue is fixed now. But how do I prevent such issues happening in the future? And if 100% code coverage is not the answer, what is? And what is my source of truth?</p><p>So I started digging. And it went way deeper than I expected, and changed the way I look at software engineering forever.</p><h2><strong>Down the rabbit hole</strong></h2><p>I started thinking about what the gold standard of software quality is. My first answer was NASA. How does NASA solve these kinds of issues?</p><p>AI now produces so much code that I feel like I am losing ownership of it. Not only of the code. Of the intent.</p><p>I wanted to understand how people work when tests passing is still not enough and the price of being wrong is huge.</p><p>The surprising thing is that a lot of NASA&#8217;s work is public. Their software engineering requirements are public. FRET is public. Kind2 is public. A lot of the case studies are public. There are papers about aircraft, Mars rovers, superconducting magnets, and formal requirements that found bugs before code existed.</p><p>I started reading all of this not as an academic exercise, but because I had a very dumb practical problem: my tests were green, my coverage looked fine, and still one missed edge case was enough to create a public CVE.</p><p>Then I went deeper into automotive and aerospace. It opened a whole new world of software engineering for me. For some reason, our world of consumer software engineering and regulated software engineering in those industries almost do not intersect. Different tools, different conferences, different language. Sometimes it feels like they live in a parallel universe.</p><p>Some of it looks archaic. Some methodologies are weird.</p><p>Our engineering progressed a lot too. We got very good at moving fast and catching damage quickly. CI/CD, linters, tests, canaries, observability, rollbacks. I don&#8217;t want to pretend every SaaS product should behave like avionics certification.</p><p>But we optimized for speed. They optimized for evidence.</p><p>Their industry spends much more time asking what evidence they need before they are allowed to trust the change. Some of it is painful. Some of it is bureaucracy. But the idea underneath is not stupid: if you claim the system should behave in some way, you need a durable chain from that statement to tests, code, and evidence.</p><p>There is real proof there, but it is not the fantasy version I had in my head, where every line of every product is mathematically proven end-to-end. They prove specifications. They use model checking. They simulate models, like with Simulink, against many input/output cases. They measure structural coverage. They use formal proof where the criticality justifies it.</p><p>And they still use testing, code review, static analysis, and all the normal engineering work around it. The difference is that proof and evidence are attached to the parts where being wrong is not acceptable.</p><p>That actually made the idea useful for normal engineering.</p><p>This is a huge topic, which I will cover in future articles. But the first concrete thing I found was MC/DC. It is one of the ways safety-critical industries look at coverage, and it made standard line coverage look very weak to me.</p><p><strong>Line coverage says a line was touched at runtime. It does not say that the decision was tested.</strong></p><h2><strong>Why 90% line coverage can still mean 60% real coverage</strong></h2><p>I still use line coverage. I still look at it.</p><p>But line coverage is bullshit. You should not trust it. Not on its own.</p><p>In Go, when you run:</p><pre><code><code>go test -cover ./...</code></code></pre><p>you mostly get statement coverage. The tool tells you whether a statement executed during the test run. That&#8217;s useful. But it doesn&#8217;t tell you whether the decision was tested.</p><p>Take a tiny parser-style example:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;go&quot;,&quot;nodeId&quot;:null}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-go">func isDigit(c byte) bool {
&#9;return c &gt;= '0' &amp;&amp; c &lt;= '9'
}</code></pre></div><p>Now test it like this:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;go&quot;,&quot;nodeId&quot;:null}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-go">func TestIsDigit(t *testing.T) {
&#9;if !isDigit('5') {
&#9;&#9;t.Fatal("5 should be a digit")
&#9;}
&#9;if isDigit('x') {
&#9;&#9;t.Fatal("x should not be a digit")
&#9;}
}</code></pre></div><p>Looks fine. The line ran. The function returned true once. The function returned false once. Your coverage report can look perfect.</p><p>But what did you actually prove?</p><p>You tested <code>'5'</code>. You tested <code>'x'</code>. You didn&#8217;t prove the lower boundary. You didn&#8217;t prove that <code>'/'</code> fails because it&#8217;s before <code>'0'</code>. You didn&#8217;t prove that <code>':'</code> fails because it&#8217;s after <code>'9'</code>.</p><p>The line is covered. The boundary is not.</p><p>MC/DC stands for Modified Condition/Decision Coverage. It asks the question line coverage does not ask: did each condition independently affect the outcome?</p><p>When your code says <code>if a &amp;&amp; b</code>, line coverage tells you the <code>if</code> was hit. MC/DC asks whether <code>a</code> alone can change the result, and whether <code>b</code> alone can change the result.</p><p>For this line:</p><pre><code><code>return c &gt;= '0' &amp;&amp; c &lt;= '9'</code></code></pre><p>there are two conditions:</p><pre><code><code>c &gt;= '0'
c &lt;= '9'</code></code></pre><p>A simplified MC/DC table looks like this:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!R0um!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75668db6-dcc9-400f-b6ad-6cd14dc87f9a_1270x332.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!R0um!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75668db6-dcc9-400f-b6ad-6cd14dc87f9a_1270x332.png 424w, https://substackcdn.com/image/fetch/$s_!R0um!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75668db6-dcc9-400f-b6ad-6cd14dc87f9a_1270x332.png 848w, https://substackcdn.com/image/fetch/$s_!R0um!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75668db6-dcc9-400f-b6ad-6cd14dc87f9a_1270x332.png 1272w, https://substackcdn.com/image/fetch/$s_!R0um!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75668db6-dcc9-400f-b6ad-6cd14dc87f9a_1270x332.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!R0um!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75668db6-dcc9-400f-b6ad-6cd14dc87f9a_1270x332.png" width="1270" height="332" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/75668db6-dcc9-400f-b6ad-6cd14dc87f9a_1270x332.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:332,&quot;width&quot;:1270,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:68532,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.reqproof.com/i/195788861?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75668db6-dcc9-400f-b6ad-6cd14dc87f9a_1270x332.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!R0um!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75668db6-dcc9-400f-b6ad-6cd14dc87f9a_1270x332.png 424w, https://substackcdn.com/image/fetch/$s_!R0um!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75668db6-dcc9-400f-b6ad-6cd14dc87f9a_1270x332.png 848w, https://substackcdn.com/image/fetch/$s_!R0um!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75668db6-dcc9-400f-b6ad-6cd14dc87f9a_1270x332.png 1272w, https://substackcdn.com/image/fetch/$s_!R0um!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75668db6-dcc9-400f-b6ad-6cd14dc87f9a_1270x332.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The table is just a way to say: these are the cases that matter. This is the part ordinary coverage does not force you to say.</p><p>This used to be mostly a safety-critical tooling conversation. DO-178C requires MC/DC for the highest-criticality aviation software. The tooling was expensive, slow, and hard for normal teams to justify.</p><p>That changed. GCC 14 has <code>-fcondition-coverage</code>. Clang 18 has <code>-fcoverage-mcdc</code>. Rust is moving in the same direction with richer branch and condition coverage work, even if I would not call Rust MC/DC stable yet. Go does not have native MC/DC support, so I ended up adding code-level Go MC/DC measurement to <a href="https://reqproof.com/">Proof</a>, and we have been extending the same direction to JavaScript and TypeScript as well.</p><p>What aerospace and automotive had because they were slow and diligent is now becoming available to normal engineering teams because AI changed the economics. You don&#8217;t need a certification lab to ask a harder question about your tests. You also don&#8217;t need to apply all of this to the whole company on day one. Start with the part where wrong behavior actually hurts.</p><h2><strong>The jsonparser numbers weren&#8217;t subtle</strong></h2><p>After the CVE fix, I wanted to understand why my previous approach didn&#8217;t make this kind of missing behavior obvious enough.</p><p>So I applied the MC/DC and requirements approach to jsonparser in a later public PR: <a href="https://github.com/buger/jsonparser/pull/281">buger/jsonparser#281</a>.</p><p>Again: this PR didn&#8217;t fix the original CVE. It was the follow-up work after the CVE fix. But it was not just a paperwork exercise. The hardening pass found and fixed more real issues and removed dead code that my previous process had not made obvious.</p><p>That was the uncomfortable part for me. I started by asking: what did my tests actually prove?</p><p>On the main branch before that work, ordinary Go statement coverage was already decent:</p><div class="captioned-image-container"><figure><div class="image-link image2" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Smuc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81b395e9-4e39-4a35-9da9-44514b5463b1_1362x338.png 424w, https://substackcdn.com/image/fetch/$s_!Smuc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81b395e9-4e39-4a35-9da9-44514b5463b1_1362x338.png 848w, https://substackcdn.com/image/fetch/$s_!Smuc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81b395e9-4e39-4a35-9da9-44514b5463b1_1362x338.png 1272w, https://substackcdn.com/image/fetch/$s_!Smuc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81b395e9-4e39-4a35-9da9-44514b5463b1_1362x338.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Smuc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81b395e9-4e39-4a35-9da9-44514b5463b1_1362x338.png" width="1362" height="338" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/81b395e9-4e39-4a35-9da9-44514b5463b1_1362x338.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:338,&quot;width&quot;:1362,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:75969,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.reqproof.com/i/195788861?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81b395e9-4e39-4a35-9da9-44514b5463b1_1362x338.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Smuc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81b395e9-4e39-4a35-9da9-44514b5463b1_1362x338.png 424w, https://substackcdn.com/image/fetch/$s_!Smuc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81b395e9-4e39-4a35-9da9-44514b5463b1_1362x338.png 848w, https://substackcdn.com/image/fetch/$s_!Smuc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81b395e9-4e39-4a35-9da9-44514b5463b1_1362x338.png 1272w, https://substackcdn.com/image/fetch/$s_!Smuc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81b395e9-4e39-4a35-9da9-44514b5463b1_1362x338.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></div></figure></div><p></p><p>85.3% coverage isn&#8217;t bad. Most teams would see that and move on. But decision coverage told a different story: only 66% of decisions were fully covered, and only 69.2% of conditions were proven independently.</p><p>And the more interesting part: some functions already looked perfect by ordinary coverage.</p><p>Examples from the before state:</p><pre><code><code>parseInt                     100% statement coverage
Unescape                     100% statement coverage
decodeSingleUnicodeEscape    100% statement coverage
</code></code></pre><p>But MC/DC still found missing independent-condition evidence:</p><pre><code><code>bytes.go:21   parseInt missing proof for c &lt; '0'
escape.go:148 Unescape missing proof for len(in) &gt; 0
escape.go:47  decodeSingleUnicodeEscape missing proof for h1 == badHex
escape.go:47  decodeSingleUnicodeEscape missing proof for h2 == badHex
escape.go:47  decodeSingleUnicodeEscape missing proof for h3 == badHex
</code></code></pre><p>100% line coverage can still leave a condition unproven.</p><p>The code ran. The decision wasn&#8217;t tested.</p><h2><strong>The bug was in what I forgot to describe</strong></h2><p>Coverage does not paint the whole picture. Even MC/DC. The bug can still be in the spec.</p><p>That is what happened with jsonparser. It was a classical case: you are building something, moving forward, and not looking back. You don&#8217;t know what you don&#8217;t know. I did not think about what would happen if this edge case appeared. I think most of us do not think about it this way.</p><p>I did not have any specs driving development or anything that forced me to think about the edge cases before writing the code. So of course I did not test for them. You cannot test for what you never described.</p><p>Testing assumes the specification is correct. That is the NASA/formal-methods lesson that changed how I think about this. The hard part is not testing the implementation. The hard part is questioning the specification itself.</p><p>This is where I found two different questions that I had been mashing together.</p><p>The first question starts from my specification: if this is what I claim the system should do, which logical cases need to be witnessed?</p><p>Not the code. The intent.</p><p>NASA built an open-source tool called FRET (Formal Requirements Elicitation Tool) that lets you write requirements in structured English and translates them into formal logic.</p><p>FRET includes an algorithm called FLIP (FuLl Independence Pair). FLIP takes a formalized requirement and generates the minimum set of test cases proving each boolean variable independently affects the outcome. Not every possible combination. Just the ones that matter.</p><p>I still have to write the requirement. I still have to decide what malformed input, boundaries, errors, and edge cases mean. FLIP does not do that for me.</p><p>But once the requirement is formalized, FLIP tells me exactly which test cases that requirement needs.</p><p>I built a tool called <a href="https://reqproof.com/">Proof</a> that implements this approach.</p><p>That is the part I care about: how many tests are enough for this requirement?</p><p>Not &#8220;how many tests did I happen to write?&#8221; Enough for what I described.</p><p>The second question starts from my actual code: did my tests exercise every boolean condition in the implementation so each one independently affects the outcome?</p><p>This side does not care what I meant. It looks at what I wrote.</p><p>And sometimes it shows that my code has many more logical cases than my spec. So maybe my spec is not accurate enough.</p><p>Or my spec says this edge case matters, but my tests don&#8217;t witness it.</p><p>Or my tests cover implementation details, but the behavior is under-described.</p><p>I learned this the hard way on jsonparser. The spec side and the code side kept disagreeing in useful ways, and that is where code drift and spec drift become visible.</p><p>The gap goes in both directions. Sometimes the code is wrong. Sometimes the tests are weak. Sometimes the spec is too vague.</p><p>Sometimes all of it combined badly.</p><h2><strong>Checklists, not memory</strong></h2><p>What can be more deterministic than a checklist? In aerospace and automotive, everything has its own checklist. The price of a mistake is too high to rely on someone&#8217;s memory. I think checklists are the driving force behind quality engineering in those industries.</p><p>When you do not have specifications, it is very hard to create a checklist. When you are building a feature, you can have test cases, but that is a moving target. The items are constantly changing. You need something that will be the same all the time.</p><p>In this context, an obligation is not a test case. It is a category of behavior you are required to describe. Malformed input is an obligation. Boundary behavior is an obligation. Error handling is an obligation. For each one that applies to your requirement, you need at least one test case that proves how the system behaves in that category. The obligation does not tell you the answer. It forces you to ask the question.</p><p>You cannot rely on humans here. Even on me, to be frank. I can miss these items too. You need deterministic checklists.</p><p>In practice, the questions are very simple:</p><p>What will happen if this is malformed data? What will happen if this is slow and the request times out? What will happen if the database is down? What will happen if you have a very large object? What will happen if the function returns different values with the same inputs?</p><p>These are the cases where security issues and data bugs tend to live. For jsonparser, these are the exact cases I had not thought about.</p><p>Without obligations, edge cases depend on memory. Maybe I remember to test malformed data. Maybe the AI remembers. Maybe a reviewer notices. Maybe no one does.</p><p>At the moment, it is just a matter of whether someone forgets or not forgets to test it.</p><p>This is where the CVE fix actually changed how I work. The fix itself was mechanical. But the obligations I wrote afterward forced me to think about the cases I had skipped. Every one of those became an explicit question I had to answer. Not &#8220;did someone remember to test this?&#8221; but &#8220;here is the list, and each item needs a witness.&#8221;</p><p><strong>Obligations turn edge cases from &#8220;someone remembered to test this&#8221; into a deterministic process.</strong></p><p>When I first started writing obligations for jsonparser, it was actually quite easy with modern AI tooling. I reviewed all of the specs. The flow is: you cannot pass this check until the checklist is green, until you define obligations for all of those cases, and until you define test cases for all of those cases as well.</p><p>This is what the double link looks like in practice:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;javascript&quot;,&quot;nodeId&quot;:null}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-javascript">// In the code &#8212; annotated with the requirement it implements:
// SYS-REQ-863
func (s *Service) lookupCache(req Request) (*Result, bool) {
    // ...
}

// In the test &#8212; annotated with both the requirement AND the specific MC/DC row:
// Verifies: SYS-REQ-863
// MCDC SYS-REQ-863: cache_lookup_requested=T, component_inputs_unchanged=F,
//                    cached_component_result_reused=F =&gt; TRUE
func TestMCDC_SYS_REQ_863_Row1(t *testing.T) {
    evalVerifyScenario(t, "SYS-REQ-863", map[string]bool{
        "cache_lookup_requested":         true,
        "component_inputs_unchanged":     false,
        "cached_component_result_reused": false,
    }, true)
}</code></pre></div><p>Each test is not just &#8220;test the function.&#8221; Each test is: &#8220;prove that this specific variable independently affects the outcome of this specific requirement.&#8221;</p><p>If I change the spec, I can see exactly which MC/DC rows are affected and which tests need to be reviewed. If I change a test, I can see which spec requirement it was proving and check whether the spec still says the same thing. If I add a new variable to the requirement, FLIP will generate new witness rows, and the missing tests become immediately visible.</p><p>This is the double link. Change the spec, review the tests. Change the tests, review the spec. If you have not touched the spec, why would you touch the test?</p><p>This is where the &#8220;how many tests are enough?&#8221; question changed for me. Before, the answer was always vibes. Write enough tests. Cover important paths. Don&#8217;t overdo it. Be pragmatic.</p><p>All true, and also not very helpful.</p><p>Now I think about it differently. Enough tests means enough evidence that every condition I described, or every condition my code actually contains, can independently affect the behavior I care about.</p><p>It is not about how many tests I have. It is about whether I really, really trust my system and whether it actually does what I described.</p><h2><strong>The true challenge is legacy</strong></h2><p>You can always start a new project and have a really nice experience with all of this. But the true challenge lies in the big legacy projects. They make up like 90% of all software. They bring the majority of the money. And they are the ones where wrong behavior actually hurts.</p><p>I work with very complex software. At Tyk, we build API gateway software used by banks, governments, and other serious enterprise customers. I am a very sceptical person. I always want some proof. At the same time, I understand that software is always about compromises.</p><p>But the game is changing. What was not possible in the past is now possible for small teams in terms of quality and processes. The wind is changing with AI.</p><p>The true power happens when you can apply some of those approaches to legacy large enterprise codebases. If it works there, it will work everywhere.</p><p>I know how challenging it is. You cannot do it in one go. You cannot just make a switch and start using a new process.</p><p>This is not only about the technical part. It is also about the people part. Even at the size of Tyk, with like a hundred people, it is not about the implementation. It is about the processes and the people. The technical part is the easiest one.</p><p>In order to convince people that you can actually make it, you need to be able to do it in parts. Start small, then scale.</p><p>Can you take small parts, turn them into a repeatable process, and then start scaling? That is how it works in the majority of cases.</p><p>So I picked the policy engine. Authorization and gateway policy decisions are obviously critical. If the policy engine behaves incorrectly, you are not talking about a cosmetic bug.</p><p>I applied the same kind of thinking to the Tyk policy package in a public PR: <a href="https://github.com/TykTechnologies/tyk/pull/7932">TykTechnologies/tyk#7932</a>.</p><div class="captioned-image-container"><figure><div class="image-link image2" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!LEmi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffab79f2d-481f-4368-b891-eb8af99b9a29_1362x334.png 424w, https://substackcdn.com/image/fetch/$s_!LEmi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffab79f2d-481f-4368-b891-eb8af99b9a29_1362x334.png 848w, https://substackcdn.com/image/fetch/$s_!LEmi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffab79f2d-481f-4368-b891-eb8af99b9a29_1362x334.png 1272w, https://substackcdn.com/image/fetch/$s_!LEmi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffab79f2d-481f-4368-b891-eb8af99b9a29_1362x334.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!LEmi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffab79f2d-481f-4368-b891-eb8af99b9a29_1362x334.png" width="1362" height="334" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fab79f2d-481f-4368-b891-eb8af99b9a29_1362x334.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:334,&quot;width&quot;:1362,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:72886,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.reqproof.com/i/195788861?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffab79f2d-481f-4368-b891-eb8af99b9a29_1362x334.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!LEmi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffab79f2d-481f-4368-b891-eb8af99b9a29_1362x334.png 424w, https://substackcdn.com/image/fetch/$s_!LEmi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffab79f2d-481f-4368-b891-eb8af99b9a29_1362x334.png 848w, https://substackcdn.com/image/fetch/$s_!LEmi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffab79f2d-481f-4368-b891-eb8af99b9a29_1362x334.png 1272w, https://substackcdn.com/image/fetch/$s_!LEmi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffab79f2d-481f-4368-b891-eb8af99b9a29_1362x334.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></div></figure></div><p></p><p>81% ordinary coverage. 64.3% decision coverage. The normal coverage number says most statements ran. The MC/DC number says a lot of policy decisions still do not have independent evidence.</p><p>For a policy engine, the second number is the one I care about.</p><h2><strong>Code coverage is not about a metric</strong></h2><p>It is about trust.</p><p>What do we trust? In classical software engineering, we say: here is the code and here are the tests, the tests are the source of truth. If you want to know how the system works, read the tests.</p><p>I do not believe that anymore. Not with AI writing code. Not with AI writing tests. Not with AI validating its own assumptions.</p><p><strong>The source of truth cannot just be tests anymore. AI can write those too.</strong></p><p>A passing test can prove that the code agrees with the test. It cannot prove that both agree with my intent.</p><p>So I moved the source of truth up. For me, it has to be the specification: the static description of what I expect the system to do.</p><p>Then code implements it. Tests witness it. Coverage measures evidence around it. Traceability keeps the chain from silently rotting.</p><p>I started this whole journey because of one CVE in a library I wrote six years ago. I ended up in a completely different place.</p><p>I thought the problem was in the code. It was in what I forgot to describe.</p><p>I thought coverage was the answer. It was the wrong question.</p><p>The <a href="https://blog.reqproof.com/p/ai-writes-your-code-nobody-verifies">first article</a> was about losing intent. This one is about binding intent back to code.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Verification Gap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Writes Your Code. Nobody Verifies the Intent.]]></title><description><![CDATA[AI made implementation faster, but it did not solve trust. In both solo projects and regulated enterprise systems, the real bottleneck is still verification of intent.]]></description><link>https://blog.reqproof.com/p/ai-writes-your-code-nobody-verifies</link><guid isPermaLink="false">https://blog.reqproof.com/p/ai-writes-your-code-nobody-verifies</guid><dc:creator><![CDATA[Leonid Bugaev]]></dc:creator><pubDate>Thu, 23 Apr 2026 15:09:01 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f7ca47e8-c7d7-48e0-80bd-7205cc410018_1451x720.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I live in two different worlds now.</p><p>In one, AI made me more productive than I have ever been.</p><p>I have written more software in the last two years than across the rest of my career. I have barely written any code manually in the last year.</p><p>That part is real.</p><p>The speed boost is real.</p><p>The weird part is what came with it.</p><p>AI helps me ship more.</p><p>But it also asks me to trust more.</p><p>That is the uncomfortable part.</p><p>I am not just delegating typing.</p><p>I am delegating thinking, validation, and judgment too.</p><p>And I am still not sure where the safe line is.</p><p>In the other world, I lead engineering for software used by banks, governments, and other regulated environments, where mistakes are expensive and confidence matters more than speed.</p><p>And if you ask whether AI made us ship features 2x faster there, the honest answer is no.</p><p>Not even close.</p><p>That does not mean AI was useless.</p><p>It helped somewhere else.</p><p>It reduced noise.</p><p>A lot of engineering time in a big system does not go into writing the feature. It goes into interruption-based work: support engineers trying to understand how a feature behaves, PMs trying to figure out whether something is a bug or intended behavior, solution architects pulling in senior engineers just to inspect a corner of the system.</p><p>Tools that let people talk to the codebase, inspect it safely, and even generate tests or benchmarks to validate a hypothesis helped a lot with that.</p><p>People were less interrupted.</p><p>Context switching got better.</p><p>Engineers were happier.</p><p>But the main bottleneck did not move.</p><p>Implementation got dramatically faster. Trust did not.</p><p>That is the wall I keep hitting in both worlds.</p><h2><strong>The part people keep smoothing over</strong></h2><p>The industry keeps talking as if faster code generation automatically means faster engineering.</p><p>It does not.</p><p>In a lot of teams, it just means mistakes can scale faster than judgment.</p><p>As an individual engineer, I can create software much faster than before. Good software too. Clean structure. Tests. Refactors. Nice terminal output.</p><p>And still I trust it less than I want to.</p><p>Maybe less than before, because I know how much invisible reasoning I no longer fully own.</p><p>As a Head of Engineering, I can see the same problem from the other side.</p><p>We can accelerate some parts of the flow.</p><p>But we still have to verify whether the thing we built is actually the right thing, and whether it behaves correctly in the bigger system.</p><p>In a complex product, implementation is a relatively small slice of the work.</p><p>Validation and verification are the bigger slice.</p><p>That is why I keep coming back to the same phrase:</p><p><strong>verification gap</strong></p><p>The verification gap is the distance between what I mean and what I can actually prove.</p><p>Between intended behavior and demonstrated behavior.</p><p>That gap always existed.</p><p>AI did not invent it.</p><p>It just made it wider, faster, and easier to ignore until production forces the issue.</p><h2><strong>Why this got worse with AI</strong></h2><p>When humans wrote the code, the same brain often held the intent, the implementation, and the validation loop together.</p><p>Not perfectly.</p><p>People still shipped bugs. Specs were incomplete. Tests missed things.</p><p>But there was at least one place where the system could be understood as a whole: the person writing it.</p><p>That is no longer the default.</p><p>Now the human writes the prompt.</p><p>The model writes the code.</p><p>The model writes the tests.</p><p>The human skims the diff.</p><p>The model writes the cleanup.</p><p>The CI passes.</p><p>The feature ships.</p><p>And if the original intent was slightly wrong, incomplete, or misunderstood, that mistake does not stay in one place anymore.</p><p>It gets propagated through the whole stack.</p><ul><li><p>The plan is based on the wrong assumption.</p></li><li><p>The implementation is based on the wrong assumption.</p></li><li><p>The tests are based on the wrong assumption.</p></li><li><p>The &#8220;manual validation&#8221; is often you asking the same model to sanity-check itself.</p></li></ul><p>And then you look at the whole thing and it feels solid.</p><p>But it is solid on top of the wrong assumption.</p><p>So what exactly are we proving at that point?</p><p>That the system is internally consistent with the assumption it invented for itself.</p><p>Not that it matches your intent.</p><p>That is why so much AI productivity discourse feels fake to me.</p><p>A lot of teams did not automate engineering.</p><p>They automated typing.</p><p>That difference matters more than most people want to admit.</p><h2><strong>Bug free is not the same as intent-correct</strong></h2><p>People keep saying: just write better tests.</p><p>I do write tests.</p><p>AI writes tests for me too.</p><p>That is not the point.</p><p>Tests verify behavior for cases somebody thought of.</p><p>That somebody used to be a human.</p><p>Now it is often a human plus a model.</p><p>That is still not the same thing as verifying intent.</p><p>You can have 100% line coverage and still completely miss the thing that matters.</p><p>You can have a green CI run and still not know whether the software behaves the way you intended.</p><p>You can even have bug-free code in a narrow sense and still have software that is wrong.</p><p>A green pipeline can still be a polished misunderstanding.</p><p>That is one of the biggest traps in the current AI coding wave.</p><p>We are getting very good at generating artifacts.</p><p>Code.</p><p>Tests.</p><p>Docs.</p><p>Migration scripts.</p><p>Benchmarks.</p><p>RFC drafts.</p><p>None of that answers the deeper question:</p><p>does the system actually do what we mean?</p><h2><strong>Software is not flat. It is layers.</strong></h2><p>The problem gets worse as the software gets bigger.</p><p>Software is not flat.</p><p>It is layers.</p><p>It is wide, deep, and full of interacting components, hidden assumptions, backwards compatibility constraints, old decisions nobody remembers, and behavior that only makes sense if you know four other subsystems.</p><p>Any project that lives long enough eventually reaches a point where one brain is no longer enough.</p><p>That was true before AI.</p><p>It is still true now.</p><p>AI does not remove that limit.</p><p>In some cases it makes you hit it faster, because you can generate change faster than you can understand its consequences.</p><p>That is why the industry created all the layers around engineering in the first place:</p><ul><li><p>CI/CD</p></li><li><p>QA</p></li><li><p>RFCs</p></li><li><p>Architecture reviews</p></li><li><p>Team ownership boundaries</p></li><li><p>Support escalation paths</p></li><li><p>Approval workflows</p></li></ul><p>These are not random rituals.</p><p>They are patches over the same underlying problem:</p><p>software complexity grows beyond what one brain can safely manage.</p><h2><strong>Where does intent live now?</strong></h2><p>I think mainstream software engineering is still missing something fundamental.</p><p>We do not maintain a real source of truth for intent.</p><p>If I ask where the intended behavior of a system lives right now, the honest answer in most teams is:</p><p>all of it combined badly.</p><p>Some of it is in source code.</p><p>Some of it is in tests.</p><p>Some of it is in RFCs.</p><p>Some of it is in Jira tickets.</p><p>Some of it is in Confluence.</p><p>Some of it is in the heads of senior engineers.</p><p>None of those is the place where I can go and see, clearly, how the system is supposed to behave right now.</p><p>That is not a source of truth.</p><p>That is archaeology.</p><p>And that feels like a drastic difference from fields like aerospace or automotive.</p><p>They have their own fragmentation problems too. Different groups write requirements, validate them, implement them, monitor them. Those worlds often barely talk to each other.</p><p>But at least intended behavior is treated as a first-class artifact.</p><p>There is an SRS.</p><p>There are explicit requirements.</p><p>There is a recognized place where intent is supposed to live.</p><p>In mainstream software, especially for something complex like an API gateway, that still feels almost unimaginable.</p><p>We mostly reconstruct intent after the fact from scattered artifacts.</p><p>And then we act surprised when regressions keep happening.</p><h2><strong>Why enterprise teams do not get the full AI payoff</strong></h2><p>This is also why the conversation about AI productivity is often too shallow.</p><p>Yes, implementation is faster.</p><p>Sometimes dramatically faster.</p><p>But if speed of implementation is no longer the hard part, then what is?</p><p>That is the real question.</p><p>If a feature can be implemented in hours instead of weeks, why have so many teams not seen the full payoff?</p><p>Because implementation was never the only bottleneck.</p><p>The harder part is deciding what should be built, making that intent explicit enough, and then verifying that the resulting system still matches it after the code, tests, and surrounding context have all changed.</p><p>That is where the time goes.</p><p>That is also where a lot of current AI hype becomes unserious.</p><p>People showcase how fast a model can produce code.</p><p>Fine.</p><p>Show me how fast your team can decide what is correct, verify that the behavior matches the intent, and avoid turning six months of hyperproductivity into twelve months of regression cleanup.</p><p>At work, we effectively built a zero-trust environment.</p><p>We do not blindly trust humans.</p><p>We do not blindly trust AI.</p><p>We review the code.</p><p>We validate the assumptions.</p><p>We check the tests.</p><p>That posture protected quality when AI adoption accelerated.</p><p>But it also meant we did not suddenly become 10x faster.</p><p>We became less noisy.</p><p>More focused.</p><p>Better at answering questions.</p><p>Faster in implementation.</p><p>Still constrained by verification.</p><h2><strong>Not everyone needs safety. Everyone needs trust.</strong></h2><p>As an individual engineer, the same tension shows up in a different shape.</p><p>I can move incredibly fast.</p><p>But I know that if I let trust slide too far, I eventually stop building and start doing bug fixing and regression management full-time.</p><p>The software turns into glue and patches.</p><p>You can feel your taste slipping if you are not careful.</p><p>It all kind of works, but you are no longer fully sure why.</p><p>Safety bar differs. Obviously.</p><p>A bank flow is not the same thing as a weekend prototype.</p><p>One component inside a product may deserve a much stricter baseline than another.</p><p>But trust? Everyone needs that.</p><p>If I built a website, a product, a service, an internal tool, whatever it is, I need to trust that it actually follows my intent closely enough for the context it lives in.</p><p>That is the standard I care about.</p><p>Not some abstract perfection.</p><p>Not a fantasy of zero bugs.</p><p>Not a productivity screenshot.</p><p>Trust.</p><p>Can I tell how my software behaves right now?</p><p>Do my docs, specs, tests, and code align with each other?</p><p>Do I know which parts are intentional, which parts are accidental, and which parts are cargo cult left over from earlier decisions?</p><p>When I change something, am I making the system better, or just shifting uncertainty around?</p><p>So what is engineering now, exactly?</p><p>Where is the place of the human?</p><p>Where is the place of judgment?</p><p>And which part should I never offload, even if AI is very good at pretending it can carry it for me?</p><p>Those were already hard questions before AI.</p><p>AI did not create them.</p><p>It amplified them.</p><p>It exposed how incomplete our current software practices already were.</p><h2><strong>Why I am writing this</strong></h2><p>That is why I do not think a smarter model or a shinier coding assistant will solve this by itself.</p><p>The missing layer is verification.</p><p>Not just whether the code runs.</p><p>Not just whether the tests pass.</p><p>Not just whether the reviewer approved.</p><p>I mean verification of intent.</p><p>That is what I have been thinking about for a long time now, and why I am starting this newsletter.</p><p>I want to write about the gap itself, what causes it, why it compounds, why mainstream software and regulated engineering barely learn from each other, and what it would take to close it.</p><p>Not with slogans.</p><p>With examples, systems, failures, tools, and uncomfortable questions.</p><p>AI did not remove the hard part of engineering.</p><p>It moved it from writing to verification.</p><p>If this problem feels familiar, subscribe.</p><p>This is what I am writing about now.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.reqproof.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Verification Gap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>