TL;DR: A new engineering profession is emerging: the Software Verification Engineer. Manual coding is dead. Agents will write the code and the tests; someone needs to build the rules, environments, and verification infrastructure that make those changes trustworthy. Human review does not disappear. But what we review becomes the intent, the impact, the evidence, and what remains unresolved—not every line of the implementation.
The agent writes all of the code, all of the tests, and everything needed to run them. It opens a pull request. The tests pass.
What am I supposed to review?
Take an installation change in a product with multiple components. The agent says it tested the installation, and I believe that it ran something. But did it use the setup I wanted it to use? Did it install everything from scratch, or upgrade an existing installation? Were the other components running the latest versions too, or the versions our customers actually have?
Those are different questions. Reading the changed code does not tell me which of those things happened.
If I am working on a user interface, I would like to see the video where the agent clicks through the changes and narrates what it is doing. For an API change, I want the actual test attached: what it called, what came back, and what environment it was running against. Basically, the test I would otherwise do manually, except I want the agent to do it and show me the result.
Now somebody has to make all of that possible. Prepare the environments, define what needs to be checked, capture the results, and make them useful to the person accepting the change.
That is the engineering function I see emerging. It is much deeper than automatically writing tests.
We still need to review something
Our end goal is essentially prompt to production. But there is still a human review piece here. Putting a human at the end to read everything the agents produced is not going to sustain it.
We need to build our engineering processes around agent-written implementation, rather than treating it as an exception that a human will carefully inspect afterwards. All of the code, all of the tests, everything. The question is how we make that process good enough to trust.
This is why I think code review is dying as the main way we accept changes. We cannot keep increasing the amount of implementation and expect the same human review process to absorb it.
And when I accept a product change, I am accepting much more than the code diff. It is about the global picture: people, processes, dependencies, and the environments where the product needs to work. Code review is just one of the places where we try to understand that picture.
So we still need to review something. We need to understand what we are accepting and why we should believe it works. The question is what we put around the change so that a human can make that decision without reconstructing the entire verification process.
A good QA agent is not the whole answer
Let’s imagine we already have the fully automated QA person. A QA agent that is actually good at its job, follows instructions, asks the uncomfortable questions, and knows how to investigate something that looks wrong.
That would be useful. But go back to the installation change. How does this agent get the environments it needs? Where does it find the supported combinations of components? How does it know which upgrade paths matter? What should it capture to demonstrate that the installation worked?
Somebody needs to prepare the infrastructure, the rules, and the evidence types. Otherwise, we are still relying on someone remembering to ask, “Did you try it with this configuration?”
When it comes to evidence, it is about trust. Different people have different levels of trust, and different companies need it done differently. I might be happy with a recorded walkthrough for one change. Another change might need compatibility tests, failure scenarios, and evidence that we can recover if something goes wrong.
The recording is not a replacement for all the other checks. It answers a particular question I have about the change. The same is true of a benchmark, an API transcript, or an installation run. Each needs to tell me something relevant.
So even with a very good QA agent, there is a lot of engineering left to do. We need to decide what evidence is required and build the means to produce it reliably.
The rules cannot live only in the prompt
It can start with something like code coverage. Let’s say the agent gets it to 100%. Fine, but what have we established? We have not established that the tests check the right behaviour. High coverage can coexist with weak tests; that distinction existed long before coding agents. martinfowler.com
Now imagine the agent misunderstood the requirement and wrote both the implementation and the tests around that misunderstanding. Everything could pass, and we would still have the wrong change.
How do we validate that the agent sticks to the original intent and does not divert from it?
I want the intended behaviour to remain something the implementation is checked against. If the agent discovers that a requirement needs to change, that needs to become an explicit decision. It should not quietly change the requirement, update the tests to match, and present the result as finished.
The same applies to the guardrails. Telling an agent to follow a rule is not the same as checking that it followed it. Wherever we can turn that rule into an executable check, we should.
For the installation change, that might mean the pipeline requires an upgrade run against a defined configuration. A fresh installation should not satisfy that requirement. A description of how the upgrade would probably work should not satisfy it either.
We need to add as much determinism to the system as possible. I am not asking for the agent to take exactly the same steps every time. I want the acceptance rules to hold regardless of which steps it takes. A missing result should remain missing. A failed check should not become a passed check because the agent found a convincing explanation for it.
And the agent doing the implementation should not have unrestricted control over the rules that decide whether its implementation is acceptable.
Security warnings need the same discipline. If an AI system says something is vulnerable, what supports that claim? Can it reproduce the problem? Can it show the relevant path and the conditions required? If it cannot establish the finding, I want to know that too. I do not want a suspicion presented as a verified problem, or a confident dismissal presented as proof that everything is fine.
Every assumption we can check with code should be checked with code. Where we cannot establish something, it should remain visible as unresolved.
The agent needs to close its own evidence gaps
The loop should not be a human saying, “You missed the test here. Please add it,” followed by another review where the human finds the next missing thing.
That is still a process where the human has to reconstruct the verification plan, one comment at a time.
I want the agent to be able to ask: what is currently missing from my evidence? Which requirement have I not demonstrated? Which configuration has not been exercised? Did I actually resolve that warning, or just explain it away?
Then it should be able to do the missing work. Start another environment, run the upgrade, capture the result, investigate the failure, and try again. The pipeline should check whether the required evidence is actually present and whether it satisfies the acceptance rules.
There has to be a way to stop and say that it cannot finish, too. If a requirement is ambiguous or an environment is unavailable, that should come back as a specific unresolved issue. The agent should not keep going until it finds a way to call the change complete.
This is where we need an almost inhuman amount of quality work. Not just more generated tests, but much more verification than a human could reasonably carry out for every change.
There is a useful precedent here. In his essay on testing, Dan Luu described hardware teams investing in custom test generators and running them across thousands of machines. The engineering effort went into systems that could investigate far more cases than people could write individually. Dan Luu
That is the kind of ambition I want for agentic software engineering. The agents should not just produce more implementation. They should do the enormous amount of work required to demonstrate that the implementation is acceptable.
I want the evidence to be thorough enough that even the most stubborn QA person can say, “Yes, I’m happy.” But that needs to mean their questions were answered, not that we produced so much output they gave up reading it.
Evidence has to make review easier
There is an obvious problem here: if we replace a huge code diff with a huge collection of logs, recordings, and reports, we have just moved the review problem somewhere else.
The evidence needs to be grouped around what the human is accepting. I should be able to understand the conclusion, inspect the relevant result, and see what was not covered. For each result, I want to know which build and configuration it came from, rather than having to guess whether it still applies to the change in front of me.
Essentially, I need to answer four questions:
What am I accepting?
What could it affect?
Why should I believe it works?
What still remains unresolved?
Some of those questions extend into production. Pre-production evidence does not remove the need to observe a rollout, limit its exposure, and recover from failures. Charity Majors made this point about gradually earning confidence in releases long before this discussion about agents. charity.wtf
So the verification system needs to be clear about what has been established before release, what can only be checked during the rollout, and what should stop that rollout. “Ready to deploy” should not quietly become “nothing can go wrong.”
This is the job of the Software Verification Engineer
The work includes the custom harnesses, the environments, the acceptance rules, the evidence formats, and the way those results reach a human. Different industries will need different systems. Different changes within the same product will need different evidence.
This is not a claim that verification itself is new. It is a change in what the function owns.
The Software Verification Engineer is responsible for the system that makes agent-produced changes acceptable. What needs to be demonstrated? Which rules must be enforced? What infrastructure does the agent need? How do we know that the evidence is valid? What happens when it is incomplete?
The agents can help build that infrastructure too. This is not about preserving a special category of code that humans must still write. The responsibility is deciding what the system must enforce and establishing that it actually does.
We have been talking about better automation, testing infrastructure, and reliable delivery for years. Agentic development is going to force us to connect those things properly. If agents do the implementation, the way forward is to dramatically increase the trust and engineering quality of the pipelines around them.
For the next installation change, I do not want to remember to ask whether the upgrade was tested. I want the system to know that the upgrade evidence is required, give the agent the means to produce it, and refuse to call the change ready when it is missing.
Building that system is the job of the Software Verification Engineer.


