I’ve been following Zach Lloyd’s series on building agentic software factories, and his latest post stopped me mid-scroll. Not because of the technical details (though those are genuinely interesting), but because of one specific capability he describes: agents that can watch their own work, confirm whether it actually solves the problem, and keep iterating until it does.

That loop, reproduce, verify, iterate, is something most executive decision-making processes never achieve. And I think that gap is worth talking about.

The verification problem in business decisions

Here’s what Lloyd describes in practical terms. His agents don’t just write code and ship it. They use computer and browser control to actually run the application, watch what happens, capture video or screenshots, and confirm whether the behavior matches the spec. If it doesn’t, the agent keeps working.

The agent isn’t just producing output. It’s checking whether the output did what it was supposed to do.

Now ask yourself: how many business decisions in your company get that treatment?

Most of the decisions I see at the CEO level follow a simpler path. Someone identifies a problem, a solution is proposed, the solution gets approved and implemented, and then… the team moves on to the next problem. Whether the original issue was actually resolved often goes unmeasured. Or it gets measured six months later, by which point the signal is buried under a dozen other variables.

Lloyd’s framing here is useful even if you never write a line of code: the verification step is where most of the real learning happens.

Reproduce before you fix

One of the most underappreciated points in the post is about the triage phase. Lloyd argues that before agents try to fix a bug, they should first reproduce it. Confirm the problem actually exists, and exists in the way it was reported, before spending resources on a solution.

This sounds obvious. In practice, it’s rare.

Think about how many initiatives in your company are “fixing” problems that haven’t been cleanly defined. Someone escalates a complaint, a meeting happens, and resources get deployed before anyone has verified the precise nature of the issue. You end up with solutions chasing descriptions instead of solutions chasing confirmed behaviors.

The discipline of reproduction, of forcing yourself to demonstrate the problem before you act on it, is one of the most valuable habits a decision-maker can build. It’s the same discipline I push with clients in a different vocabulary: diagnose before you spend. And it requires good data infrastructure underneath it, which is why I keep coming back to this theme at JLytics. You cannot reproduce a business problem if you don’t have reliable, accessible data that shows you what’s actually happening.

The self-improving loop

The part of Lloyd’s post that I find most strategically relevant is his description of what happens when you combine verification with spec-driven development. If the agent has a detailed product specification, it can run computer use at every implementation pass, compare what it sees to what the spec requires, and keep iterating until the two match.

That is a closed feedback loop. The spec is the decision criteria. The verification is the measurement. The iteration is the response.

Your business decisions have the equivalent of all three of these components available to you, but most organizations haven’t connected them into a loop. The spec exists as a strategic plan or a quarterly goal. The verification exists as a report somewhere. The iteration happens, if at all, at the next planning cycle.

The gap between those cycles is where value leaks. A loop only compounds if it remembers: a team can review numbers every single month and still learn nothing, because nobody recorded what changed next to the results it produced. Cause logged beside effect is what turns activity into learning, for an AI agent and for a leadership team alike.

From my perspective doing data strategy work with mid-market companies, the organizations that close this loop fastest are the ones that have invested in two things: clear KPIs that actually define “done,” and data infrastructure that lets them check those KPIs on a reasonable cadence. Not quarterly. Not monthly. Often weekly or faster, depending on what’s being tracked.

Proof reduces review burden

Lloyd makes another point worth lifting out: when an agent provides a video showing a feature working end-to-end, the human reviewer is more likely to trust the underlying code. For low-risk UI changes, he suggests the video alone might be sufficient and you may not need to review the code at all.

The principle translates directly to executive reporting. When a recommendation comes to you with a clear verification artifact, here is the problem we confirmed, here is what we changed, here is the measured result, you can review it faster and with more confidence. You’re not starting from scratch trying to assess whether the solution worked. The work of verification has already been done.

This is what good analytics infrastructure actually delivers. Not dashboards for their own sake, but decision artifacts that let you allocate your review time where it matters most.

What this means for how you run your organization

A few practical takeaways from Lloyd’s framework, translated for a CEO audience:

Require reproduction before resourcing. Before you approve budget or headcount to fix a problem, ask for a concrete demonstration that the problem is confirmed. Data showing the issue, not just a description of it.

Define verification criteria at the start. When you approve an initiative, specify in advance what “working” looks like. Lloyd calls this the acceptance criteria. It’s the equivalent of your KPI definition, and it needs to exist before the work begins, not after.

Close the loop on every major decision. Build a habit of returning to the original problem definition after implementation and asking whether the behavior changed. This requires having baseline data before the fix, which means your data infrastructure needs to capture state over time.

Parallelize verification where you can. Lloyd notes that fanning out verification across multiple cloud agents reduces latency significantly. The business equivalent is building teams and systems that can run multiple verification tracks simultaneously rather than sequentially.

The underlying theme across this entire series by Lloyd is that intelligent automation is not about replacing judgment. It’s about systematizing the habits that good judgment requires, reproduction, specification, implementation, verification, and observation after release. Those habits produce better outcomes whether the agent is a language model or a human team.

Your best next decision deserves better inputs. Start with a JLytics data assessment to see clearly where your business actually stands.

Original source: The computer use verification skill that every agent needs

Start the Conversation

Interested in exploring a relationship with a data partner dedicated to supporting executive decision-making? Start the conversation today with JLytics.