What a Planner-Executor Agent Actually Does in AI QA Testing
TL;DR: A planner-executor agent splits one job into two roles: a planning model looks at the current screen and decides the single next action, and a separate system actually carries that action out in a real browser. The split exists because one model doing both reasoning and precise clicking is worse at both. WayRunner runs exactly this loop against your live app, plus a third role most write-ups skip entirely: a model whose only job is explaining why a step failed in words you can act on.
You searched this term because you kept seeing “planner” and “executor” mentioned like they’re obvious, and they’re not. Most of what comes up is either an academic paper describing the general pattern for any multi-step AI task, or a post from a testing vendor using the same two words to mean something closer to validating a pull request before it merges. Neither one tells you what’s actually happening when a tool like this points a browser at your app and reports back pass or fail. That’s the gap this post is for.
Why one model can’t do both jobs well
Ask a single large language model to both figure out what to do next on a page and then precisely click the right pixel, and you get a model that’s mediocre at both. Deciding “the next step is to submit this form” requires broad reasoning over a screenshot and the surrounding context; actually finding the submit button and firing a real click event is a narrower, more mechanical job, and it usually needs different tooling entirely, not just a different prompt. So the architecture splits it: a Planner reasons over what it sees and emits one concrete action at a time, click this, type that, navigate here. An Executor takes that single instruction and carries it out against a real browser, then reports back what changed. Neither one tries to do the other’s job.
This isn’t unique to testing. The same planner-executor split shows up anywhere an AI agent has to take more than one step toward a goal, and it’s the reason “one giant prompt that does everything” tends to fall apart past a few steps. What differs is what the Executor is actually driving. For a coding agent, it’s a shell and a filesystem. For a QA agent, it has to be a real, rendered browser, because the only honest way to know if a signup flow works is to watch it happen the way a visitor would.
How the loop runs against your actual app
Here’s what that looks like end to end. You paste your app’s URL and a plain-English instruction, something like “sign up with a new email, verify it, and confirm the dashboard loads.” Before spending a full run, a quick feasibility check does a basic reachability probe on the URL and a sanity pass on the instruction, so an unreachable site or an obviously impossible ask fails fast instead of burning a run.
Then the loop starts. The Planner (Claude Haiku 4.5, specifically, chosen for being cheap enough to call on every single step) looks at a screenshot plus a summary of the page’s structure and decides exactly one next action. That action gets dispatched to the Executor, a real Chromium browser driven by Magnitude Core, an open-source vision-first browser framework built for exactly this: reading a page visually and acting on it directly rather than depending on a prewritten script or a stable set of selectors. The Executor carries out the action and hands back a fresh observation, and the Planner reasons over that to decide the next one. This repeats until the flow is done or a hard limit trips: 40 steps, the plan’s browser-ready time limit, or a cap on how many tokens the run is allowed to burn through. Those limits exist so a confused agent fails loudly and quickly instead of quietly looping.
The role most explanations of this leave out
Here’s where the generic version of this pattern usually stops, at “planner decides, executor acts, done.” That’s fine if the audience already knows how to read a stack trace. It’s useless if you’re a founder who just wants to know what broke. So WayRunner adds a third role on top: when a step fails, a Failure Analyst gets the full, uncapped reasoning transcript, the screenshot, and the DOM at the moment of failure, and turns that into a plain explanation of what went wrong. That’s the actual difference between “step 14 failed” and “the confirmation email link led to a page that never finished loading, so the account stayed unverified.” One of those tells you something you can fix.
It’s worth being direct about how this differs from another architecture you’ll find using nearly identical language: some AI coding tools describe a Planner, Generator, and Evaluator loop that reviews AI-written code before it merges, structured that way because a model grading its own generated code tends to rate it positively even when it’s broken. That’s a real and useful pattern, but it’s solving a different problem for a different audience: it assumes a GitHub repo, a pull request, and a CI pipeline to gate. WayRunner’s loop never touches code at all. It watches the live, rendered app the way a user would, which is the only option left once there’s no repo to read in the first place, the exact situation covered in QA testing for AI-built apps.
Where two models actually cost you something
None of this is free, and it’s worth saying so plainly. Two models in the loop means two round trips of latency per step instead of one, and a cheap Planner can still misread a cluttered or ambiguous screen, especially on pages with a lot of overlapping modals or dynamic content that hasn’t finished rendering. A vision-based Executor is more resilient to a redesigned button than a hardcoded CSS selector would be, but it still depends on the page actually looking the way it’s supposed to; if your app renders broken on a slow connection, the agent sees the same broken render a real visitor would. This architecture buys you resilience to change, not immunity to a genuinely confusing UI. If your flow is that ambiguous to a model looking at it fresh, it’s probably also going to confuse a first-time user, which is its own useful thing to learn from a failed run.
What this means if you’re the one running the test
You don’t have to think about any of this to use it; you paste a URL, describe a flow in a sentence, and read a report. But the two-role split, plus the third role that translates a failure into English, is why the report says something more useful than pass or fail. We walk through what that setup actually looks like on a real signup flow in how to test your Lovable, Bolt, or Replit app before you ship it, and if you’re wondering what the broader “AI QA agent” label actually covers beyond this one architecture, that’s laid out in AI QA agents, explained for indie founders.
If you want to see the loop run against your own app: wayrunner.run/#signup.