QA Testing for Bolt.new Apps: The Preview Is Not the Test

TL;DR: Bolt’s preview and your deployed site are two different environments, so “it works in the preview” is not evidence that it works for users. The Bolt-specific testing tools you’ll find each check one narrow slice: security scanning, usability with real people, or a generic rendering audit. None of them check whether the specific flow you care about still works after your last prompt. That check has to be built from what your app is actually supposed to do, and it has to be cheap enough to re-run every time you ship.

You prompted, Bolt built, the preview looked right, you hit deploy. Then someone tried to sign up and nothing happened. The button did its little press animation and the page just sat there.

The confusing part is that it still works in the preview. You go back, click through it yourself, and everything’s fine. So now you have a bug that only exists somewhere you can’t see it, and no idea whether the next thing you ship will do the same thing.

Why the preview isn’t the thing you’re shipping

This is worth understanding once, because it explains most of the “works here, not there” bugs you’ll hit with Bolt specifically.

Bolt runs your app inside StackBlitz’s WebContainers, an in-browser development environment that runs Node, npm, and a dev server inside the browser tab you have open. That’s the whole trick behind Bolt feeling instant; there’s no remote machine spinning up, it’s all right there.

Your deployed app is not that. It’s a build output sitting on a host, usually Netlify, being fetched by a stranger’s browser over a real network. Different runtime, different origin, different environment variables, different cold-start behavior on any API it calls.

Most of what breaks between those two lives in the gaps:

  • Environment variables. Values you set for the preview don’t automatically exist on the host. Netlify’s docs are explicit that build-time variables have to be defined in the site’s own settings. An API key that was present in the preview and absent in the build produces a page that renders fine and silently fails on every request.
  • Anything pinned to a URL. OAuth redirect URIs, email confirmation links, CORS allowlists, and webhook endpoints all reference a specific host. The preview host is not your deploy host, and your deploy host is not your custom domain.
  • Real network latency. In the preview the server is inside your browser, so responses are effectively instant. Deployed, they’re not. Code that never waits properly for a response works in the preview because there was never anything to wait for.

None of these are Bolt doing something wrong. They’re what happens when the place you look at your app and the place your users run it are genuinely different places.

What the Bolt testing tools you’ll find actually check

Search “test my Bolt app” and you get three tools that all sound like they do the same thing. They don’t, and the differences matter more than the marketing suggests.

Vorota scans Bolt apps for security vulnerabilities before production. That’s a real problem; AI builders leak keys and skip authorization checks constantly. It’s also not behavior. A perfectly secure app can have a signup button that goes nowhere.

Great Question covers user research: recruit five people, watch them use it, find out whether your flow makes sense to a human who isn’t you. Genuinely valuable, and it answers a question no automated tool can. But it answers it once, before launch, and it tells you whether your design is confusing, not whether your code broke.

QAlaunch is the closest to what you probably meant. Paste a deployed URL and it runs a set of automated checks across desktop and mobile viewports, then sells you a PDF report. It catches rendering problems and browser-specific breakage, which is real value for nine dollars.

Here’s the limit, and it applies to every generic audit of this shape: a tool running a fixed list of checks can only find problems it already knows how to describe. It doesn’t know that in your app, a user has to pick a plan before the dashboard will load, or that your invite flow only works if the email matches an existing workspace. It can tell you a button has poor contrast. It can’t tell you the button doesn’t do the thing your app exists to do.

That’s not a knock on any of them. It’s just that the failure you’re actually afraid of is specific to your app, so the check has to be specific to your app too.

What actually breaks in Bolt apps, and why it keeps happening

Bolt regenerates code in response to prompts, and a prompt about one feature can rewrite files that feature doesn’t obviously touch. This is the mechanism behind almost every regression in an AI-built app: you asked for a new settings page, you got a new settings page, and somewhere in that same edit the auth redirect changed.

You won’t notice, because you’ll test the settings page. That’s the feature you were thinking about. Nobody re-tests signup after adding a settings page.

There’s a second thing about Bolt that sharpens this. Bolt’s pricing runs on tokens, so every fix has a price. Discovering a broken signup flow four prompts later means the fix now has to work around three intervening changes, and you pay for the diagnosis and the fix and usually a retry. Catching it immediately after the prompt that caused it is not just faster; it’s the difference between one small correction and an expensive untangling.

So the useful habit isn’t a bigger pre-launch checklist. It’s a short list of flows you re-run after every prompt, which is a different discipline entirely. We wrote that process up in how to do regression testing without a QA engineer, and it’s the piece that matters most once your app is live. If you’re still before launch, the pre-launch checklist for Lovable and Bolt apps covers the one-time checks this post assumes you’ve done.

Where WayRunner fits, and one thing we had to build for apps like yours

WayRunner tests the deployed URL, not the preview. You write the check as a sentence, the way you’d say it to a person: “sign up with a new email, confirm it, and check the dashboard loads.” A Planner model looks at a screenshot of the current page and decides the next single action; a real Chromium browser carries it out and screenshots the result. When something fails, a separate model reads the whole transcript and writes out what went wrong in plain language, which is the part that matters if you’re not going to read a stack trace.

Here’s the specific thing, and it exists because of apps built exactly the way yours was.

Bolt apps are single-page apps, usually Vite and React. When you load one, the server sends back an HTML shell with an empty <div id="root"> in it, and the browser reports the page as loaded before any of your actual app has rendered. A naive test harness takes its first screenshot at that moment, sees a blank white page, and reports that your entire app is broken.

We hit this early and it made the tool useless on precisely the apps we built it for. So the browser controller doesn’t trust the page-loaded signal on the first navigation. It samples the DOM to check whether the root element has any child elements or visible text yet. If the app hasn’t mounted, it waits two seconds and samples again. If it’s still empty after that, it reloads once and gives it one more chance, because an SPA that lost a chunk request on the first try will usually mount fine on the second. Only after that does it call the page genuinely blank.

That’s a small piece of code and a boring bug. It’s also the difference between a testing tool that works on a Bolt app and one that reports false failures on every single run.

Credentials work the way you’d want if you’re testing a real login: they’re held in an encrypted vault and injected into the browser at the moment of typing, so the model choosing what to click never sees your password.

The takeaway

Test the deployed URL, not the preview; they’re not the same environment and the gap between them is where your first user-facing bug is going to live. Run a security scan, because those tools catch a category this one doesn’t. Then write down the three or four flows your app can’t work without, and re-run them after every prompt, not just before launch. The bug you’re worried about isn’t exotic. It’s the signup flow you stopped checking two features ago.

If you’d rather not click through those flows yourself every time Bolt regenerates something: wayrunner.run/#signup.