How to Test an App Built with AI Before Launch
A practical pre-launch checklist for apps built with Claude Code, Cursor, Lovable and other AI coding tools: what breaks, what to test and what to automate.
AI coding tools such as Claude Code, Cursor, Lovable, Bolt, Replit and v0 have changed how fast an app gets built. A founder can go from an idea to a working product in days. The demo works, the main screens look right, and the happy path does what it should.
The trouble is everything around the happy path. AI-generated code does what the prompt asked for. It doesn’t know your business rules, your users or what happens when two people use the app at once, unless someone told it. This guide is the checklist we use before an AI-built app goes live.
Why AI-built apps need a different kind of testing
Three things make apps built with AI coding tools fail in their own way:
- The code follows the prompt, not the product. A feature can be complete for the request and still miss a rule that was only in your head: who may edit an order, what happens to a refund after 30 days.
- Generated tests share the code’s blind spots. If the same tool writes the feature and its tests, the tests usually check what the code already does, not what it should do.
- Fast iteration breaks old features. Each new prompt can rewrite files that earlier features depend on. Without regression tests, nobody notices until a user does.
None of this means AI-built apps are bad. It means the testing has to come from outside the code: from the business, the users and the ways things go wrong.
1. Write down the critical paths first
Before testing anything, list the five to ten journeys your app cannot fail at. For most apps they look like this:
- Sign up, confirm the account, log in.
- The core action: place an order, book a slot, publish a post.
- Pay, and get a receipt.
- Reset a forgotten password.
- An admin finds and fixes a user’s problem.
Every later step is measured against this list. If you only have time for one thing, test these paths end to end, on a phone and on a desktop.
2. Log-in, sessions and passwords
Authentication is usually generated quickly and checked rarely. Test:
- Password reset links: do they expire, and can they be used twice?
- Log out on one device: is the session gone, or does an old tab still work?
- Expired sessions: does the app send the user back to log-in, or show a broken screen?
- Sign-up with an email that already exists, with uppercase letters, with spaces at the end.
3. Permissions and data isolation
This is where AI-built apps most often leak data. The app works for every single user, but nobody checked whether user A can see user B’s data.
Create two test accounts and try, as user A:
- Opening user B’s order, profile or invoice by changing the ID in the URL.
- Calling the same API request with user B’s ID.
- Reading a list endpoint and checking that it returns only your own records.
If the app talks to its database directly from the browser, as many Supabase and Firebase apps do, the database’s own access rules (for example row-level security) are the real lock. Check that every table has them, not just the ones the prompt mentioned.
4. Inputs and edge cases
Generated forms tend to handle the input in the example and little else. Try:
- Empty fields, very long text, emoji and Vietnamese diacritics.
- Pressing the submit button twice, quickly.
- A slow or dropped connection in the middle of saving.
- Dates across midnight and time zones, and amounts with decimals.
5. Payments and money
If money moves, test it like money:
- Can one click create two charges or two orders?
- What happens if the payment provider’s confirmation arrives twice, or late?
- Do refunds, discounts and rounding produce the right totals?
- Does the order state match the payment state after a failure?
6. Errors and empty states
Look at the screens the demo never showed: no results, no network, a server error, a deleted item. A blank page or a raw error message is a bug, even if nothing crashed.
7. Real phones and screen sizes
Open the app on an actual phone, not just a resized browser. Check that buttons are reachable, nothing scrolls sideways and the keyboard doesn’t cover the field you are typing in.
8. Secrets and configuration
Search the built app for API keys and service credentials. Anything sent to the browser is public. Server keys belong on the server, and test and production settings must not be mixed.
9. Automate the critical paths
Once the critical paths pass by hand, automate them so every future prompt is checked. A short end-to-end test with Playwright, Cypress or Selenium is enough to start. This one checks that one user cannot open another user’s order:
import { expect, test } from '@playwright/test';
test("a user cannot open someone else's order", async ({ page }) => {
await logInAs(page, 'user-a@example.com');
const response = await page.goto('/orders/ORDER-OF-USER-B');
expect(response?.status()).toBe(404);
await expect(page.getByText('ORDER-OF-USER-B')).toHaveCount(0);
});
Run these tests on every change, before anything is deployed.
10. Check again after deployment
A release that passed every test can still fail in production: a missing environment variable, a database migration, a payment setting. Right after each deployment, walk through the critical paths on the live app. It takes minutes and catches the problems users would otherwise find first.
The short version
- List the critical paths and test them end to end.
- Test permissions with two accounts: this is where data leaks hide.
- Try the inputs and timings the demo never used.
- Treat payments, errors and phones as first-class tests.
- Automate the critical paths, and check them again after every deployment.
This is the process we follow at Red Giant, including on ShipMe, the app we built with AI coding tools and run in production.