Skip to content
wins.solutions

The Production Checklist for Vibe-Coded Applications

A practical checklist for reviewing AI-assisted applications before putting real users, data or money behind them.

FreeGuide15 min readIntermediateBy wins.solutions teamUpdated

An app built with Cursor, Claude Code, Codex or Copilot can go from idea to working demo in a day. "It works on my machine and the demo looks right" proves that the happy path works for one trusted user. It says nothing about a second user signing up, someone editing a request in the network tab, a deploy failing halfway through a migration, or a server-only key ending up in a JavaScript bundle.

AI assistants are good at satisfying the prompt and much less reliable at honouring requirements nobody wrote down, such as "users must not see each other's invoices". This checklist is a structured review for that gap. It does not replace a security review or a standard such as the OWASP Application Security Verification Standard, but it covers the problems that commonly separate a prototype from something you can put real users, data or money behind.

How to use it:

  1. Work top to bottom. Later sections assume the earlier ones are done.
  2. Tick items honestly. "The model probably handled that" is not a tick. Open the code, send the request, look at the table.
  3. Write down what you skip and why, in the repository. A skipped item with a reason ("OAuth only, no passwords") is a decision. One without a reason is a gap you will rediscover later.

Scope

Before reviewing code, decide what the application is supposed to do. AI-assisted projects accumulate features nobody designed: an admin page scaffolded "for completeness", a public API route generated alongside a form, a webhook handler left over from an abandoned prompt. Each is reachable code that accepts input and needs the same authentication, authorization and validation as everything else. Code you did not ask for is code you have not reviewed, and it is still attack surface.

Product requirements

Write a few sentences on what the app does, who uses it, what data they create and who may see it. Every authorization rule is judged against this. "It seems fine" is not a requirement.

Critical flows

List the flows where a failure costs a user something real: sign-up and login, payments, creating and deleting the core record, sharing, data export, account deletion. These get the most scrutiny in testing.

Out-of-scope functionality

Inventory every entry point. In a Next.js app, that means every route.ts file, every "use server" function, every scheduled job and every database function callable from the client. Anything outside your requirements gets deleted, disabled, or promoted into scope and reviewed. Deleting is usually right: code that no longer exists needs no securing.

Architecture

You should be able to say where each responsibility lives. Generated code tends to blur boundaries: a database client next to UI code, a permission check in the browser, an SDK called from the client with a secret key.

  1. Browser
  2. Server action / API
  3. Auth check
  4. Database
Every read or write of user data passes a server-side check before it reaches the database.

Boundaries

The browser is untrusted: anything that runs there can be read and modified by the user. The server verifies identity, enforces permissions and validates input. The database is the last line: constraints and, where applicable, row-level policies that hold even when application code has a bug.

For each endpoint that touches user data, point to the line where the auth check happens. "The page redirects when you're logged out" protects the page, not the server action or API route it calls. Server actions accept direct POST requests, and the Next.js data security guide says to verify authentication and authorization inside each one.

Dependencies

Assistants add packages freely. Confirm that each entry in package.json is imported, maintained and the package you meant: lookalike names are a known supply-chain risk, and a model can suggest a name that does not exist or belongs to someone else. Commit the lockfile, run your package manager's audit, and decide on each finding.

Third-party services

List every external service (auth, database, payments, email, storage, analytics, AI APIs) with its environment variables, whether its key is server-only, what your app does when it is down, and whether its webhooks are signature-verified. An unverified webhook endpoint accepts a fake "payment succeeded" event from anyone.

Authentication

Authentication answers "who is making this request?" Use your framework's established auth library or a hosted provider, not a generated implementation. Password hashing, session tokens and reset flows are exactly the code that looks correct in a demo and fails under attack.

Sessions

  • Session cookies are HttpOnly, Secure in production, and have an explicit SameSite attribute.
  • Sessions expire, and logout invalidates the session on the server, not only in the browser.
  • Every handler takes the user ID from the verified session, never from the body, query string or headers.

Password and reset flows

If the app has password login, check the parts a demo never exercises:

  • Passwords are hashed with Argon2id, scrypt or bcrypt, ideally by your auth library.
  • Reset tokens are random, single-use, short-lived and stored hashed.
  • Login and reset give the same response for known and unknown emails, so they cannot be used to enumerate accounts.
  • Login and reset endpoints are rate limited.
  • Changing the email or password requires re-authentication.

With OAuth or magic links only, the provider handles most of this. Still check that its redirect URL allowlist contains only your real domains, not localhost or an old preview URL.

Authorization

Authorization decides whether this user may do this to this specific record. Prompts rarely state it, so generated code can easily omit it: the model builds updateProject(id, data) as asked, and nothing said the project had to belong to the caller. The result is an insecure direct object reference: change an ID in the request, get someone else's data.

UI checks are not authorization

Showing the Delete button only when project.ownerId === user.id is good UX and no security at all. Anyone can copy the button's request from the network tab, change the ID and send it again. The UI decides what people see. The server decides what they can do.

Ownership check on the server

Load the record, compare its owner with the session user, and only then mutate:

app/projects/actions.ts
"use server";
 
import { z } from "zod";
import { db } from "@/lib/db";
import { getSession } from "@/lib/auth";
 
const RenameProjectInput = z.object({
  projectId: z.string().uuid(),
  name: z.string().trim().min(1).max(100),
});
 
export async function renameProject(input: unknown) {
  const session = await getSession();
  if (!session) {
    throw new Error("Unauthorized");
  }
 
  const { projectId, name } = RenameProjectInput.parse(input);
 
  const project = await db.project.findUnique({
    where: { id: projectId },
    select: { ownerId: true },
  });
 
  // Same response for "doesn't exist" and "not yours",
  // so the endpoint can't be used to probe for valid IDs.
  if (!project || project.ownerId !== session.user.id) {
    throw new Error("Not found");
  }
 
  await db.project.update({
    where: { id: projectId },
    data: { name },
  });
 
  return { ok: true };
}

The user ID comes from the session, not the input. The input is typed unknown and parsed, because a server action receives whatever the caller sends. Missing records and other people's records get the same error.

Alternatively, scope the write itself: an updateMany filtered by both id and ownerId, treating zero updated rows as not found. That also closes the gap between check and write.

Roles and tenants

Check roles on the server from data you load there, never from a role field in the request. Keep the check in one helper, such as requireRole(session, "admin"), so it is easy to review. In multi-tenant apps, scope every query by the tenant ID from the session. List endpoints count too: returning every row and filtering in the browser has already leaked everything.

Database

Constraints

The database should reject bad data even when application code lets it through. Check for:

  • NOT NULL on required columns.
  • Foreign keys with a deliberately chosen ON DELETE behaviour.
  • Unique constraints where uniqueness matters: email addresses, one membership per user per team.
  • CHECK constraints for simple invariants, such as non-negative quantities.
  • Money stored as integer minor units or fixed-precision decimals, never floating point.

Migrations

Every schema change should be a migration in version control. If the production schema was built in a dashboard or by an assistant running SQL directly, capture a baseline migration now and prove it by building a fresh database from migrations alone. Review generated migrations like generated code: changing a column type by dropping and re-adding the column deletes its data.

Row Level Security

If the browser queries the database directly, as it does with Supabase's client libraries, Row Level Security is your authorization layer. Supabase's RLS documentation states that a table in an exposed schema without RLS is readable and writable by any role with a grant on it, and the key carrying those grants ships in your bundle by design. Enable RLS on every exposed table with policies scoped to the authenticated user; using (true) on user data is the same as no policy. Common Supabase RLS mistakes covers the failure patterns in detail.

If all database access goes through your server with a privileged connection, RLS is optional defence in depth, and your server-side checks have to be complete.

Backups

Confirm that backups run, how long they are kept, and whether your plan includes point-in-time recovery. Then restore one into a separate database and check the data is there. A backup you have never restored is an assumption.

Secrets

Server and client boundary

Secrets belong in per-environment variables, never in source files. The subtler failure is a secret crossing into the browser.

Frontend frameworks inline some environment variables into the client bundle at build time, selected by a name prefix: NEXT_PUBLIC_ in Next.js, VITE_ in Vite, EXPO_PUBLIC_ in Expo. The Next.js environment variable docs describe this as replacing each reference with a hard-coded value in the JavaScript sent to the browser. The prefix does not mean "semi-private". It means published.

Only publishable values get the prefix: a Supabase URL and publishable key, a Stripe publishable key, an analytics ID. Database URLs, Supabase secret keys (sb_secret_..., or the legacy service_role key), Stripe secret keys, webhook secrets and AI provider keys never do. Watch generated changes here: when a server-only variable is undefined in the browser, adding the prefix silences the error by publishing the key. The fix is to move the call to the server. In Next.js, import "server-only" in modules that read secrets turns an accidental client import into a build error.

.env.example
# Committed with names only. Real values live in your hosting
# provider and in git-ignored .env.local files.
 
# Server-only: never add a public prefix to these
DATABASE_URL=
SUPABASE_SECRET_KEY=
STRIPE_SECRET_KEY=
STRIPE_WEBHOOK_SECRET=
OPENAI_API_KEY=
SESSION_SECRET=
 
# Inlined into the client bundle at build time: must be safe to publish
NEXT_PUBLIC_SUPABASE_URL=
NEXT_PUBLIC_SUPABASE_PUBLISHABLE_KEY=
NEXT_PUBLIC_STRIPE_PUBLISHABLE_KEY=

Leaked keys in git history

Deleting a key from a file leaves it in every earlier commit, clone and fork. Before a repository goes public or gains collaborators, scan its full history with a tool such as gitleaks or TruffleHog, and enable your Git host's secret scanning where available.

Rotating keys

Rotate by issuing a new credential, deploying it everywhere, confirming the app works, and only then revoking the old one. Practise once before launch to learn where each key lives (hosting provider, CI, teammates' machines) and which need a redeploy. After a real exposure, check the provider's usage logs for activity you do not recognise.

Validation

Validate every input on the server: form data, JSON bodies, query strings, route parameters, headers, webhook payloads, and model output your app acts on. Client-side validation is user experience; it protects nothing, because requests can be sent without your UI. Validate bounds as well as types, or a notes field becomes a multi-megabyte row or an expensive AI prompt.

app/api/invoices/route.ts
import { z } from "zod";
import { getSession } from "@/lib/auth";
import { createInvoiceForUser } from "@/lib/invoices";
 
const CreateInvoice = z.object({
  customerId: z.string().uuid(),
  currency: z.enum(["usd", "eur", "gbp"]),
  dueDate: z.coerce.date(),
  notes: z.string().trim().max(2000).optional(),
  lineItems: z
    .array(
      z.object({
        description: z.string().trim().min(1).max(200),
        quantity: z.number().int().min(1).max(1000),
        unitAmount: z.number().int().min(0).max(10_000_000), // cents
      }),
    )
    .min(1)
    .max(50),
});
 
export async function POST(request: Request) {
  const session = await getSession();
  if (!session) {
    return Response.json({ error: "Unauthorized" }, { status: 401 });
  }
 
  const body = await request.json().catch(() => null);
  const parsed = CreateInvoice.safeParse(body);
  if (!parsed.success) {
    return Response.json({ error: "Invalid request" }, { status: 400 });
  }
 
  // parsed.data is typed, bounded and stripped of unknown keys.
  // createInvoiceForUser still checks that customerId belongs to
  // this user, and computes the total itself.
  const invoice = await createInvoiceForUser(session.user.id, parsed.data);
  return Response.json({ id: invoice.id }, { status: 201 });
}

Stripping unknown keys prevents mass assignment: a raw body passed into db.user.update lets a user add "role": "admin" to a profile edit. Zod object schemas drop unknown keys by default, so write only the parsed data. For payments, never trust a client-sent amount; look the price up on the server.

While you are in the handlers, also search for SQL built by string concatenation, dangerouslySetInnerHTML on user content, and error responses that expose stack traces.

Testing

You do not need full coverage before launch. You need tests for what would hurt most.

Critical user paths

Turn your critical flows into end-to-end tests (Playwright or similar) that run against a production build in CI on every pull request. That way an assistant's edit to a shared component cannot quietly break sign-up or checkout.

Failure cases

Generated test suites tend to cover the happy path. Add the unhappy cases deliberately:

  • User B reads, updates and deletes user A's records by ID, for every resource type, and gets a 403 or 404.
  • Unauthenticated requests to every mutation are rejected.
  • Malformed and oversized input returns a 400, not a 500.
  • Badly signed webhooks are rejected, and duplicate deliveries are not applied twice.
  • A failing third-party call, such as an email provider timeout, leaves the database consistent.

Cross-user tests are cheap (two test users and a loop over your endpoints) and cover the failure that matters most in a multi-user app.

Regression tests

When you fix a bug, add a test that fails without the fix. A model editing nearby code does not know why a line exists and can reintroduce a fixed bug; the test is that memory. Watch for failing tests "fixed" by weakening the assertion, and review test diffs as carefully as code diffs. The AI code review checklist covers reviewing generated changes more broadly.

Observability

Error tracking

Set up error tracking on server and client, upload source maps so stack traces are readable, and route alerts somewhere you actually look. After the first production deploy, trigger a deliberate error and confirm it arrives.

What to log

  • Authentication events: logins, failed logins, password resets, email or MFA changes.
  • Authorization failures. A burst of requests for records a user does not own is what probing looks like.
  • Payment and webhook events, with the provider's event ID.
  • Errors, with a request ID that is also returned in the response.
  • Admin actions and data exports: who did what, to which record, and when.

Use structured JSON logs so you can filter by user ID or request ID.

What never to log

  • Passwords, including attempted passwords on failed logins.
  • Session tokens, JWTs, API keys, cookies and Authorization headers.
  • Password reset and magic-link tokens, or URLs containing them.
  • Full card numbers and security codes. With a hosted checkout, these never reach your server.
  • Full request bodies on authentication and payment routes.

These rarely leak through an explicit console.log(password). They leak through generic logging: a console.log(req.body) left from debugging, an error handler that serialises the whole request, an error tracker capturing headers by default. Search for those patterns and configure your error tracker's scrubbing. The OWASP Logging Cheat Sheet has a fuller exclusion list.

Deployment

Production environment

  • Production has its own database, keys and auth provider configuration, shared with no other environment.
  • Preview deployments cannot reach the production database.
  • Debug modes, verbose error pages, seed scripts and test accounts are disabled or removed.
  • Auth redirect URLs and CORS allowlists name only real origins.

Schema changes on deploy

Migrations usually run before the new code goes live, so each must also work with the code currently running. Expand, then contract: add a nullable column, deploy code that writes it, backfill, and only in a later release add the NOT NULL constraint or drop the old column. This keeps code rollback possible. If a release drops a column the previous version reads, redeploying that version will not help.

Rollback plan

Write the plan down before launch. It should answer:

  1. Trigger. Which signals mean roll back (error rate, a broken critical path, suspected data corruption), and who makes the call.
  2. Code. The exact command or dashboard action that redeploys the previous build. Many hosts can promote a previous deployment directly; confirm yours can.
  3. Database. Whether the release included migrations, whether they are backward-compatible, and if not, the restore procedure and how much data it would lose.
  4. Configuration. Environment variables changed in the release and their previous values. Public-prefixed variables are inlined at build time, so changing one needs a rebuild.
  5. Verification. Critical-path smoke tests pass and the error tracker goes quiet.
  6. Communication. Who needs to be told, and where status updates go.

Rehearse the rollback once on staging and time it. Your first rollback should not happen during an incident.

Final release gate

Copy this into your launch pull request. Tick an item only once you have verified it, and note why anything stays unticked.

Scope

  • The app's users, data and access rules are written down in the repository
  • Every entry point is inventoried, and out-of-scope ones are removed or disabled

Architecture

  • Every dependency is intentional, the lockfile is committed, and audit findings are resolved or noted
  • Every incoming webhook verifies the provider's signature

Authentication

  • Every handler takes the user ID from the verified session, never from the request
  • Session cookies are HttpOnly, Secure and SameSite, and logout invalidates the session server-side
  • Login and password reset are rate limited and do not reveal whether an email exists

Authorization

  • Every mutation re-checks ownership or role on the server before writing
  • Every list and detail query is scoped to the session user or tenant
  • Roles are read from server-side data, never from the request

Database

  • Required columns are NOT NULL, and foreign keys and unique constraints match the data model
  • A fresh database built from migrations alone matches production
  • Every client-exposed table has RLS enabled with user-scoped policies
  • A backup has been restored into a separate database and checked

Secrets

  • No secret has a client-exposed prefix such as NEXT_PUBLIC_ or VITE_
  • The built client assets (for example .next/static) contain none of your secret values
  • Git history has been scanned, and every secret ever committed has been rotated
  • .env.example lists names only, and real .env files are git-ignored

Validation

  • Every server entry point parses input with a schema that sets length and size limits
  • No handler writes a raw request body to the database
  • Prices and totals are computed on the server

Testing

  • End-to-end tests for every critical flow pass in CI against a production build
  • Cross-user access tests exist for every resource type and pass

Observability

  • A deliberate production error appears in error tracking with a readable stack trace
  • Sampled production logs contain no passwords, tokens, cookies or card data

Deployment

  • Production has its own database and keys, unreachable from preview deployments
  • Every migration in the release works with the previously deployed code
  • The rollback plan is written down and has been rehearsed
  • Free

    How to Review AI-Generated Code Before Shipping It

    A practical process for reviewing AI-generated diffs for behaviour, security and maintainability before they reach production: what to read first, what to distrust, and what to ask the agent.

    GuideWebIntermediate

    FreeRead
  • Free

    Frontend Developer Roadmap 2026

    A step-by-step path from your first web page to job-ready frontend work: HTML and CSS, JavaScript, Git, React, TypeScript, testing, accessibility, performance, deployment and responsible use of AI coding tools. Each step has a project to build and free official resources.

    RoadmapWebBeginner

    FreeView roadmap
  • Free

    Supabase RLS Mistakes That Can Expose Your Application

    Your Supabase key ships in every browser bundle, so Row Level Security decides what each request can touch. These are the policy mistakes common in Supabase apps, especially AI-generated ones, and how to verify policies before launch.

    GuideWebIntermediate

    FreeRead