Skip to content
wins.solutions

How to Review AI-Generated Code Before Shipping It

A practical process for reviewing AI-generated diffs for behaviour, security and maintainability before they reach production: what to read first, what to distrust, and what to ask the agent.

FreeGuide16 min readIntermediateBy wins.solutions teamUpdated

The most dangerous AI-generated code isn't obviously broken. It's code that appears reasonable enough that nobody reviews it carefully.

Obviously broken code is the easy case: it fails the build, the tests or the first click. The expensive problems are diffs that compile, pass CI, read cleanly and quietly do something other than what you asked. Whether the change came from Cursor, Claude Code, Codex or GitHub Copilot, the problem is the same: the code is cheap to produce and still expensive to understand.

This guide is a review process for a single AI-generated change, with the highest-yield checks first. For the wider question of whether an app built this way is ready for real users, see the production checklist for vibe-coded applications.

Why AI diffs are different

Reviewers of human-written code lean on signals they rarely notice: the author's track record, rough edges that show where they struggled, a size limited by typing speed. AI-generated diffs remove most of them.

  • Volume. An agent writes a thousand lines in the time it takes you to read fifty. The reviewer becomes the bottleneck, and bottlenecks get skipped.
  • Plausible-looking code. Consistent formatting and idiomatic structure hide the usual tells of careless work.
  • Confident naming. A function called requireAdmin or sanitizeHtml describes what the model intended, not what the body does.
  • Changes outside the requested scope. You asked for a new filter; the agent also widened a type to silence an error and edited a shared helper.
  • New dependencies. Adding a package is often the shortest path to working code.
  • Deleted checks. A validation that failed a test, or a guard that caused a type error, can disappear while "making it work".
  • Tests that assert the implementation. Tests written alongside the code tend to confirm what it does, not what it should do.

None of this makes AI-generated code worse by default. It means the review has to be deliberate.

Understand the change

Before reading any code, establish which files changed, why each one changed, and what behaviour is now different. Start with the shape of the diff, not its contents.

# Uncommitted agent work: untracked files, then staged and unstaged changes
git status --short
git diff --stat HEAD
 
# A branch compared with the point where it left main
git diff --stat main...HEAD
git diff --name-status main...HEAD

--name-status marks each file as added (A), modified (M), deleted (D) or renamed (R), so deletions and renames surface before you read a line. New files do not appear in git diff until they are tracked; git add -N <path> makes them show up without staging their content.

Then ask the agent to explain itself:

Summarise the change you just made:
1. Every file you created, modified or deleted, and why.
2. Any behaviour that changed beyond what I asked for.
3. Any dependency, config, migration or environment variable you added or changed.
4. Anything you removed, disabled or skipped: checks, tests, error handling.
5. What you did not verify.

Finally, write one sentence yourself: "After this change, X happens when Y." If you cannot, you do not understand the change well enough to approve it.

Review the diff

Read in an order that gives you context before detail:

  1. Tests, to see what behaviour the change claims.
  2. Types, schemas and migrations, to see the shape of the data.
  3. Entry points: routes, server actions, handlers, jobs and webhooks, where outside input arrives and access checks belong.
  4. Core logic, then UI, then configuration.

While reading, look for:

  • Out-of-scope edits. Changes unrelated to the task, often hidden in formatting churn; git diff -w ignores whitespace-only changes.
  • Removed code. A missing if, throw or await requireUser() produces no error anywhere, so deleted lines deserve more attention than added ones.
  • Renames. A renamed export, route, column or environment variable breaks every caller the agent did not see, including clients already deployed.
  • Suppressed warnings. A new @ts-ignore, as any, eslint-disable or loosened tsconfig usually means a problem was hidden, not fixed.
  • Generated files. Lockfiles, snapshots and generated clients. Review the input that produced them and confirm they regenerate identically. Marking them linguist-generated in .gitattributes collapses them in GitHub pull request diffs.

A quick scan for common red flags:

# Added lines that suppress checks or skip tests
git diff main...HEAD | grep -E '^\+.*(@ts-ignore|@ts-expect-error|as any|eslint-disable|\.skip\(|\.only\()'
 
# Removed lines that look like guards
git diff main...HEAD | grep -E '^-[^-].*(auth|session|role|owner|throw|valid)'

These are heuristics, not a substitute for reading.

The most effective fix is upstream: keep AI changes small. Ask for one concern per change, have the agent plan before editing, and commit after each step. If you cannot hold the whole change in your head, it is too large to approve.

Dependencies

Every new package is code you did not review, running with your application's permissions. Check package.json and the lockfile in every AI-generated diff, even when the task was unrelated.

For each new package:

  • Does it exist, and is it the one you meant? Models sometimes suggest names that do not exist or sit one character from a popular package. Attackers register both lookalikes (typosquatting) and commonly hallucinated names (slopsquatting). Open the registry page and confirm the repository link points where you expect.
  • Is it maintained? Check the last publish date, open issues and whether the repository is archived.
  • Is it necessary? lodash for one groupBy, uuid where crypto.randomUUID() exists, or a second date library next to the one you already use.
  • Is the license compatible with how you distribute your product?
  • Does it run code on install? preinstall, install and postinstall scripts run on every developer machine and CI runner. npm runs them by default; recent pnpm versions block them unless you allow them.
npm view <pkg> version license repository.url time.modified
npm view <pkg> scripts      # look for install-time scripts
npm ls <pkg>                # where it sits in your dependency tree
npm explain <pkg>           # why it is installed
npm audit                   # known vulnerabilities

Read the lockfile diff too. Adding one package should not rewrite hundreds of entries; unexplained churn usually means the lockfile was regenerated, silently upgrading transitive dependencies. Check that resolved URLs point at the registry you expect.

When a package passes these checks, consider committing it as its own change so the decision stays visible in history.

Authentication

Authentication answers "who is this?". AI changes tend to weaken it in recognisable ways:

  • New entry points without the existing guard. A route that falls outside your middleware matcher, or a server action that assumes the page rendering it is protected. In Next.js, server actions can be invoked with a direct POST request, so each one must check the session itself.
  • Unverified tokens. Decoding a JWT is not verifying it. Supabase's documentation, for example, warns against trusting getSession() in server code and recommends methods that validate the token, such as getUser() or getClaims().
  • Fallback identities. session?.user.id ?? "demo-user" or a SKIP_AUTH flag added to make something run locally, then forgotten.
  • Weakened session settings. Changed cookie flags (httpOnly, secure, sameSite), longer token lifetimes, or removed rate limits on login and password reset.

Call every new entry point without a session. It should be rejected, not return data or crash with a 500.

Authorization

Authorization answers "may this user do this to this record?". Generated code often looks correct here and is not, because the missing piece is a condition that never appears in the diff.

The classic failure is an insecure direct object reference (IDOR): the handler checks that someone is logged in, then fetches whatever ID the request names.

lib/invoices.ts
// Before: any logged-in user can read any invoice
export async function getInvoice(invoiceId: string) {
  await requireUser();
  return db.invoice.findUnique({ where: { id: invoiceId } });
}
lib/invoices.ts
// After: ownership is part of the query
export async function getInvoice(invoiceId: string) {
  const user = await requireUser();
  const invoice = await db.invoice.findFirst({
    where: { id: invoiceId, ownerId: user.id },
  });
  if (!invoice) throw new NotFoundError(); // same response for "missing" and "not yours"
  return invoice;
}

With ownership in the query, no code path loads another user's row, and "not found" in both cases avoids confirming the ID exists. Also watch for:

  • Identifiers from the client. A userId, orgId or role in the request body is a claim, not a fact. Derive them from the session.
  • Role checks moved to the client. Hiding the delete button from non-admins is UX. If the endpoint does not check the role, anyone with curl is an admin.
  • Missing tenant scoping. In multi-tenant apps every query needs the organisation condition, including counts, exports and search.
  • Service credentials as a fix. When an agent hits a permission error and switches to an admin client or service-role key, the error disappears because authorization disappeared.

The OWASP Authorization Cheat Sheet covers the principles in more depth.

Input validation

Validate at the server boundary every time, even when the form already validates in the browser. AI-generated code gets this wrong in repeatable ways: validating only on the client, calling a schema and then using the raw body anyway, or casting parsed JSON with as SomeType, which checks nothing at runtime.

const UpdateProfile = z.object({
  displayName: z.string().trim().min(1).max(80),
  bio: z.string().max(500).optional(),
});
 
const user = await requireUser();
const input = UpdateProfile.parse(await req.json());
await db.user.update({ where: { id: user.id }, data: input });

Zod's z.object strips unknown keys by default, so a role: "admin" field in the payload never reaches the database. Passing the raw body into an update is mass assignment, and it is easy to miss because the code looks tidy.

Also check for:

  • Raw SQL built with string interpolation, including ORM escape hatches such as Prisma's $queryRawUnsafe.
  • User-controlled HTML reaching dangerouslySetInnerHTML or a Markdown renderer with raw HTML enabled.
  • User-supplied URLs that the server fetches (server-side request forgery) or redirects to (open redirects via parameters like ?next=).
  • Missing upper bounds on strings, arrays, file sizes and limit parameters.

Database migrations

Migrations are where a small-looking diff can do irreversible damage, because they run against real data, often automatically on deploy.

  • Edited history. To resolve schema drift, an agent may edit an already-applied migration or suggest prisma migrate reset or db push. An existing migration file marked M in --name-status is a red flag.
  • Destructive changes. Dropped columns and tables, truncating type changes, and renames. During a rollout the previous version is still serving traffic, so a rename behaves like a drop plus an add.
  • Locks on large tables. Many ALTER TABLE forms take an ACCESS EXCLUSIVE lock, and while one waits for it, every query behind it waits too. A plain CREATE INDEX blocks writes until it finishes. A new foreign key or check constraint scans the whole table unless added as NOT VALID and validated separately.
  • Irreversible steps. A transformation that discards data cannot be undone by a down migration. Back up first.
  • Defaults. A new column's default is a product decision applied to every existing row. is_public boolean default true publishes old data.
  • Row Level Security. New tables without RLS enabled, policies with using (true) or with check (true), and views or security definer functions that bypass policies. The guide to common Supabase RLS mistakes covers these in detail.
-- Blocks writes to orders until the index is built
create index orders_customer_id_idx on orders (customer_id);
 
-- Allows writes during the build, but cannot run inside a transaction block
create index concurrently orders_customer_id_idx on orders (customer_id);

Many migration tools wrap each file in a transaction, so check how yours handles this. The PostgreSQL ALTER TABLE documentation lists the lock level for each form.

Secret handling

  • Hardcoded values. API keys in source, test fixtures, seed scripts, or a .env.example containing real values. Run a scanner such as gitleaks or TruffleHog in CI; GitHub push protection blocks many common key formats.
  • Secrets moved into the client bundle. Variables prefixed with NEXT_PUBLIC_, VITE_ or EXPO_PUBLIC_ are inlined into code shipped to browsers and devices. When a server-only value is undefined in a client component, the shortest fix is to add the prefix, and renaming STRIPE_SECRET_KEY to NEXT_PUBLIC_STRIPE_SECRET_KEY publishes it. Importing a server module into a client component has the same effect; the server-only package turns that mistake into a build error.
  • Secrets in logs. console.log(req.headers), full request bodies, or HTTP client errors that include the request configuration and its Authorization header. Logs often have wider access than the database.
git diff main...HEAD | grep -iE '^\+.*(api[_-]?key|secret|token|password|private[_-]?key)'
git diff main...HEAD | grep -E '^\+.*(NEXT_PUBLIC_|VITE_|EXPO_PUBLIC_)'

If a secret was committed or shipped, rotate it. Deleting it in a later commit does not remove it from history, anyone's clone or a bundle already deployed.

Failure behaviour

Generated code is usually written for the path where everything works. Review what happens when it does not:

  • Swallowed errors. Empty catch blocks, catch blocks that only log, and .catch(() => null).
  • Fallbacks that lie. data ?? [] after a failed request renders "You have no invoices" instead of an error. The user now believes something false.
  • Unbounded retries. Retry loops without a maximum, backoff or idempotency, especially around payments, emails and webhooks.
  • Missing timeouts. An outbound call without a timeout can hold your request open far longer than a user will wait. For fetch, AbortSignal.timeout() is a one-line fix.
  • Partial writes. Related writes without a transaction, so a failure halfway leaves inconsistent data.
// Swallowed: the order is marked paid even when the charge fails
async function completeOrder(orderId: string) {
  try {
    await payments.charge(orderId);
  } catch (err) {
    console.error(err);
  }
  await db.order.update({ where: { id: orderId }, data: { status: "paid" } });
}
// Handled: the failure is recorded, surfaced and safe to retry
async function completeOrder(orderId: string) {
  try {
    await payments.charge(orderId, { idempotencyKey: `order-${orderId}` });
  } catch (err) {
    await db.order.update({ where: { id: orderId }, data: { status: "payment_failed" } });
    throw new Error("Payment failed", { cause: err });
  }
  await db.order.update({ where: { id: orderId }, data: { status: "paid" } });
}

For each failure path, ask what the user sees. An accurate error and a way to retry are part of the feature, and so is not leaking stack traces or internal IDs into the response.

Testing

Passing tests are only evidence if they could have failed.

  • Does the test fail without the change? Keep the new test, temporarily restore the old implementation and run it. A test that passes either way is not testing the change.
  • Tests that mirror the implementation. Tests that mock every collaborator and assert the mocks received exactly the arguments the code passes. They break on every refactor and catch few bugs. Prefer observable outcomes: return values, persisted rows, HTTP responses.
  • Edited assertions. When an agent is asked to make tests pass, changing the expected value is often the shortest path. Every modified assertion in an existing test is a behaviour change you are approving, and updated snapshots deserve the same scrutiny.
  • Weak assertions. toBeDefined(), toBeTruthy() and "does not throw" pass for many wrong answers.
  • Edge cases. Empty input, boundaries, duplicates, time zones, concurrent requests, and a request from a different user. The agent covers the cases you described; the rest are yours to add.
# With the change uncommitted: set aside only the implementation, keep the new test
git stash push -- src/lib/pricing.ts
npm test -- pricing    # this should fail
git stash pop

Performance

AI-generated code is usually fine at demo scale. Review it as if the tables were a hundred times larger:

  • N+1 queries. A query inside a loop, or inside Promise.all(items.map(...)), is still one query per item. Load related data with a join, an include, or a single IN query.
  • Unbounded queries. findMany() without take, select * pulling large columns, a client-controlled limit with no maximum, or no pagination on lists, exports and search.
  • Missing indexes for new filter and sort columns. EXPLAIN ANALYZE shows the real plan, but it executes the statement, so do not run it on writes against production.
  • Work in render paths. Sorting large arrays on every render, effects that refetch in a loop because of unstable dependencies, or the same query issued several times for one page.
  • Per-request setup. Creating a new database client inside a request handler, which can exhaust connections under load, particularly on serverless platforms.

Most of these are invisible against a local seed database with a handful of rows. Seed realistic volumes before reviewing anything that lists, searches or aggregates.

Maintainability

The agent will not maintain this code; you will. Look for:

  • Duplication. A new formatCurrency next to the one already in lib/format.ts. Search for the concept before accepting a new helper.
  • Dead code. Leftovers from earlier attempts in the same session, unused exports, commented-out blocks. Knip, or TypeScript's noUnusedLocals and noUnusedParameters, will find much of it.
  • Inconsistent patterns. A different HTTP client, error shape or folder convention from the rest of the codebase. Consistency lets the next person, or the next agent, work without guessing.
  • Over-abstraction. A generic BaseRepository<T> or a factory for a feature with one caller. Abstractions should follow repetition, not anticipate it.
  • Comments that restate code. // increment the counter adds nothing, and comments describing the change ("updated to use the new API") go stale immediately. Keep comments that explain why.

Final checklist

Use the full list for anything touching money, permissions or stored data. When time is limited, start with the 30-minute pass.

A 30-minute review pass

  1. Scope
  2. Access
  3. Data
  4. Secrets
  5. Failure
  6. Tests
What to check first when time is short
  1. Scope. Compare --name-status with the agent's summary and note every unexpected file.
  2. Access. For each new or changed entry point, confirm the session check and query scoping.
  3. Data. Read migrations and anything that writes.
  4. Secrets and dependencies. Scan added lines for keys and public prefixes, and look up any new package.
  5. Failure. Search the diff for catch and read every block.
  6. Tests. Confirm one test fails without the change and no existing assertion was quietly edited.

If anything in the first three steps worries you, stop and do the full review.

The full list

Scope and diff

  • Read git diff --stat and --name-status before any code
  • The agent's summary matches the files that actually changed
  • Every out-of-scope edit is justified or reverted
  • Deleted lines were read as carefully as added ones
  • No new @ts-ignore, as any, eslint-disable or loosened config

Dependencies

  • Every new package exists, is the intended one and is maintained
  • Install scripts and licenses of new packages were checked
  • Lockfile changes are proportional to the dependency change

Security

  • Every new entry point rejects requests without a valid session
  • Queries are scoped to the caller's user or organisation on the server
  • No IDs, roles or ownership fields are trusted from the request body
  • Input is validated on the server and only the parsed result is used
  • No secrets in source, fixtures, logs or public environment variables
  • No service-role credentials added to fix a permission error

Data

  • No applied migration was edited
  • Destructive or locking migrations have a rollout plan
  • New column defaults are safe for existing rows
  • New tables have RLS enabled, with policies tested as a second user

Failure and tests

  • No empty or log-only catch blocks on important paths
  • Retries are bounded and idempotent; outbound calls have timeouts
  • Related writes happen in a transaction
  • Users see an accurate error, not a misleading empty state
  • At least one new test fails when the change is reverted
  • No existing assertion or snapshot changed without a reason

Performance and maintainability

  • No queries inside loops; lists are paginated, bounded and indexed
  • No duplicated helpers, dead code or comments that restate the code

Questions to ask

Ask the agent, and treat the answers as leads to verify:

  • What did you change that I did not ask for, and why?
  • What did you remove, disable or skip to make this work?
  • Which inputs would make this fail, and what happens then?

Ask yourself:

  • What happens if this request is sent by a different logged-in user?
  • What happens if it is sent twice, or twice at the same moment?
  • What happens when the third-party call times out?
  • What does this migration do to the rows already in production?
  • Could I explain every line of this diff without the agent's help?

Once you approve a change, the code is yours, whichever tool wrote it.

  • Free

    The Production Checklist for Vibe-Coded Applications

    A practical checklist for reviewing AI-assisted applications before putting real users, data or money behind them.

    GuideWebIntermediate

    FreeRead
  • Free

    Frontend Developer Roadmap 2026

    A step-by-step path from your first web page to job-ready frontend work: HTML and CSS, JavaScript, Git, React, TypeScript, testing, accessibility, performance, deployment and responsible use of AI coding tools. Each step has a project to build and free official resources.

    RoadmapWebBeginner

    FreeView roadmap
  • Free

    Supabase RLS Mistakes That Can Expose Your Application

    Your Supabase key ships in every browser bundle, so Row Level Security decides what each request can touch. These are the policy mistakes common in Supabase apps, especially AI-generated ones, and how to verify policies before launch.

    GuideWebIntermediate

    FreeRead