Engineering in the AI Era/Responsible Shipping
Lesson 1 of 1 · Episode 1

AI Wrote the Code. Ship It?

Build confidence in AI-written code through better plans, deliberate review, layered evidence, and controlled rollout.

AI code reviewVerificationGuardrailsRelease readiness
Watch on YouTube ↗

AI can produce a convincing pull request in minutes. That changes the bottleneck, not the responsibility: the engineer still owns whether the change is correct, safe, and worth releasing. The goal is not to trust AI; it is to build enough independent confidence to make a sound decision.

Review studioTurn an AI-generated diff into an engineering decision
date-range.ts · AI generated
18  export function lastSevenDays(now) {
+  const start = subHours(now, 168);
+  return { start, end: now };
21  }
➤
CONFIDENCE SIGNALS
◇Plan challenged
⌁Scope inspected
✓Evidence layered
◒Rollout controlled
BUILDING CONFIDENCE…
The decision

Confidence is accumulated, not declared

“The tests passed” is one useful fact. It is not a complete argument for shipping. Confidence grows when several different layers agree: the plan matches the product intent, the implementation matches the plan, the evidence covers the important risks, and the release limits the cost of being wrong.

Confidence modelSpeed comes from layers, not blind trust
AI can accelerate every stage. Ownership of the decision still stays with the engineer.
The ownership rule
Delegating code generation does not delegate accountability. If you approve the pull request, you own the reasoning behind that approval — including the assumptions the model made and the evidence you accepted.
Before implementation

Review the plan while change is still cheap

The most leveraged review happens before a large diff exists. First form your own view of the problem. Then give the model the relevant product, architecture, and repository context and ask it for a plan. Challenge the plan: what will change, what must remain unchanged, which edge cases matter, and how will the result be verified?

Plan modeChallenge an ambiguous requirement before implementation
“Add a filter for the last seven days.”
What does “seven days” mean?
The plan cannot be reviewed until this product decision is explicit.

Define success as observable behavior

A useful plan is testable. “Implement the filter” is a task. “A user can select a rolling seven-day window; dates use the account timezone; the export keeps its existing range; boundary cases have tests” is a success definition. The second version gives both the implementer and reviewer a shared contract.

The implementation

Direct attention without surrendering the whole diff

AI makes it tempting to ask another model to review everything and then accept the summary. A stronger workflow separates two questions: did the implementation touch the expected scope, and do the critical parts deserve deeper inspection? Start with the file list, dependencies, data flow, permissions, and integration points. Unexpected scope is itself a review signal.

Diff mapRead the shape of the change before reading every line
4 FILES · +86 −14
ƒdateRange.ts

The range calculation is the behavioral center — and has six existing consumers.

DEPENDENCY SURFACE6 consumers
Expected changeReview question
UI filterControl, state, query parametersDoes selection map to the agreed date semantics?
Shared date utilityParser and boundary helpersWhich other screens and exports call it?
TestsHappy path and boundariesWould a plausible wrong interpretation still pass?
Unexpected export fileNot mentioned in the planNecessary integration, accidental scope, or hidden coupling?
Small diff ≠ small risk
A two-line permission change may expose private data. A large generated fixture may have almost no production impact. Review attention should follow behavior and blast radius, not line count.
Verification

Ask what every green check actually proves

Automate recurring checks aggressively: repository rules, linters, types, tests, scripts, and review prompts turn repeated human feedback into durable guardrails. But no check is universal. Each one observes a different slice of reality, under specific assumptions.

EvidenceGreen checks answer different questions
What it proves

Known examples still behave as asserted.

What it cannot prove

Missing assertions and misunderstood requirements.

Strong evidence is intentionally redundant. A unit test can verify the cutoff rule, a typecheck can verify the contract, an independent review can inspect the integration, and a screenshot or recording can show the actual flow. Agreement across unlike checks is more valuable than running the same kind of check five times.

Review feedback

Investigate findings; do not count them

An AI reviewer can surface excellent leads and confident nonsense. For each finding, reproduce the behavior, decide whether it is real, check whether this change introduced it, and separate severityfrom scope. A severe issue might be local; a moderate issue might affect every consumer of shared logic.

TriageA finding is a lead, not a verdict
STEP 1 / 4
AI reviewer: empty dates may bypass the filter
Turn discoveries into guardrails
When a finding is valid, do more than patch the line. Add the missing test, lint rule, script, repository instruction, or review checklist so the same class of mistake becomes cheaper to catch next time.
Learning loopThe best review finding improves the next review too
Confidence grows because the system remembers what the reviewer learned.
Human advantage

Bring in context outside the diff

The repository rarely contains the whole decision. A related pull request may be changing the contract, another service may own the real constraint, or product and design may have revised the expected behavior yesterday. Humans add disproportionate value by connecting this outside context to the implementation under review.

Context radarThe decisive fact may live outside the repository
PR
#482
SIGNAL FROM RELATED PR

A parallel migration renames the API contract next week.

The diff cannot reveal this on its own. The reviewer connects it to the code and changes the decision.
Risk

Let impact determine the depth of review

Review depth should grow with exposure, sharedness, sensitivity, and irreversibility. A change behind an internal flag has a different confidence requirement from one that moves money, deletes data, changes permissions, or alters a shared library used by every client.

ImpactReview depth follows blast radius, not line count
1
Focused review
Target the changed behavior and its contract.
Shipping

Merge readiness is not release readiness

A pull request may be correct enough to merge while still needing a controlled release. Feature flags, employee access, limited cohorts, monitoring, and a recovery plan let you acquire production evidence without immediately exposing everyone.

Release readinessMerged is a code state; released is an exposure decision
Recovery check: turning off a feature flag stops new exposure; it may not undo writes, messages, or side effects that already happened.
Merge readyRelease ready
CodeReviewed, tested, and integratedDeployed in the intended environment
ExposureMay still be unreachableAudience and rollout stage are explicit
SignalsCI and review evidenceProduct, error, latency, and business metrics
RecoveryRevert or follow-up is possibleStop, repair, and data-recovery paths are rehearsed
Practice

Check your understanding

Q1Multiple choice
Why is reviewing the plan before implementation especially valuable with AI-generated code?
Q2Multiple choice
All tests, types, and lint checks pass. What can you safely conclude?
Q3Multiple choice
An AI reviewer reports a high-severity bug. What is the best next step?
Q4Sort each scenario
Classify each signal as primarily merge readiness or release readiness.
Typecheck and focused tests pass
Employee cohort and expansion criteria are defined
Unexpected file changes were investigated
Existing side effects can be repaired after rollback
Apply it

A review brief for your next AI-assisted change

Before you approve the next AI-written pull request, write down:

  1. The product behavior, unchanged behavior, edge cases, and success evidence.
  2. The expected files and integrations — then compare them with the actual scope.
  3. What each automated check proves and which important question remains unanswered.
  4. The impact factors that justify your chosen review depth.
  5. The rollout stages, monitoring signals, stop condition, and recovery path.
Key takeaways
  • →AI changes implementation speed; it does not move accountability away from the approving engineer.
  • →Confidence begins with a challenged plan and observable success criteria, not with a polished diff.
  • →Inspect scope first, then spend deep attention where behavior and impact justify it.
  • →Tests, static checks, independent review, and behavioral proof answer different questions.
  • →Treat AI findings as hypotheses to reproduce and classify, then convert valid discoveries into reusable guardrails.
  • →Review depth follows blast radius, sensitivity, and reversibility — never line count alone.
  • →A merge-ready change still needs deliberate exposure, monitoring, and recovery before broad release.