Continuous integration (CI) regularly integrates changes and checks them automatically. Continuous delivery keeps a tested release ready for deployment. Continuous deployment automatically releases changes that pass the configured gates. The release authorization policy determines which deployment steps are automatic.
An LLM application still needs ordinary software testing. It also needs checks for behavior affected by prompts, model settings, retrieval, tools and policies. Test the consequences of a change rather than assigning risk from its file extension or line count.
Interview exercise: release updated interview feedback
Suppose a coaching service changes its rubric and the model that generates feedback. The new release must improve useful guidance while preserving access rules, saved sessions and response-time limits.
Functional requirements
- Build a versioned candidate and identify all behavior-changing artifacts.
- Compare it with the deployed baseline on representative and critical cases.
- Package the release evidence and apply the repository's release policy.
- Expose real traffic gradually with explicit stop/expansion conditions.
- Restore a compatible prior release or contain the failure when needed.
Non-functional requirements
- Keep evaluation data and copied production traffic within their access rules.
- Bound test spending and keep a fast feedback path for ordinary changes.
- Prevent evaluation/shadow tools from causing duplicate production effects.
- Preserve reproducible build inputs and observable model configuration.
- Define recovery for in-progress workflows and state migrations.
Start with static checks, deterministic contract tests and a small maintained behavior suite. Add broader evaluation, load tests and staged exposure according to the change's risk. A full expensive suite on every typo is wasteful; a prompt change that enables a write tool deserves action-level tests.
What is a release?
A prompt edit can change which tools run. An index rebuild can change answers without any application code changing. A provider alias can move to another model. Therefore, record a release manifest with code, prompt, model/version/settings, tool schemas, policy, retrieval/index version, and evaluation/grader versions.
Some providers do not offer immutable versions forever. Record what is observable and monitor behavior; do not promise reproducibility beyond the provider's contract.
A practical pipeline
Read diagram source
flowchart LR
A[Versioned change] --> B[Static and contract tests]
B --> C[Behavior and risk evaluations]
C --> D[Release decision]
D --> E[Isolated shadow or canary]
E --> F[Expand or revert]
| Gate | Purpose | Examples |
|---|---|---|
| Static | Catch cheap defects | Template variables, schema, configuration |
| Contract | Protect system boundaries | Authorization, parsing, tool arguments, idempotency |
| Regression | Preserve useful behavior | Previously fixed cases and representative slices |
| Adversarial | Challenge severe risks | Cross-tenant retrieval, malicious tool output |
| Performance | Verify operating envelope | Queueing, latency tails, token/cost budget |
| Release | Assign accountability | Evidence, unresolved risks, rollback readiness |
Version the evaluation cases and graders separately from prompts under development. A development set used to tune the prompt is not an untouched holdout. Exact match fits identifiers, enums and deterministic contracts; semantic answers need criteria that accept multiple correct formulations. A phrase-presence test can pass an incorrect answer that merely repeats expected words.
Traditional services also have probabilistic tests and data dependencies. AI systems add more behavior-changing artifacts; the response is disciplined versioning and evaluation, not abandoning normal engineering tests.
Protect the release pipeline
Use trusted build inputs, scoped CI credentials and immutable artifact identifiers. Treat issue bodies and generated content as data, not shell source. Keep candidate code and model tools from rewriting the independent acceptance checks. Dependency/model changes should be visible in the review even when application code is unchanged. See framework migration discipline.
Make quality gates meaningful
Use hard gates for unacceptable known failures and baseline-relative tolerances for ordinary quality tradeoffs. A mean judge score above 4/5 cannot make a cross-tenant leak acceptable. Set thresholds before examining the candidate, and include uncertainty and important slices.
Repeated generation can vary even with temperature zero due to implementation and environment differences. Run sufficient trials for the decision. If a test is flaky, investigate whether the cause is the product, grader, environment, or a bad test. Do not delete every difficult case until the release passes.
Shadow and canary differ
Shadow: the candidate sees copied inputs but its outputs do not control the user experience. Writes must be simulated or isolated; otherwise two agents can both refund the customer. Apply privacy and retention policy to copied traffic.
Canary: real users receive candidate behavior. Start with appropriate users/tasks and an exposure budget. Observe enough volume and outcome delay to detect the relevant regression. A fixed 30-minute window is inadequate when harmful outcomes appear days later.
A/B experiments estimate product effects under a defined design. A safety canary is primarily about limiting and detecting harm; do not confuse the two purposes.
Rollback is a compatibility problem
Revert the complete compatible release, including prompt and tool contracts. A blue/green switch can make traffic routing fast, but it does not solve data compatibility or reverse external actions. A new index schema may not work with old code. Running workflows may need versioned execution paths. Test rollback and recovery before the incident.
Stop new consequential actions when a critical invariant fails; do not wait ten minutes because a generic threshold says to. Quarantine affected jobs, reconcile external effects, inform the appropriate owner, and preserve necessary evidence. A reverted deploy cannot unsend an email.
Assign operating ownership
Assign an approver, on-call owner, and exception process. Keep the release evidence concise and reviewable. Reserve time to maintain datasets and graders; a gate nobody trusts becomes a ritual. Measure escaped defects, false blocks, rollback success, lead time, and time to detect harmful regressions.
Quick recall
“Only the prompt changed. Full process?” Select tests by behavioral risk, not file extension.
“Offline scores improved but canary complaints rose?” Pause expansion, inspect traffic mismatch and failure slices, and compare actual outcomes.
“Can we automatically roll back every quality dip?” Only with validated signals and a safe compatible rollback; noisy graders can otherwise create instability.
See LLM evaluation and Google SRE canary guidance.
A release you can identify, gate, and roll back
Use a manifest whose fields refer to immutable artifacts. The values below are illustrative identifiers, not deployable production references.
release: support-2026-09-18-candidate
application_commit: example-commit-b
model_snapshot: evaluated-model-revision-b
prompt_digest: sha256:example-prompt-b
retrieval_snapshot: policy-index-19
tool_schema_version: refunds-v3
policy_version: returns-19
evaluation_dataset: support-holdout-12
grader_version: support-rubric-4
previous_release: support-2026-09-10
owner: support-platform
The following support-service example makes an action contract concrete. A deterministic contract test should cover a consequence, not merely compare a prompt string. This self-contained example checks that an approval cannot authorize a changed amount. The toy predicate omits authentication and persistence; those need their own integration tests.
def approval_matches(proposal, approval):
fields = ("order_id", "amount_cents", "currency", "proposal_version")
return all(proposal[k] == approval[k] for k in fields)
proposal = {"order_id": "o42", "amount_cents": 4000,
"currency": "USD", "proposal_version": 3}
approval = dict(proposal)
assert approval_matches(proposal, approval)
assert not approval_matches({**proposal, "amount_cents": 5000}, approval)
assert not approval_matches({**proposal, "proposal_version": 4}, approval)
| Result on the same 500 held-out tasks | Baseline | Candidate | Gate interpretation |
|---|---|---|---|
| Verified resolutions | 460/500 | 470/500 | Inspect paired changes and uncertainty; ten more alone is not the whole decision |
| Unauthorized effects in risk suite | 0 | 1 | Block release and investigate; an average cannot offset this |
| p95 completion latency | 3.0 s | 3.4 s | Pass only if the predeclared budget permits 3.4 s |
| Total cost/verified resolution | $0.040 | $0.038 | Include retries and review; compare after hard gates |
Define the gate before running: deterministic contracts pass, no unresolved critical-risk failure, required slices meet agreed bounds, and operational budgets hold. If all pass, deploy to a small persistently assigned canary and monitor real outcomes. The table's candidate does not ship because of the unauthorized effect.
For rollback, point traffic back to support-2026-09-10, verify its model and index artifacts still exist, and run a smoke test against its output/tool contracts. Keep backward-compatible database changes or a separately tested migration plan. Do not replay action-taking shadow traffic against production tools, and do not assume rollback reverses already-issued refunds. Existing workflows may need their original compatible release until they finish or migrate safely.
Interview questions with developed answers
Q1: How do you test prompt changes before production?
Sample answer: I version the prompt with its model, tool, policy, and retrieval configuration. I first validate template rendering and deterministic contracts, then run representative regression cases and targeted risk cases against the baseline. I inspect the changed behaviors and grade semantic outcomes with calibrated methods. For consequential actions, I verify external state and authorization. After offline gates pass, I use an isolated shadow where useful and a bounded canary with predefined stop conditions. The amount of testing follows the behavior the change can affect, not merely how many words changed.
Follow-up: Does a prompt-only change need code review? It can alter product behavior and authority, so it needs an appropriate review path.
Q2: What triggers should cause automatic rollback?
Sample answer: I define triggers from the service's consequences and measurable signals. A known cross-tenant exposure or unauthorized action may require immediate containment rather than waiting for an aggregate threshold. Operational regressions such as error rate, latency, and cost can have statistically and operationally appropriate windows. Quality signals may be delayed and need human review. I test that the rollback target is compatible, and I distinguish rolling back a release from disabling a risky action. There is no universal number of minutes of harmful output that is acceptable.
Follow-up: What if the old version is also affected? Use containment or a safe degraded mode instead of switching blindly.
Q3: How do shadow and canary deployments differ for agents?
Sample answer: A shadow observes copied inputs and produces results that do not affect the user or real external systems. Its tools must be isolated, especially writes. A canary receives real work and therefore needs bounded exposure, monitoring, and rollback authority. Shadowing helps compare behavior but cannot prove all production interaction effects. A canary tests more of the real experience but exposes users to risk. I choose the sequence based on the workflow and verify the isolation rather than assuming “not displayed” means “no side effects.”
Follow-up: Can read tools have side effects? Some do, such as marking an item read or triggering metering; inspect their contracts.
Q4: Why is a release manifest useful?
Sample answer: It records the combination of code, prompt, model settings, tools, policies, index, and evaluation versions that produced the tested behavior. When quality changes, we can identify which components changed and reproduce the relevant configuration. It also makes rollback meaningful: we return to a compatible combination, not just an old prompt file. For external provider aliases, I record the identifiers and evidence available and acknowledge limits to reproducibility. The manifest connects a deployment decision to the evidence supporting it.
Follow-up: Where do source-data versions belong? In the manifest or linked provenance so retrieval behavior can be explained.
Q5: What happens to long-running jobs during a deployment?
Sample answer: I define compatibility explicitly. Jobs may remain on an old workflow version, resume with compatible new code, or migrate through a tested process. I test representative stored histories and state formats before changing step order or business meaning. Approval records remain bound to their original proposals, and uncertain external effects require reconciliation. A deployment should not silently reinterpret a paused approval. The release plan includes owners for stuck or incompatible runs, not only a successful rollout of new workers.
Follow-up: Does reverting the deployment reverse completed actions? No; those need separate business remediation.
Final notes
I treat prompts, model settings, retrieval indexes, tools, policies, and graders as release artifacts alongside code. Fast deterministic tests catch contracts and permissions; representative evaluations measure behavioral changes and severe failure cases. A release must pass predefined gates, then move through isolated shadow testing or a bounded canary with monitoring and a rollback plan. I choose exposure and observation time from risk and traffic, not a fixed five-percent recipe. Rollback restores a compatible system configuration; it does not undo actions already taken.