Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

AI Anti-Patterns: Diagnose the Assumption Before Replacing the Design

By Anup Rai10 min readReviewed September 2026

An anti-pattern is a recurring approach that appears useful but produces harmful consequences in a particular context. The lesson is not “never use a long prompt” or “always add agents.” Identify the failed assumption, demonstrate the consequence, and compare a repair with the cost of leaving the design alone.

In a Learnastra design review, use four questions: What is the symptom? Which assumption failed? What mechanism repairs it? When would the original approach still be reasonable? This is a recall aid, not a formal industry standard.

1. Recognize the failure, not the fashionable technology

Anti-pattern Observable symptom Repair When the original choice can work
Prompt grows without a behavior contract Contradictory instructions; fixes break unrelated tasks Prioritize requirements, version the prompt, test interactions A long but coherent prompt for a bounded task
Every task becomes an agent Simple requests trigger planning, tools and retries Start with one call or a fixed workflow Dynamic steps are genuinely necessary
Retrieve more whenever quality drops Relevant facts drown in duplicates and stale evidence Diagnose indexing, retrieval, packing and generation separately Additional evidence actually resolves a known gap
Use the largest model for every request Higher cost without a measured outcome gain Compare task-specific configurations The stronger configuration consistently meets a difficult requirement
Fine-tune before diagnosing errors Training memorizes facts that soon change Separate missing knowledge from missing behavior Repeated behavior errors and suitable training data justify adaptation
Depend on a provider without a failure plan Outage becomes an unbounded queue Define deadlines, degradation and an exit strategy One validated provider may be appropriate for a small product
Add persistent memory by default Stale or sensitive facts reappear in unrelated sessions Define scope, correction, deletion and retention Multi-turn work needs durable preferences or task state
Omit required state Agent forgets completed actions and repeats them Persist operation IDs, task status and necessary constraints Independent read-only requests can remain stateless

Each repair adds obligations. Splitting an unclear prompt into five agents distributes the ambiguity and adds coordination. A second provider helps only if its capability, policy and capacity are validated. Persistent state improves continuity while increasing privacy and correctness responsibilities.

See design patterns for the corresponding mechanisms and their selection criteria.

2. Find the broken stage before adding components

Architecture / visual model
flowchart TD A[Production failure or representative test] --> B[Locate the failing stage] B --> C[State a falsifiable cause] C --> D[Smallest repair and expected tradeoff] D --> E[Compare with the baseline] E --> F{Outcome and safety criteria met?} F -->|Yes| G[Canary rollout and monitor] F -->|No| H[Reject repair or revise diagnosis] G --> I{Regression or cost breach?} I -->|Yes| J[Rollback and preserve evidence] I -->|No| K[Keep and document decision] H --> B
Read diagram source
flowchart TD
    A[Production failure or representative test] --> B[Locate the failing stage]
    B --> C[State a falsifiable cause]
    C --> D[Smallest repair and expected tradeoff]
    D --> E[Compare with the baseline]
    E --> F{Outcome and safety criteria met?}
    F -->|Yes| G[Canary rollout and monitor]
    F -->|No| H[Reject repair or revise diagnosis]
    G --> I{Regression or cost breach?}
    I -->|Yes| J[Rollback and preserve evidence]
    I -->|No| K[Keep and document decision]
    H --> B

Example: the assistant gives an outdated cancellation rule.

  1. Verify whether the current policy entered the document store.
  2. Check the indexed revision and deletion/update propagation.
  3. Inspect authorized retrieval results for the exact request.
  4. Check whether context packing discarded the current exception.
  5. Check whether generation ignored present evidence or cited it incorrectly.
  6. Repair the failing stage and test old/new policies together.

A new embedding model does not fix a document that was never ingested. A larger context window does not resolve conflicting policy versions automatically. A new answer format does not make the answer current.

3. The invisible data pipeline

A query diagram that assumes clean, current and authorized documents omits much of the real system. Include ingestion, structural parsing, source revisions, permissions, deletion, deduplication and freshness monitoring.

Before and after: chunking without a contract

Before: split every 500 characters and store only {text, vector}. A table's “14 days” separates from “refurbished,” and the application cannot identify chunks belonging to a deleted source.

After: preserve useful structural units and store a record such as:

{
  "source_id": "returns",
  "revision": 19,
  "section": "refurbished",
  "span": [420, 690],
  "acl_version": 8,
  "parser_version": 3,
  "embedding_version": "embedding-config-7",
  "text": "...authorized source passage..."
}

Define the span's coordinate system, such as Unicode code points in a specific normalized source revision. Preserve a required table header or exception. Test that the answer remains recoverable after chunking, and recheck access when expanding surrounding context. Metadata makes lifecycle handling possible; it does not enforce permissions by itself.

Cost/benefit: metadata and version tracking add ingestion/storage work but make updates, deletions, debugging and reproducibility possible. The exact chunk size remains a workload choice, not a universal best practice.

4. Prompts and JSON are not authority boundaries

Common failures include vague tool descriptions, a prompt that claims to prohibit all injection, and accepting anything that parses as JSON.

Before: “Return JSON with the refund amount.” The executor accepts {"amount": "all of it", "skip_approval": true}.

After: validate a versioned schema using integer cents, bounded values and forbidden extra fields. Then check account ownership, policy, current refundable balance, approval and operation identity. A schema-valid request can still be unauthorized or stale.

The guardrails walkthrough separates structure from business rules. Limit repair attempts. Validation errors should describe the allowed contract without revealing another user's records or secret configuration.

Tempting shortcut Why it fails Concrete boundary
“The system prompt forbids disclosure” Retrieved content can influence generation Permission checks before retrieval/tool execution; scoped data
“The critic approved the action” A model judgment is not authorization Trusted policy decision and current business-state check
“A human will catch mistakes” Review can be rushed or unstaffed Clear review evidence, capacity, decision authority and escalation
“The sandbox makes every tool safe” Allowed network calls can still leak data or cause effects Tool allowlists, egress limits, scoped credentials and action approval
“Valid JSON means a correct answer” Syntax does not establish truth or entitlement Semantic validation and evidence checks

Use tools with narrow operations and explicit errors. Avoid a general shell or arbitrary SQL executor when a constrained read API satisfies the requirement. Where broad tools are necessary, isolation, resource limits and authorization must match their actual power.

5. Retry-until-success and unlimited waiting queues

A loop that keeps trying can amplify an outage, spend without progress, or repeat a side effect. Bound elapsed time, steps, tokens, retries and parallelism outside the model. Detect repeated ineffective actions and return an honest partial result or handoff.

A timed-out refund may already have succeeded. Starting a new workflow with a new payment key can duplicate it. Preserve the business operation ID across retries and reconciliations. A durable checkpoint does not undo an external action. A refusal or authorization failure should not trigger a search for a provider willing to perform the prohibited action.

Worked capacity example

Assume a service receives 20 requests/second and each occupies a worker for an average of five seconds. Under stable conditions, Little's law gives an average of 20 × 5 = 100 requests in service. This is a sizing clue, not a p95 latency guarantee.

If only 50 workers can run, their simplified maximum completion rate is 50 / 5 = 10 requests/second. At an arrival rate of 20/second, the backlog grows by roughly ten per second while those assumptions hold—about 600 waiting requests after a minute.

Proposed fix Benefit Remaining cost or limitation
Semaphore of 50 Caps active provider calls Does not cap the queue or waiting memory
Bounded queue with admission control Limits waiting work and fails promptly Some callers must retry later or accept rejection
Per-tenant limits and fair scheduling Prevents one tenant exhausting capacity Scheduling complexity; spare capacity policy needed
More workers Raises throughput when downstream capacity exists Spend, provider quotas and database capacity may become limiting
Faster acceptable path Reduces occupancy Must preserve the required outcome and authorization

Use measured service-time distributions and load tests for burst behavior. A rate limit caps starts over time; a concurrency limit caps simultaneous work. They are complementary. Retrying rejected requests immediately can recreate the same overload.

6. The benchmark-only release

A demo proves that some examples work. It does not estimate representative production quality. A single benchmark average can conceal rare permission failures, poor language/domain slices, contaminated cases and a judge that rewards verbosity.

Before and after: a random split that leaks incidents

Before: randomly split 10,000 support messages so turns from the same incident and near-identical templates appear in both training and test.

After: group related incidents before splitting, and use a time boundary if future behavior is the target. Check source and near-duplicate overlap across splits. Tune on development data; preserve an independent test set for the release decision. An illustrative 70/15/15 allocation is only a split ratio—it does not establish independence or sufficient rare-case coverage.

  1. Define the user outcome and unacceptable failure modes.
  2. Build representative cases plus targeted adversarial and rare-risk cases.
  3. Separate task-quality scores from permission and side-effect invariants.
  4. Calibrate model graders against human labels and objective checks.
  5. Report uncertainty, disagreement and failure slices.
  6. Evaluate the final configuration, then canary and monitor outcomes.

Repeatedly inspecting a holdout and tailoring the implementation to it turns it into development data. Replace or refresh the independent evaluation accordingly. Detailed methods are in evaluation frameworks.

7. The fake savings calculation

Count unsuccessful calls, retries, tool charges, review and infrastructure. “The token price is half as much” does not imply the task costs half as much.

Suppose 1,000 tasks use $0.01 of model calls each. If 10% need a two-minute review and reviewer time costs an illustrative $30/hour:

  • Model cost: 1,000 × $0.01 = $10.
  • Review time: 100 × 2 / 60 = 3.33 hours.
  • Review cost: 3.33 × $30 = $100.
  • Combined modeled cost: $110, before infrastructure, failures and support.

A second configuration costs $0.03/task but needs the same review on only 2% of tasks: $30 + $20 = $50. This comparison assumes equal final quality, review effectiveness and no other cost differences. Measure those assumptions; do not treat the hypothetical as a vendor forecast.

A global token counter sampled before and after each request also misattributes overlapping requests. Record usage per call, link calls to the task, and reserve shared budget before parallel work starts. Alerts explain overspend after it happens; admission controls prevent new work from exceeding a reservation policy.

8. Unsafe caching and needless cache avoidance

Caching every answer risks serving stale policy or another user's private result. Never key a personalized answer by question text alone. Include the relevant authorization scope, data revision and model/prompt configuration, and define invalidation.

Semantic similarity is not equivalence. Negation, dates, units and requested actions can reverse meaning. A threshold cannot guarantee safe reuse across embedding models and workloads. See semantic caching.

Conversely, refusing all caching can repeatedly compute stable public reference material with no benefit. Cache safe immutable artifacts or retrieval intermediates when the personalized final answer cannot be reused. Measure hit rate and the total cost of lookup, validation, storage and invalidation.

9. Organizational failures belong in the design

Missing responsibility Failure Required operating decision
Evaluation owner Every team uses a different success definition Version the rubric and name the release decision-maker
Exception-queue staffing “Human in the loop” becomes an indefinite wait Capacity, working hours, deadline and escalation
Data lifecycle owner Deleted or obsolete evidence remains searchable Update/delete SLO and reconciliation
Incident recovery owner Engineers stop the model but leave user harm unresolved Business-state repair and user communication procedure
Maintenance budget Prompts, policies and providers drift without review Scheduled reassessment and regression ownership

Do not invent production stories or quantified savings to sound experienced. In an interview, distinguish a hypothetical design from your own work and explain what evidence you would collect.

Interview practice

1. Is a single provider always an anti-pattern? No. One validated provider may fit the product's risk and team capacity. The failure is ignoring dependency risk. Define degradation, monitoring and an exit plan; add failover when its benefit exceeds compatibility and operating costs.

2. How do you distinguish missing memory from a reasonable stateless design? Ask whether future steps require facts or operation identity from earlier work. Persist necessary state with scope and retention. Independent restartable requests may need none.

3. How do you challenge a complex architecture constructively? Tie each component to a requirement, inspect evidence of benefit, and propose a smaller experiment. Name the replacement's costs as well. A working complex system may not justify migration just to reduce box count.

4. A bigger model improves the benchmark but doubles review effort. Ship it? Compare representative final outcomes and full cost per success. Inspect why review grew: uncertainty, verbosity or new high-risk mistakes may explain the discrepancy. Benchmark improvement alone is insufficient.

5. Is “use RAG before fine-tuning” a universal sequence? No. Retrieval supplies external evidence; fine-tuning changes learned behavior. Diagnose the problem and evaluate the relevant intervention. Some tasks need both, neither, or an authoritative API.

6. Why does a successful retry not prove recovery? The original write may also have succeeded. Verify the final business state and operation identity, then reconcile duplicates or unknown outcomes. Infrastructure success and user recovery are different outcomes.

Final summary and review notes

Remember the missing contract: behavior, data, authority, execution budget, evaluation or ownership. Name which one failed before prescribing a framework.

A strong closing answer identifies a concrete symptom, a testable cause, a repair and its cost, a counterexample where the original choice is reasonable, and a rollback condition. Improvements should survive representative tests and production measurement—not only a persuasive diagram.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← AI Design Patterns: Choose the Mechanism That Solves the Failure
NEXT LESSONDesign an Enterprise Knowledge Assistant →

Explore the diagram