Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

LangSmith Observability

By Anup Rai8 min readReviewed September 2026

Observability is the ability to understand a system's behavior from its emitted telemetry. In an AI application, useful telemetry connects the user's task to retrieval, model calls, tool actions, state transitions, failures and outcomes. LangSmith provides tracing, debugging and evaluation tooling for such applications; it is usable beyond LangChain applications. LangSmith observability.

The goal is to answer a concrete question: Where did this run go wrong, what evidence supports that diagnosis, and did the proposed fix improve the result? A dashboard of token counts alone cannot answer it.

Separate signals and their purposes

Signal Example Useful question
Metric Request count, error rate, latency distribution Is the system meeting its target?
Trace Parent run with retrieval/model/tool child operations Which steps contributed to this request?
Log/event Timeout category or approval-state change What specific event occurred?
Evaluation result Evidence support or task-success score Was the observed behavior acceptable?
Business outcome Refund ledger or completed exercise record Did the intended external result happen?

A trace captures instrumented execution, not private internal model reasoning. An explanation in a model output is not guaranteed to describe the actual cause of its decision. Use observable inputs, tool results and outcome records for diagnosis.

Design the telemetry contract

For a documentation assistant:

Functional requirements

  1. Correlate the request with retrieval, generation and validation steps.
  2. Identify document, prompt, code and model versions.
  3. Record known failure categories and incomplete/cancelled runs.
  4. Attach evaluation and user-feedback results to the relevant run.
  5. Turn selected failures into reviewed regression cases.

Non-functional requirements

  1. Exclude secrets and prohibited customer content before export.
  2. Bound telemetry overhead, sampling and retention cost.
  3. Keep trace access scoped to authorized reviewers.
  4. Preserve useful signals during exporter/network failure.
  5. Distinguish sampled observations from whole-population metrics.
Architecture / visual model
flowchart LR A[Application request and child operations] --> F[Allowlist, mask and sample telemetry] F --> T[Trace collector and searchable runs] A --> M[Aggregate metrics and outcome records] T --> D[Diagnose a concrete failure] M --> D D --> R[Curate permitted regression cases] R --> E[Evaluate baseline and candidate versions] E --> G[Release decision under quality and cost gates] G --> P[Bounded production rollout] P --> T
Read diagram source
flowchart LR
    A[Application request and child operations] --> F[Allowlist, mask and sample telemetry]
    F --> T[Trace collector and searchable runs]
    A --> M[Aggregate metrics and outcome records]
    T --> D[Diagnose a concrete failure]
    M --> D
    D --> R[Curate permitted regression cases]
    R --> E[Evaluate baseline and candidate versions]
    E --> G[Release decision under quality and cost gates]
    G --> P[Bounded production rollout]
    P --> T

LangChain integrations can generate useful trace structure once configured. Custom code and other SDKs may need explicit instrumentation. Verify parent/child correlation across asynchronous work and queues; do not assume every node and background action appears automatically.

Record enough to diagnose a run

Field Reason to retain it Boundary
Run/trace and parent IDs Connect the execution tree IDs alone should not grant access
Application/prompt revision Compare releases Avoid secrets embedded in prompts
Model/provider and settings Identify behavior/cost differences Label unavailable usage rather than inventing it
Retrieval IDs and corpus revision Reproduce evidence selection Content access must remain controlled
Tool identity and outcome category Separate rejection, failure and uncertainty Redact credentials and sensitive arguments
Timing and usage Attribute resource consumption Distinguish inclusive/exclusive time
Pseudonymous account/cohort metadata Diagnose scoped behavior Minimize personal identifiers
Final business outcome reference Verify the task's effect Read the authoritative system when needed

Use trace names and metadata with bounded cardinality for aggregation. Put unique request IDs in correlation fields rather than creating one metric label value per user or request. Retain the mapping only where authorized.

Read a trace before changing the prompt

Suppose the assistant cites an obsolete API parameter. Investigate in order:

  1. Did the request resolve to the correct product and version?
  2. Did retrieval include the updated document?
  3. Did filtering or reranking discard it?
  4. Did context assembly include its relevant section?
  5. Did the model contradict the supplied evidence?
  6. Did citation/output validation catch the discrepancy?

If the current document never reached retrieval, adding “be accurate” to the prompt is unlikely to repair the cause. Conversely, a correct retriever can still feed evidence that the model misuses. This is failure localization, not proof of every internal causal mechanism.

Inspect failures and successes together. A fix that improves one memorable trace can regress a broader query class. Keep the problematic case as evidence, then evaluate a representative set.

Attribute latency without double-counting

Illustrative request timings:

Stage Duration
Admission/queue wait 50 ms
Retriever A 80 ms
Retriever B, parallel with A 120 ms
Merge/context assembly 10 ms
Model response 700 ms
Output checks 40 ms

Ignoring other overhead, the request takes 50 + max(80, 120) + 10 + 700 + 40 = 920 ms. Summing both parallel spans gives 1,000 ms and overstates elapsed time. Adding a parent span to its children double-counts again.

Distinguish time to first token from completion time. Streaming can improve perceived responsiveness while total generation remains long. Averages also hide tail failures; inspect latency distributions by task and release, without calculating end-to-end p95 by simply summing stage p95 values.

For cost, include retries, failed calls, embeddings, tools, cached input, evaluation and infrastructure where applicable. Reducing eight equal-cost model calls to five reduces that model-call component by 37.5%. The total bill can fall by a different percentage because other costs remain or change. In real runs, token counts, models and tools often differ too.

Protect telemetry before it leaves the process

LangSmith supports hiding or transforming inputs, outputs and metadata, and selectively disabling tracing. Configure those controls for the actual SDK and data policy. Sensitive-data controls.

  1. Prefer an allowlist of useful fields over logging every object.
  2. Remove credentials, session tokens and prohibited document content before export.
  3. Check nested metadata, exception strings and tool outputs as well as prompts.
  4. Restrict project access and define retention/deletion procedures.
  5. Test the emitted payload using representative sensitive fixtures.

A trace-sampling rate is not a privacy control: sampled traces can still contain sensitive data. Hiding inputs does not automatically remove the same information from metadata or errors. Do not route production conversations into permanent datasets by default.

Use sampling without misleading yourself

LangSmith trace sampling can reduce recorded trace volume. Choose sampling deliberately and preserve the fields needed to understand its bias. Separately instrument aggregate request/error metrics when full-population rates are required.

For illustration, sampling 5% of 100,000 requests/day yields about 5,000 root traces. At 12 spans per traced request and 2 KB per span, that is approximately 120 MB/day of raw span payload, before indexing, replicas and retention. Root traces and child spans are different billing/storage units; inspect the actual service plan instead of assuming one request equals one stored item.

Rare failures may need additional collection under a supported error/tail-sampling or application policy. If all errors are retained but only 5% of successes are retained, the error percentage in the trace list is not the production error rate. State the sampling rule when interpreting dashboards.

Turn traces into evaluations carefully

LangSmith evaluation supports datasets, evaluators and experiments. Offline evaluation compares versions against prepared examples; online evaluation examines selected production behavior. Neither automatically makes its labels correct.

Evaluator Good use Limitation
Deterministic check Schema, exact IDs, recomputed totals Only checks encoded properties
Reference-based check Known correct facts or task results References may be incomplete or stale
Model grader Semantic criteria difficult to encode Bias, inconsistency and prompt injection
Human review Ambiguous outcomes and rubric calibration Time, disagreement and expertise

A larger model is not automatically a superior judge. Calibrate against qualified human labels, define the rubric, test adversarial answers and preserve disagreement. A tool-action evaluation should inspect authorization and the actual effect, not just polite wording in the final message.

Curate a regression dataset:

  1. Select representative failures and successful controls.
  2. Remove or transform disallowed content.
  3. Add required outcomes, constraints and necessary context.
  4. Review labels and record dataset versions.
  5. Keep tuning cases separate from held-out evaluation cases.
  6. Run enough repeated trials to characterize variability for the decision.

Negative user feedback is a useful signal, but not automatically a correct answer label. The request may be ambiguous, the feedback may concern interface latency, or the submitted text may be adversarial.

Connect evaluation to release controls

The application's release system routes canary or A/B traffic, applies stop criteria and rolls back. An observability experiment does not automatically provide those controls for every deployment.

For an A/B comparison, define the unit of assignment—often account or conversation, not arbitrary individual turns—and keep it stable where carry-over matters. Compare task mix, quality, cost and latency over an adequate sample. A universal 2% canary fraction or a single unexplained failure threshold is not a rollout design.

Set severe-failure gates separately from average quality. A small average improvement does not compensate for exposing another customer's data. Missing evaluator results or exporter failures should be visible, not counted as successful evaluations.

Interview practice

Q1: What does a trace tell you that metrics do not?

It connects the steps of a specific request, helping localize a failure or bottleneck. Metrics reveal population behavior. Use both, plus authoritative outcomes for external actions.

Q2: Can you see the model's full reasoning through LangSmith?

No. You see instrumented calls, observable messages, tool results and state. Generated explanations are not guaranteed faithful accounts of hidden computation.

Q3: Why not copy every negative-feedback trace into a dataset?

It can contain private data, incorrect labels or unrepresentative failures. Curate, redact, label and version the cases before using them for training or evaluation.

Q4: How can tracing make the dashboard's error rate misleading?

Biased sampling, missing child spans and excluded failed exports can change the observed population. Use reliable aggregate counters and state the trace sampling policy.

Q5: Does the evaluation product deploy and roll back the application?

Only if the chosen deployment integration explicitly implements that behavior. Otherwise the release system must consume evaluation results and enforce rollout/rollback policies.

Q6: How do you justify observability cost?

Measure time to diagnose incidents, avoided repeat failures and useful resource optimizations against telemetry, evaluator and operating costs. Do not claim a fixed subscription saving or API-cost reduction without workload evidence.

Final notes

Recall card: Instrument → protect → correlate → diagnose → curate → evaluate → control release. A useful trace leads to a testable explanation and a measured improvement.

Next: LlamaIndex.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← LangGraph Orchestration
NEXT LESSONLlamaIndex: document retrieval and event-driven workflows →

Explore the diagram