What problem observability solves
A poor response does not, by itself, identify the component that failed. It may have been an unsupported generation, but it may also have been a retrieval query that returned poorly relevant documents, an external tool that returned an error, a validation step that accepted an incorrect output, or a retry that changed both behavior and cost. Logging only the final text and an error marker leaves too many hypotheses open.
Observability turns an interaction into a reconstructible sequence of technical facts. Its practical purpose is to answer, for a specific execution and for a population of executions, which version handled the request, which structured inputs it received, which steps it performed, how long each took, which resources it consumed, and what verifiable outcome it produced. That evidence makes it possible to distinguish an isolated incident from a regression following a change.
It does not replace pre-deployment evaluations or guarantee that a response is correct. Its role is to complement those practices with production signals. The supplied sources describe the need to connect model, prompt, stage-level timings, retries, tokens, cost, and quality when the use case requires it. This guide translates that framework into a minimum instrumentation contract without assuming that an aggregate metric alone explains root cause.
To navigate from this framework to related materials, the content architecture can link to [Learn](route:learn.index), [Compare](route:compare.index), and [Discover](route:discover.index). Those links do not replace the evidence collected in the trace of each execution.
The unit of analysis: an interaction trace
The minimum useful unit is one trace per business interaction, not one trace per isolated model call. It should begin when the application accepts an identifiable task—for example, answering a question or processing a request—and end when it delivers, rejects, or abandons an outcome. The trace has a stable identifier and contains child spans for its component stages.
A span represents a bounded operation: input normalization, retrieval, a model call, tool execution, a retry, validation, post-processing, or delivery. Each span retains its parent relationship, start time, end time, status, and specific attributes. This makes it possible to break down high total latency without automatically assigning the delay to the model provider.
The trace should be associated with a pseudonymized request or conversation identifier, a channel, a task type, and a deployment cohort. Those attributes should not be used to store free text or direct identifiers of people. The purpose is to segment a degradation by version, task, environment, or channel without re-identifying the person who used the service.
Reconstruction must include immutable references to the versions used: the model and relevant parameters, the prompt template, retrieval configuration, tool definitions, validation rules, and workflow version. If an identifier points to mutable content, a later investigation may reconstruct a configuration different from the one that produced the incident.
The minimum telemetry contract
Define the contract before instrumenting. At trace level, record an identifier, time, environment, task type, channel, cohort, workflow version, final status, and a business outcome where one exists. At span level, record operation type, status, duration, attempt number, dependencies, and attributes that make configuration comparisons possible. Use controlled vocabulary for statuses such as success, technical error, validation error, safe rejection, cancellation, and timeout.
For a model call, minimum attributes include the provider or model family, resolved version or alias where available, parameters that materially alter the output, the template identifier, input and output tokens, and calculated cost or enough data to calculate it under the applicable pricing. For retrieval, include the versioned index or collection, strategy, filters, candidate count, and selected documents through non-sensitive identifiers.
Tools need a name and version, requested operation, result code, error class, duration, and an idempotency or correlation key if they perform actions. For validations, record the rule name and version, outcome, and a failure category. Do not confuse syntactically valid JSON with a response that is acceptable for the business: they are different signals.
The contract must document what is excluded. By default, avoid prompts, responses, retrieved documents, complete tool arguments, email addresses, phone numbers, addresses, secrets, access tokens, and any data that is not essential for diagnosis. When text is necessary for authorized debugging, apply minimization, redaction, access controls, and differentiated retention.
Fields and logging decision
| Field | Diagnostic use | Recommended handling |
|---|---|---|
| trace_id and span_id | Reconstruct the sequence | Log |
| Model, prompt, and workflow version | Compare changes | Log |
| Tokens, duration, and status | Measure cost and performance | Log |
| Retrieved document ID | Review relevance | Log without content by default |
| User text or response | Analyze individual cases | Exclude or minimize; restrict access |
| Credentials and secrets | Provide no legitimate diagnostic value | Do not log |
How to measure real cost per interaction
Cost per interaction is not equivalent to the average cost of one primary call. Add model input and output calls, retries, tool calls with their own billing, retrieval where it generates attributable spending, and failed steps. Keep observed cost separate from an estimate: the former comes from usage data and applied prices; the latter may depend on rates, rounding, or incomplete information.
Attribute every component to the same trace and retain the currency, calculation date, and version of the price table or estimation method. This matters because a pricing, model, or processing-mode change may mean that two periods are not directly comparable. If no verifiable cost exists for a component, mark it as unknown rather than assigning zero.
Measure distributions, not only averages. A stable average can conceal a tail of interactions with multiple retries or excessive context. Segment by task, channel, version, and final outcome. The cost of an execution that ends in error is still operating cost and should appear both in the total and in waste analysis.
When human review is involved, it is preferable to record it as a separate cost or operating effort with an explicit estimation method. Mixing it with inference spending can hide the fact that a latency optimization shifted work to people.
Auditable calculation of cost per trace
- 01Group all child spans by trace_id, including those that ended in error or cancellation.
- 02Add token costs for each call using the applicable rate and date; retain the calculation source as an internal attribute.
- 03Add attributable tool and retrieval costs without replacing unknown values with zero.
- 04Separate inference cost, attributable infrastructure, and estimated human-review cost.
- 05Publish the total, components, percentage of failed executions with cost, and percentiles by cohort.
Separating sources of latency
End-to-end latency is the time perceived by the person using the application, but it does not identify the component that needs correction. Record spans for queueing or admission, context preparation, retrieval, model calls, network time where observable, tools, retries, validation, and output serialization. Their sum may not exactly match the total if operations run in parallel; for that reason, the temporal relationship between spans must also be recorded.
Compare duration percentiles by stage, not only the total average. An increase in the high percentile for tools while model latency remains stable directs the investigation toward an external dependency. An increase in retrieval time may come from an index, a filter, or candidate growth. A slow response caused by retries requires reviewing both the condition that triggers them and the limit policy.
Avoid simplistic attribution. The fact that a trace contains a model call does not demonstrate that the model was the bottleneck. Appropriate evidence is a time distribution segmented by the same version, task type, and comparable conditions. Changes in traffic, content, or user mix are confounding factors that should be noted.
Production quality signals and their limits
Quality should not be reduced to one score. User feedback reflects perceived experience, but it may be sparse, biased toward extreme cases, or lack context. Human review can apply defined criteria and identify subtle failures, although it has a cost and limited coverage. Deterministic rules are reproducible for verifiable requirements, such as a schema or authorization, but they do not independently capture usefulness, factual accuracy, or contextual appropriateness.
Automated evaluators can help prioritize samples and track trends if they are versioned, calibrated against human review, and used with explicit limits. They do not constitute independent proof of truth, especially when evaluating ambiguous tasks or sharing biases with the system being evaluated. Record their version, available input, criterion, outcome, and confidence level where the method produces one.
Link every signal to the trace and clearly distinguish its origin. A decline in thumbs-up ratings is not equivalent to a breached business rule; output with valid JSON does not demonstrate that its values are correct. The dashboard should make each signal visible separately and then allow examination of agreements or divergences.
Sample cases without feedback in addition to negative ones. Otherwise, the team learns about people who report problems but not about silent failures. The sample design should be documented: population, period, strata, size, and review criterion.
Interpreting quality signals
| Signal | What it contributes | Main limitation |
|---|---|---|
| User feedback | Perceived experience | Coverage and response bias |
| Human review | Contextual judgment using a rubric | Cost and reviewer variability |
| Deterministic rule | Reproducible compliance | Covers only defined conditions |
| Automated evaluator | Monitoring and prioritization | Requires calibration and can be wrong |
Diagnosing four common incidents
An invented response should be investigated backward from the outcome. Check whether the task required grounding, whether retrieved context was available, which documents were selected, which source-use instructions applied, and whether a rule or review identified unsupported claims. If retrieval configuration or prompt version was not recorded, the cause may remain indeterminate; do not attribute the failure to the model by elimination.
For irrelevant context, compare the normalized query, filters, versioned collection, strategy, candidate count, and final selection. The problem may lie in indexing, filters, metadata, corpus changes, or ranking strategy. Low relevance can also originate because the request was classified into the wrong task before information retrieval.
For invalid JSON, separate syntax from semantics. Review the required schema, structured-output method, parser version, retries, and validation response. A successful retry can hide a growing rate of invalid first responses while increasing cost and latency.
For a failed tool action, determine whether the tool was called, whether it received permitted arguments, whether the dependency responded, whether a timeout occurred, and whether there were partial effects. Actions with consequences should incorporate idempotency keys, confirmation states, and retry limits. A satisfactory final response does not prove that the action was executed.
Incident investigation routine
- 01Bound the incident by period, cohort, task, and outcome; retain the trace identifier for a representative sample.
- 02Compare affected traces with a comparable earlier or control group without mixing different channels or tasks.
- 03Locate the first anomalous span and review its versions, status, duration, retries, and dependencies.
- 04Test the hypothesis against quality signals, business rules, and tool outcomes.
- 05Apply a reversible mitigation, verify its effect in the cohort, and document the evidence and remaining uncertainties.
Alerts, thresholds, and stop decisions
A useful alert specifies a metric, window, segment, threshold, owner, required evidence, and initial action. “Quality is down” does not meet that standard. An operational formulation could monitor an increase in validation errors for a specific version during a defined window, require sample traces, and assign review to the workflow owner.
Use thresholds as attention rules, not as proof of causality. They should be based on the service’s own baseline and reviewed when volume, task mix, or product changes. Combine availability and security alerts with quality and cost review: optimizing one metric in isolation can worsen another.
Stop, roll back, or limit a change when it breaches a safety condition, a critical business rule, or an agreed economic limit, or when the evidence shows a material degradation against a comparable cohort. In other cases, reduce exposure through gradual deployments and increase sampling before drawing a conclusion.
The supplied sources recommend relating performance, cost, and quality and comparing versions or environments. They do not establish a universally applicable choice of thresholds; those must be derived from risk, baseline, and service commitments.
Alert template
| Case | Metric and window | Initial action |
|---|---|---|
| Abnormal cost | Cost per trace and high percentile by version | Limit cohort and review retries |
| Invalid output | Validation failure rate by workflow | Roll back parser or configuration if it affects the contract |
| Failed tool | Errors and timeouts by dependency | Enable safe degradation and review partial effects |
| Degraded quality | Breached rules and comparable human sample | Pause expansion and compare with control |
Retention, minimization, and access
A detailed trace can become a sensitive repository if it is designed without limits. Start with the diagnostic question and record only the attributes needed to answer it. Pseudonymized identifiers, appropriately managed hashes, and error categories often provide more operational value than indiscriminately storing prompts, responses, or complete documents.
Separate operational data from exceptional debugging data. The former may include durations, versions, counters, statuses, and technical references; the latter, if justified, requires limited access, redaction, access logging, and defined short retention. Also review which data instrumentation libraries send to third parties before enabling them.
Define who can view traces, who can access exceptional content, and who can modify redaction or retention rules. Deletion requests, regulatory requirements, and internal policies may vary by jurisdiction and use case. This guide does not determine legal obligations; they require an assessment appropriate to the organization’s context.
The supplied New Relic documentation mentions filters for excluding sensitive data before it is sent. That supports the technical feasibility of filtering, but it does not demonstrate that a specific configuration is sufficient for every data category or every regulatory framework.
Final launch template and minimum dashboard
Before releasing a version, verify that each execution can be linked to an immutable version of the workflow, prompt, model, retrieval, tools, and validations. Confirm that retries appear as distinct spans or attributes, that tokens and costs are attributed to every attempt, and that business outcomes are not confused with technical errors.
The minimum dashboard should combine trace volume, final-status rate, total and stage-level latency, tokens and cost per interaction, retries, validation outcomes, tool errors, and quality signals separated by source. Every chart should support filters by period, environment, version, task, channel, and cohort to avoid comparisons across heterogeneous populations.
For the release, select a control cohort or earlier baseline and define expansion, pause, and rollback conditions in advance. Retain representative trace samples, including successful, failed, and high-cost executions. Document traffic, corpus, pricing, or policy changes that may affect interpretation.
The expected outcome is not an automatic explanation for every incident. It is a system that reduces dependence on memory, screenshots, and isolated impressions, and explicitly shows when the evidence supports a conclusion and when it does not yet do so.
Launch checklist
- 01Assign immutable versions to the model, prompt, retrieval, tools, validation, and workflow.
- 02Run test traces covering success, retry, tool failure, validation rejection, and cancellation.
- 03Check that cost, tokens, and latency are broken down by stage and retained for failed attempts.
- 04Verify sensitive-data filters, access permissions, and the retention policy.
- 05Define a control cohort, alert owners, reviewable thresholds, and a rollback condition.
- 06Review a human sample and document uncertainties before expanding deployment.
Open questions
- The supplied sources are largely editorial guides or implementation cases; they do not establish a universal standard for schemas, thresholds, or retention.
- No independent comparative data is provided to set specific alert values, quality targets, or acceptable costs.
- The Buk case URL contains a future date relative to some possible publication contexts; it is used only as supplied material about the practices described, not to infer timeliness or general adoption.
- Privacy, security, and retention obligations depend on jurisdiction, processed data, and organizational context, and cannot be determined from the supplied sources.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction