Skip to main content
ZICQ

Wiki Infrastructure

Observability

Infrastructure
Aliases: Observability Tracing LLM Observability 可观测性 ·2026-10-07

Observability

Observability uses runtime information to understand system behavior. For LLM applications, it helps locate problems in retrieval, model calls, tool execution, or post-processing. OpenTelemetry introduces traces, metrics, and logs.

Signal Question Example
Metrics How often, costly, or slow? Error rates, token use, latency distributions
Traces Which steps ran for one task? Retrieval → model → tool → response
Logs What happened at a step? Timeout, validation failure, authorization denial

Engineering recommendations

Correlate steps with a request identifier. Record model and prompt versions, timing, status, token use, and retries. Track tool outcomes and document versions.

Avoid recording every raw prompt or parameter by default. Redact personal data and credentials; restrict access, retention, and sampling. Establish an authorized scope before retaining sensitive reproduction samples.

Performance is not correctness

A fast, error-free request may still cite outdated evidence. Link runtime failures to evaluation categories: missing evidence, unsupported claims, execution failures, or unmet requirements.

GenAI field conventions evolve. Use the version supported by your implementation; the current entry point is the GenAI semantic-conventions repository.

See tool use, RAG, agent memory, and inference.