LLM observability for SaaS: traces, metrics and privacy
Instrument AI features with useful traces and metrics for latency, tokens, tools and failures while keeping prompts, customer data and tenant identifiers protected.
In this guide
What should LLM observability measure?
LLM observability connects a customer request to the model calls, retrieval steps, tool operations and application decisions that produced the result. Start with traces for one request and metrics that explain reliability, latency, usage and outcome. Capture enough metadata to locate a slow or failed step, but do not default to storing full prompts, completions or retrieved documents: AI telemetry can contain sensitive customer data and become a separate privacy and access-control risk.
Trace each request across model, retrieval and tool boundaries
Create a parent trace for the user-visible operation and child spans for model calls, retrieval, reranking, moderation, tools, retries and post-processing. Record duration, outcome, model or provider identifier, token usage when available, retry count and a correlation ID. This lets an operator distinguish slow retrieval from provider latency or a tool loop without guessing from one total request time.
Track service health and product outcomes separately
Use metrics for request volume, error and timeout rates, latency distributions, cancellations, token and cost estimates, queue depth, refusal rates and quality or user-feedback signals. Separate technical success from task success: an HTTP 200 can contain an unhelpful answer. Define dashboards around an owner and an action, and set alerts for changes that affect customers rather than every harmless fluctuation.
Adopt a shared naming scheme, then pin its version
OpenTelemetry semantic conventions can make model, operation and token attributes easier to correlate across instrumentations and backends. GenAI conventions are evolving; record the convention and instrumentation version, review opt-in attributes before enabling them and check migration notes before upgrading. Do not make dashboards depend on undocumented provider fields without a fallback plan.
| Signal/span | Operational question | Fields needed | Sensitive data risk | Owner and retention |
|---|---|---|---|---|
How do you protect customer privacy in AI traces?
Make prompt and response capture opt-in and purpose-bound
Keep content recording off by default unless a documented debugging, quality or legal purpose requires it. Prefer metadata, redacted examples or a separately consented evaluation sample. If content is captured, explain the purpose, limit who can see it, set a short retention period and ensure vendor telemetry settings and contracts match the product's commitments.
Redact secrets and personal information before export
Remove authorization headers, API keys, session tokens, passwords, payment details and unnecessary direct identifiers before telemetry leaves the service. Apply redaction to tool arguments, exceptions, retrieval payloads and model outputs, not only the initial user prompt. Test redaction against real-shaped synthetic examples because a missed field can persist across multiple observability systems.
Keep tenant identity out of broad metric labels
Use low-cardinality dimensions such as product feature, model route, region and outcome for shared metrics. Avoid raw tenant IDs, user IDs or prompt text as metric labels; they can create enormous series counts and expose customer information. Put necessary tenant-level investigation data in access-controlled traces or a separate audit store with authorization and retention controls.
How should teams use AI traces during incidents?
Propagate correlation IDs without treating them as identity
Carry a request ID across application and provider boundaries where supported so teams can connect their own spans. Do not put secrets or personal data in IDs, and do not treat a correlation value supplied by a client as authentication or tenant authorization. Use trusted identity context separately from operational trace context.
Alert on patterns and preserve a useful investigation trail
Look for provider error increases, latency shifts, unusual token use, repeated retries, missing retrieval results and spikes in tool denials. Keep the trace fields needed to reconstruct sequence and policy outcomes, then sample or redact verbose events. Verify dashboards and alerts after instrumentation changes so missing telemetry is not mistaken for a healthy service.
Control access and test deletion across telemetry backends
Give trace access only to roles with an operational need, audit sensitive lookups and define a deletion schedule across collectors, dashboards, exports and backups. Ensure a customer-data deletion or retention request is reflected in the observability architecture. Test the whole path; deleting one UI row may leave copies in a queue or archive.
LLM observability FAQs
Should we log every prompt and completion?
No. Start with operational metadata and capture content only for a documented purpose with appropriate minimization, access controls, notice and retention. Full transcripts can contain customer secrets and personal data.
Which LLM metrics should a SaaS team watch first?
Start with request volume, success and error rates, latency, timeout and cancellation rates, token usage or cost, and the task-specific outcome that indicates user value. Add dimensions only when they support a clear operational action.
Is OpenTelemetry's GenAI convention stable?
The conventions are actively evolving. Check the current specification and instrumentation version, record what your pipeline emits and review changes before upgrades. Avoid assuming an attribute name or default remains unchanged.
Can trace IDs identify the customer who made a request?
A trace ID helps correlate operations; it is not a user identity or permission check. Keep authorization identity in trusted, access-controlled application context.
Related practical guides
Related issue guides
Sources and publication record
Draft prepared 27 September 2026; engineering, security and editorial review pending · Sources checked .
- Generative AI semantic conventionsOpenTelemetry
- Inside the LLM Call: GenAI Observability with OpenTelemetryOpenTelemetry
- OWASP Cheat Sheet: LoggingOWASP Foundation
- Production best practicesOpenAI API documentation
- Your data and model usage policies by endpointOpenAI Platform Documentation
- AWS SaaS Lens: Testing multi-tenant SaaS reliabilityAmazon Web Services
- Error codesOpenAI API documentation