By Greg Nowak. Updated 21 September 2026.
An AI-assisted workflow approves the wrong request, retrieves an obsolete document or sends an inappropriate instruction to another system. Every component reports success. The dashboard is green. Yet nobody can explain why the business received the wrong result.
That is the useful test of AI observability: can your team reconstruct one consequential decision from its original trigger, through retrieval and model processing, to the final action? If the answer requires several engineers comparing timestamps across unrelated logs, you have data—but not an investigation trail.
A successful request can still be a bad decision
Traditional application logs are good at exposing exceptions, timeouts and failed API requests. AI workflows introduce another failure mode: every technical step can work as designed while the combined outcome is unsuitable.
The retrieval layer may return a valid but irrelevant document. The model may use a different version than expected. A tool call can contain structurally valid arguments that should not have been used for this case. The destination system may correctly accept an action that a person later has to reverse.
You therefore need both operational evidence and decision evidence. Operational evidence covers latency, errors, model identity and token usage. Decision evidence identifies the sources, workflow branch, selected tool, governed outcome category and subsequent human correction. The objective is not to record the model’s private reasoning. It is to preserve enough observable facts to explain what the system received, selected and did.
| Workflow stage | Useful evidence | Question it should answer |
|---|---|---|
| Trigger | Trace ID, workflow version, event type and time | What started this run? |
| Retrieval | Data-source ID, approved document identifiers, result count and latency | Which information influenced it? |
| Model call | Operation, provider, requested and returned model, duration and token usage | What actually ran? |
| Tool execution | Tool name, call ID, status and safe argument classification | What action was attempted? |
| Downstream API | Service, operation, status and correlated request ID | Where did the action finish? |
| Outcome | Outcome category, review state, correction and replay result | Was the business result acceptable? |
Use one trace to connect the journey
A shared trace identifier turns separate operations into a navigable sequence. The current W3C Trace Context specification defines the interoperable traceparent and optional tracestate HTTP headers. Each component you control should extract the incoming context, create its own child span and inject the updated context into the next request.
Queues, background jobs and webhook handlers need the same attention as synchronous APIs. Where a third-party service cannot accept trace context, retain its safe request or job identifier on the surrounding span. Do not place customer details or other personal information in tracestate; the specification explicitly treats it as propagation metadata, not a place for business content.
OpenTelemetry’s GenAI conventions now describe spans for model operations, agent invocation, workflows and tool execution, with shared attributes for providers, models, data sources and usage. They remain under development, so pin the convention and instrumentation versions you deploy. Otherwise, a library update can quietly rename fields or break dashboards during the incident you are trying to investigate.
Make metadata the default, not raw content
Copying every prompt, response, retrieved document, tool argument and result into telemetry feels convenient. It can also create an uncontrolled second repository containing customer records, personal data, secrets and commercially sensitive instructions.
Start with an allowlist of metadata: stable identifiers, categories, counts, versions, durations, statuses and error types. OpenTelemetry marks fields such as tool arguments and results as potentially sensitive. Its Collector redaction processor can remove attributes outside an allowlist and mask matching values, but redaction should be treated as one control—not proof that arbitrary content is safe to collect.
When raw content is genuinely necessary, make the exception explicit. Define which workflow and failure class justify capture, who can access it, where filtering happens and when the content is deleted. Test those controls with representative secrets and personal data before enabling them in production. OWASP also recommends excluding or transforming access tokens, passwords, connection strings, encryption keys and sensitive personal information, while sanitising untrusted event data to prevent log injection.
Build the investigation path in five passes
- Choose one real decision. Use a reviewed poor outcome or a plausible high-impact case, not an abstract architecture diagram.
- Draw the execution path. Include retrieval, model calls, workflow branches, tools, queues and external systems as they actually run.
- Define the minimum evidence. For every step, write down the question an investigator must answer and retain only the fields needed to answer it.
- Propagate and verify context. Run a test case, open its trace and confirm that parent-child relationships survive every boundary you control.
- Close the loop. Turn a reviewed failure into a replay case, then rerun it after changing the prompt, model, retrieval rule, guardrail or integration.
Decide retention and sampling only after this path works. Routine successful runs may need less retention than errors, human escalations or rare workflow branches. Whatever policy you choose, preserve the cases that matter long enough for operations and business owners to review them together.
The deliverable is an answer, not another dashboard
A useful observability setup should let someone open one trace and answer: what triggered this decision, what information shaped it, which model and tools ran, what changed downstream, and did the correction work?
If your current logs cannot tell that story, Greg can help map a real workflow, identify the evidence worth keeping and coordinate the instrumentation, privacy and operational decisions across your team. The first step can be a focused review of one workflow and one outcome—before an incident forces everyone to reconstruct it under pressure.
Related on GrN.dk
- AI Agents Need a Spending Brake, Not Just a Billing Dashboard
- AI automations need a spend dashboard before the first runaway bill
- OpenAI Computer Use: Browser Agents Need Credentials, Not Demos
Need help with this kind of work?
Map your AI workflow with Greg Get in touch with Greg.