Your AI workflow has logs. Can they explain one bad decision?

Illustrated infographic summarizing: Your AI workflow has logs. Can they explain one bad decision?

By Greg Nowak. Updated 21 September 2026.

An AI-assisted workflow approves the wrong request, retrieves an obsolete document or sends an inappropriate instruction to another system. Every component reports success. The dashboard is green. Yet nobody can explain why the business received the wrong result.

That is the useful test of AI observability: can your team reconstruct one consequential decision from its original trigger, through retrieval and model processing, to the final action? If the answer requires several engineers comparing timestamps across unrelated logs, you have data—but not an investigation trail.

A successful request can still be a bad decision

Traditional application logs are good at exposing exceptions, timeouts and failed API requests. AI workflows introduce another failure mode: every technical step can work as designed while the combined outcome is unsuitable.

The retrieval layer may return a valid but irrelevant document. The model may use a different version than expected. A tool call can contain structurally valid arguments that should not have been used for this case. The destination system may correctly accept an action that a person later has to reverse.

You therefore need both operational evidence and decision evidence. Operational evidence covers latency, errors, model identity and token usage. Decision evidence identifies the sources, workflow branch, selected tool, governed outcome category and subsequent human correction. The objective is not to record the model’s private reasoning. It is to preserve enough observable facts to explain what the system received, selected and did.

Workflow stage Useful evidence Question it should answer
Trigger Trace ID, workflow version, event type and time What started this run?
Retrieval Data-source ID, approved document identifiers, result count and latency Which information influenced it?
Model call Operation, provider, requested and returned model, duration and token usage What actually ran?
Tool execution Tool name, call ID, status and safe argument classification What action was attempted?
Downstream API Service, operation, status and correlated request ID Where did the action finish?
Outcome Outcome category, review state, correction and replay result Was the business result acceptable?
A minimum evidence map for investigating one AI-assisted decision without storing every prompt, document and payload.

Use one trace to connect the journey

A shared trace identifier turns separate operations into a navigable sequence. The current W3C Trace Context specification defines the interoperable traceparent and optional tracestate HTTP headers. Each component you control should extract the incoming context, create its own child span and inject the updated context into the next request.

Queues, background jobs and webhook handlers need the same attention as synchronous APIs. Where a third-party service cannot accept trace context, retain its safe request or job identifier on the surrounding span. Do not place customer details or other personal information in tracestate; the specification explicitly treats it as propagation metadata, not a place for business content.

OpenTelemetry’s GenAI conventions now describe spans for model operations, agent invocation, workflows and tool execution, with shared attributes for providers, models, data sources and usage. They remain under development, so pin the convention and instrumentation versions you deploy. Otherwise, a library update can quietly rename fields or break dashboards during the incident you are trying to investigate.

Make metadata the default, not raw content

Copying every prompt, response, retrieved document, tool argument and result into telemetry feels convenient. It can also create an uncontrolled second repository containing customer records, personal data, secrets and commercially sensitive instructions.

Start with an allowlist of metadata: stable identifiers, categories, counts, versions, durations, statuses and error types. OpenTelemetry marks fields such as tool arguments and results as potentially sensitive. Its Collector redaction processor can remove attributes outside an allowlist and mask matching values, but redaction should be treated as one control—not proof that arbitrary content is safe to collect.

When raw content is genuinely necessary, make the exception explicit. Define which workflow and failure class justify capture, who can access it, where filtering happens and when the content is deleted. Test those controls with representative secrets and personal data before enabling them in production. OWASP also recommends excluding or transforming access tokens, passwords, connection strings, encryption keys and sensitive personal information, while sanitising untrusted event data to prevent log injection.

Build the investigation path in five passes

  1. Choose one real decision. Use a reviewed poor outcome or a plausible high-impact case, not an abstract architecture diagram.
  2. Draw the execution path. Include retrieval, model calls, workflow branches, tools, queues and external systems as they actually run.
  3. Define the minimum evidence. For every step, write down the question an investigator must answer and retain only the fields needed to answer it.
  4. Propagate and verify context. Run a test case, open its trace and confirm that parent-child relationships survive every boundary you control.
  5. Close the loop. Turn a reviewed failure into a replay case, then rerun it after changing the prompt, model, retrieval rule, guardrail or integration.

Decide retention and sampling only after this path works. Routine successful runs may need less retention than errors, human escalations or rare workflow branches. Whatever policy you choose, preserve the cases that matter long enough for operations and business owners to review them together.

The deliverable is an answer, not another dashboard

A useful observability setup should let someone open one trace and answer: what triggered this decision, what information shaped it, which model and tools ran, what changed downstream, and did the correction work?

If your current logs cannot tell that story, Greg can help map a real workflow, identify the evidence worth keeping and coordinate the instrumentation, privacy and operational decisions across your team. The first step can be a focused review of one workflow and one outcome—before an incident forces everyone to reconstruct it under pressure.

Related on GrN.dk

Need help with this kind of work?

Map your AI workflow with Greg Get in touch with Greg.

Sources

Latest articles

An internal AI assistant can cite an obsolete handbook with confidence. Here is how to manage document ownership, updates, deletions, access and answer review.

Cloudflare Free provides useful website protection, but its rate limiting and bot controls have limits. Here is how to assess them for a WordPress site.

An AI assistant can answer questions and guide customers to a booking. Here are practical boundaries for prices, delivery times, personal data, and contact with a staff member.

Google and Bing now offer first-party AI search visibility reports. Here’s how to build a useful baseline without inventing a misleading GEO score.

AI crawlers can copy a familiar name. Here’s how to verify signed agents at the edge while keeping legitimate automated traffic moving.

A critical Webform release is a reminder to audit every Drupal codebase, configuration and deployment—not just the main production website.

A secure AI workflow can turn Meet and Teams transcripts into approved decisions and tasks in Jira or Asana—without giving up control.

NGINX 1.31.5 can route on JSON body values. Here’s how to weigh the performance, security, and operational trade-offs before using it.

OpenAI can keep agent sessions running, but reliable workflows still depend on clear failure states, safe retries, validation, limits and human fallback.

AI can identify termination deadlines and price adjustments in supplier contracts, route uncertain findings for approval and create the right reminders.