Your AI workflow has logs. Can they explain one bad decision?

Illustrated infographic summarizing: Your AI workflow has logs. Can they explain one bad decision?

By Greg Nowak. Last updated 2026-08-22.

An AI workflow approves the wrong request. Or it retrieves an outdated document, then sends an unsuitable action to another system. Every service has logs, and every request may even show as successful. Still, nobody can explain how the workflow arrived at the result.

That is the useful test of AI observability. Can the team take one consequential decision and reconstruct it from the original trigger through retrieval, model processing and tool calls to the final change in a business system?

An API log usually answers a narrower question: did this request complete? Investigating an AI-assisted decision requires connected evidence about what happened along the way. Simply collecting more log entries will not provide that connection.

AI telemetry is becoming more structured

In August 2026, Google introduced experimental Gemini Enterprise agent telemetry aligned with OpenTelemetry's generative AI semantic conventions. Its trace spans and log entries can carry standardized attributes for agents, conversations, models, tools and token use. Message content is included only when the relevant observability settings permit it; otherwise, Google says it is redacted or omitted. The conventions are still being developed, so implementations need to allow for change.

This is useful beyond Gemini Enterprise. The OpenTelemetry GenAI semantic conventions establish shared terminology for GenAI clients, Model Context Protocol activity and provider-specific integrations. The related span conventions cover model inference, retrieval, workflow invocation and tool execution.

A shared vocabulary will not make an incomplete trace useful. It does remove a mundane but costly problem: translating differently named fields every time an investigation crosses a system boundary.

A successful request can still end badly

Operational logs are good at showing familiar technical failures: HTTP errors, exceptions, timeouts and unavailable dependencies. Teams still need those signals. The awkward part is that an AI-assisted workflow can complete every technical step successfully and produce the wrong business result.

The retrieval service might return a valid but irrelevant record. A model call could use a different model version from the one expected. A tool might accept arguments that are structurally correct but inappropriate for the case. Even the downstream API can return a success status for an action that later has to be corrected.

To diagnose such a run, the trace needs two kinds of evidence. Technical evidence covers timing, errors, model identity, usage and the relationship between operations. Decision evidence records what shaped the outcome, such as the retrieval source, selected tool, governed outcome category and any later correction. That second group deserves particular care: use deliberately chosen, low-risk metadata instead of copying prompts and business records wholesale.

Stage Evidence worth keeping Investigation question
Trigger Workflow name, event type, trace ID, timestamp What started this run?
Retrieval Data-source identifier, document identifiers, result count, latency Which source influenced the next step?
Model call Operation, provider, requested model, response model, token use, duration Which model ran, and at what operational cost?
Tool execution Tool name, status, duration, safe argument classification What action did the workflow attempt?
Downstream API Service, operation, status, error type, correlated trace context Where did the action finish or fail?
Outcome Approved result category, review state, correction or replay result Was the business outcome acceptable?
An evidence map for following one AI-assisted decision without making raw content the default debugging record.

The trace ID connects the story

An investigation becomes much easier when the same trace context survives the journey across applications, queues, model providers, tools and downstream APIs. The W3C Trace Context specification defines the interoperable traceparent and tracestate headers used for that job. traceparent carries the trace identifier, current parent operation and trace flags. tracestate can include optional vendor-specific information.

That propagation has to be designed. Each component should accept the context it receives, create a child operation and forward the updated context. If a third-party system cannot participate, the integration layer can retain a safe correlation identifier that ties its request and response back to the surrounding trace.

The result is a sequence the team can navigate, rather than a story assembled by comparing timestamps from unrelated systems. Because the correlation mechanism is standardized, it also reduces reliance on any single observability vendor.

Content capture needs an explicit policy

Recording every prompt, response, retrieved document and tool payload can look like the fastest route to easier debugging. It can also turn the observability platform into a new store of confidential and personal information.

OpenTelemetry warns that input and output messages, prompt variables, system instructions, tool definitions, tool arguments and tool results may contain sensitive data. Some content-bearing retrieval and memory attributes are opt-in rather than default fields. Google's Gemini Enterprise release note likewise describes controls that redact or omit prompt and response content when content logging is not permitted.

Metadata first is the sensible starting point. Identifiers, categories, timings, counts, versions, statuses and error types often provide enough information to locate a problem. When actual content is necessary for a defined diagnostic purpose, the policy should answer four practical questions: who can see it, how it is filtered, how long it is retained and who reviews that decision.

The OWASP Logging Cheat Sheet advises that access tokens, passwords, connection strings, encryption keys, payment data and sensitive personal information should generally be removed, masked, sanitized, hashed or encrypted instead of logged directly. It also recommends sanitizing event data to prevent log injection. In an AI workflow, redaction therefore needs to deal with both secrets and untrusted text that could corrupt or mislead the logging system.

Start with an investigation, not a dashboard

A practical observability engagement can start with one business workflow and one blunt question: can we explain an unacceptable outcome from its trigger to the final API call?

Map the execution path as it actually runs, including retrieval, model calls, tool execution and external services. At each stage, decide on the minimum evidence needed to investigate a problem. Then propagate trace context across every boundary the team controls. Operational fields such as latency, token use, model identity and errors can sit alongside carefully governed classifications for sources and outcomes.

Sampling and retention come next. A routine successful run may not warrant the same retention as an error, a reviewed outcome or a rare workflow branch. The rules need to preserve cases that matter for investigation without allowing the observability system to become permanent content storage by accident.

The trace should also lead to action. Alerts can surface unusual error rates, latency or usage. Once a poor outcome has been reviewed, it can become a replay test for the relevant workflow path after a prompt, model, retrieval rule or integration changes. The point is not to present the AI as infallible. It is to make failures explainable, corrections testable and ownership unambiguous.

What the logs should let you do

When the next bad decision appears, the team should be able to open one trace and follow the parent-child sequence. Safe metadata should show what the workflow retrieved and invoked, where its behaviour departed from expectations and whether the eventual correction works.

Greg can help map that path for a real workflow, define the evidence needed at each step and turn disconnected logs into an investigation the team can actually follow. For workflows that affect customers, records or business operations, doing this before an incident is considerably easier than reconstructing it under pressure.

Related on GrN.dk

Need help with this kind of work?

Map your AI workflow Get in touch with Greg.

Sources

Latest articles

Build a weekly marketing report from GA4 and Google Ads with verified calculations, clear data caveats and a short AI draft to support your Monday meeting.

Before buying a GPU, test one real team workflow on existing hardware. A Linux pilot can show whether quality, memory, response times, and running costs add up.

Planning a Drupal relaunch? Set clear rules for content, translations, media and old URLs, with a practical checklist for approving the migration and launch.

Use AI for your online store’s alt text with a manageable pilot: map the images, generate suggestions in Danish, and check the results in WordPress and WooCommerce.

Supplier files need more than extraction. Here’s how to check coverage, match SKUs, resolve unclear units and prices, and test product data before a catalogue import.

Shorter TLS certificates leave less room for renewal problems. Check domain validation, scheduling, deployment and the certificate your customers actually receive.

AI image credentials can disappear during routine website processing. Learn how to test your CMS, optimizer, CDN, and publishing workflow end to end.

AI-based ticket analysis can uncover recurring complaints, product defects and gaps in documentation—without the company needing yet another chatbot.

OpenAI’s X.509 workload identity can replace API keys for the right workloads. This practical framework helps teams decide where to start safely.

WordPress 7.1 helps AI agents discover and invoke site abilities. Here is how to keep exposure, authentication and permission firmly separate.

Review Greg on Google

Greg Nowak Google Reviews

 

Written recommendations from Trafik og Veje, Aarhus Municipality (2011) and AgroTech (2010) — read them on LinkedIn.