Your AI workflow has logs. Can they explain one bad decision?

Illustrated infographic summarizing: Your AI workflow has logs. Can they explain one bad decision?

By Greg Nowak. Updated 21 September 2026.

An AI-assisted workflow approves the wrong request, retrieves an obsolete document or sends an inappropriate instruction to another system. Every component reports success. The dashboard is green. Yet nobody can explain why the business received the wrong result.

That is the useful test of AI observability: can your team reconstruct one consequential decision from its original trigger, through retrieval and model processing, to the final action? If the answer requires several engineers comparing timestamps across unrelated logs, you have data—but not an investigation trail.

A successful request can still be a bad decision

Traditional application logs are good at exposing exceptions, timeouts and failed API requests. AI workflows introduce another failure mode: every technical step can work as designed while the combined outcome is unsuitable.

The retrieval layer may return a valid but irrelevant document. The model may use a different version than expected. A tool call can contain structurally valid arguments that should not have been used for this case. The destination system may correctly accept an action that a person later has to reverse.

You therefore need both operational evidence and decision evidence. Operational evidence covers latency, errors, model identity and token usage. Decision evidence identifies the sources, workflow branch, selected tool, governed outcome category and subsequent human correction. The objective is not to record the model’s private reasoning. It is to preserve enough observable facts to explain what the system received, selected and did.

Workflow stage Useful evidence Question it should answer
Trigger Trace ID, workflow version, event type and time What started this run?
Retrieval Data-source ID, approved document identifiers, result count and latency Which information influenced it?
Model call Operation, provider, requested and returned model, duration and token usage What actually ran?
Tool execution Tool name, call ID, status and safe argument classification What action was attempted?
Downstream API Service, operation, status and correlated request ID Where did the action finish?
Outcome Outcome category, review state, correction and replay result Was the business result acceptable?
A minimum evidence map for investigating one AI-assisted decision without storing every prompt, document and payload.

Use one trace to connect the journey

A shared trace identifier turns separate operations into a navigable sequence. The current W3C Trace Context specification defines the interoperable traceparent and optional tracestate HTTP headers. Each component you control should extract the incoming context, create its own child span and inject the updated context into the next request.

Queues, background jobs and webhook handlers need the same attention as synchronous APIs. Where a third-party service cannot accept trace context, retain its safe request or job identifier on the surrounding span. Do not place customer details or other personal information in tracestate; the specification explicitly treats it as propagation metadata, not a place for business content.

OpenTelemetry’s GenAI conventions now describe spans for model operations, agent invocation, workflows and tool execution, with shared attributes for providers, models, data sources and usage. They remain under development, so pin the convention and instrumentation versions you deploy. Otherwise, a library update can quietly rename fields or break dashboards during the incident you are trying to investigate.

Make metadata the default, not raw content

Copying every prompt, response, retrieved document, tool argument and result into telemetry feels convenient. It can also create an uncontrolled second repository containing customer records, personal data, secrets and commercially sensitive instructions.

Start with an allowlist of metadata: stable identifiers, categories, counts, versions, durations, statuses and error types. OpenTelemetry marks fields such as tool arguments and results as potentially sensitive. Its Collector redaction processor can remove attributes outside an allowlist and mask matching values, but redaction should be treated as one control—not proof that arbitrary content is safe to collect.

When raw content is genuinely necessary, make the exception explicit. Define which workflow and failure class justify capture, who can access it, where filtering happens and when the content is deleted. Test those controls with representative secrets and personal data before enabling them in production. OWASP also recommends excluding or transforming access tokens, passwords, connection strings, encryption keys and sensitive personal information, while sanitising untrusted event data to prevent log injection.

Build the investigation path in five passes

  1. Choose one real decision. Use a reviewed poor outcome or a plausible high-impact case, not an abstract architecture diagram.
  2. Draw the execution path. Include retrieval, model calls, workflow branches, tools, queues and external systems as they actually run.
  3. Define the minimum evidence. For every step, write down the question an investigator must answer and retain only the fields needed to answer it.
  4. Propagate and verify context. Run a test case, open its trace and confirm that parent-child relationships survive every boundary you control.
  5. Close the loop. Turn a reviewed failure into a replay case, then rerun it after changing the prompt, model, retrieval rule, guardrail or integration.

Decide retention and sampling only after this path works. Routine successful runs may need less retention than errors, human escalations or rare workflow branches. Whatever policy you choose, preserve the cases that matter long enough for operations and business owners to review them together.

The deliverable is an answer, not another dashboard

A useful observability setup should let someone open one trace and answer: what triggered this decision, what information shaped it, which model and tools ran, what changed downstream, and did the correction work?

If your current logs cannot tell that story, Greg can help map a real workflow, identify the evidence worth keeping and coordinate the instrumentation, privacy and operational decisions across your team. The first step can be a focused review of one workflow and one outcome—before an incident forces everyone to reconstruct it under pressure.

Related on GrN.dk

Need help with this kind of work?

Map your AI workflow with Greg Get in touch with Greg.

Sources

Seneste artikler

En AI-assistent kan svare på spørgsmål og føre kunder til booking. Her er de konkrete grænser for pris, levering, personoplysninger og kontakt med en medarbejder.

Et sikkert AI-workflow kan omsætte Meet- og Teams-transskripter til godkendte beslutninger og opgaver i Jira eller Asana – uden at slippe kontrollen.

AI kan finde opsigelsesfrister og prisreguleringer i leverandørkontrakter, sende usikre fund til godkendelse og oprette de rette påmindelser.

Sådan automatiserer danske virksomheder Gmail og Microsoft 365 med hurtig sortering, begrænsede rettigheder og menneskelig godkendelse.

Samme kunde på flere kort i HubSpot? Se, hvordan CVR-match, AI-forslag og menneskelig godkendelse kan bruges til at rydde op med styr på felter, relationer og kundehistorik.

Få en ugentlig marketingrapport fra GA4 og Google Ads med kontrollerede beregninger, tydelige dataforbehold og et kort AI-udkast, der hjælper jer på mandagsmødet.

Brug AI til webshoppens alt-tekster med en overskuelig pilot: kortlæg billederne, få danske forslag, og kontrollér resultatet i WordPress og WooCommerce.

AI-baseret ticketanalyse kan afsløre gentagne klager, produktfejl og huller i dokumentationen – uden at virksomheden behøver endnu en chatbot.

OpenSSH 10 fjerner DSA og advarer om nøgleudveksling, der ikke er post-kvantesikker. Her får du en metode til at afgrænse SFTP-oprydningen uden at svække alle SSH-forbindelser.

Botforespørgsler overstiger nu menneskelig webtrafik. Lær at auditere AI-crawlere, fastsætte regler på stiniveau, håndhæve robots.txt og måle det forretningsmæssige afkast.