OpenAI Presence Is Here: Is Your Workflow Ready for an Agent?

Illustrated infographic summarizing: OpenAI Presence Arrived—But Is Your Workflow Ready for an Agent?

By Greg Nowak. Updated 25 September 2026.

OpenAI Presence is not simply a chatbot with a new label. It is a managed enterprise product for voice and chat agents that can use company systems, follow policies, take approved actions and transfer work to people. OpenAI currently offers it through limited general availability to eligible enterprise customers, with deployments led by its Forward Deployed Engineers and selected systems integrators. It is not a self-service product.

That delivery model may reduce some technical work, but it does not make a confused process safe to automate. Before discussing models, integrations or procurement, ask a more useful question: is the workflow clear enough to give an agent authority?

Presence makes operational gaps visible

A conventional chatbot usually answers a question. An agent may retrieve an account, interpret policy, update a record, send a message or begin another process. Once software can act, small inconsistencies become operational risks.

The customer status may differ across two systems. An experienced employee may rely on an exception that was never documented. A handoff may send the conversation to a queue without the information already collected. Presence includes policies, guardrails, approved actions, simulations, evaluations and escalation rules, but the business still has to define the job those controls support.

This is why an agent project is partly a process-design project. The technical implementation can only be as coherent as the operating decisions behind it.

Start with one job that has a recognisable finish

“Improve customer service with AI” is too broad. “Answer delivery-status enquiries using approved order data, while transferring damaged or disputed orders to the fulfilment team” is testable.

A sensible first workflow has a clear trigger, a defined outcome and one accountable owner. It occurs often enough to justify the effort, while remaining narrow enough to inspect its normal route and important exceptions.

Do not use an agent merely because a task is repetitive. Conventional automation will usually be cheaper and easier to audit when the inputs, rules and outputs are predictable. Agents become more useful when work involves unstructured information, contextual decisions or rules that have grown difficult to maintain.

Readiness test Ready to prototype Fix before building
Scope One trigger, outcome and owner Several teams or unrelated outcomes are bundled together
Knowledge Current policies and common exceptions are documented Success depends on unwritten staff knowledge
Data Authoritative sources and freshness requirements are known Employees routinely reconcile conflicting records
Authority Read, draft, write and approve permissions are separated Broad access is requested for convenience
Handoff The receiving person gets the history and required next action The customer or employee must explain everything again
Evidence Representative cases have expected outcomes Approval rests on a polished demonstration
A workflow is ready when the team can produce evidence for each row—not merely answer yes during a meeting.

Map the real workflow, including the awkward cases

Follow the work from the event that starts it to the final recorded outcome. Identify what information arrives, which systems an employee checks, which decisions they make, what action follows and how the organisation knows the work is complete.

Then investigate the cases that procedures often omit: missing information, conflicting records, unavailable tools, suspected fraud, a request outside policy or a customer who rejects the proposed answer. For every meaningful step, record:

  • the data and context required;
  • the policy or procedure that governs the decision;
  • the tools available to the agent;
  • the actions it may take independently;
  • the point where approval or human ownership becomes mandatory; and
  • the context that must travel with a handoff.

This work frequently exposes improvements worth making even if the agent project stops.

Design permission around consequences

Give the agent the minimum access required for its defined job. Separate reading from acting: an agent might inspect an account and draft a resolution while a person remains responsible for changing the record or contacting the customer.

Classify each action by consequence and reversibility. Read-only retrieval may run automatically. A reversible update might require logging, validation and volume limits. Payments, cancellations, legal commitments and sensitive communications may require explicit approval or remain outside the agent’s authority.

Authentication and access control still matter; guardrails are an additional layer, not a replacement. The business should be able to determine which identity acted, which information informed the action, what changed and whether the action stayed within policy.

Build the test pack before the prototype impresses anyone

A useful prototype investigates failure rather than presenting only a smooth demonstration. Can the agent find the authoritative record? Does it choose the correct tool? Will it stop when systems disagree? Does the human handoff include the relevant history? Can a proposed action be prevented or reversed?

Create an evaluation set from representative, appropriately handled business cases. Include routine requests, incomplete information, ambiguous instructions, policy exceptions, attempted manipulation and tool failures. Remove or protect personal data as required.

Score the complete run: interpretation, source use, tool choice, policy compliance, action and escalation. OpenAI’s current evaluation guidance distinguishes individual traces, which help diagnose behaviour, from repeatable datasets and evaluation runs used to compare changes and detect regressions. That is a stronger acceptance process than judging whether the final response sounds convincing.

Assign production ownership before launch

Pre-launch testing cannot reproduce every real condition. NIST’s 2026 work on deployed AI monitoring stresses the need to check whether systems continue to operate reliably, detect unforeseen outputs and identify unexpected consequences in their actual environment.

Decide who reviews failed runs and escalations, which signals trigger an investigation and who can approve a change. Track measures tied to the job: successful completion, human handoff, incorrect tool use, approval requests, tool failures and new request types. Proposed changes should pass the representative test set before a controlled release.

Choose the delivery route after the readiness review

Presence may suit an eligible enterprise with a substantial voice or chat workflow and a preference for a managed deployment. A custom implementation using OpenAI’s developer platform may fit organisations that need more control over channels, integrations, hosting or application architecture. OpenAI’s current public documentation directs new agent applications towards the Agents API, while leaving the surrounding application, tool connections and execution choices to the developer.

Sometimes the correct answer is neither route. If nobody owns the policy, the records conflict or the handoff is broken, fix that foundation first.

Greg can help map the process, define permissions and handoffs, assemble evaluation cases and turn the findings into a practical buy, build or wait decision. If your agent project is moving faster than its operating model, start with a workflow-readiness conversation.

Related on GrN.dk

Need help with this kind of work?

Assess your agent workflow with Greg Get in touch with Greg.

Sources

Latest articles

Improve site search for customers who use everyday language. See where synonyms, semantic retrieval and behavioural signals help, while keeping product codes reliable.

Follow customer data through n8n, OpenAI and your CRM. Check who can access retained copies, what redaction hides and whether deletion works before scaling.

When checkout fails, your operations provider needs concrete evidence to work with. See how AI, dmesg and journalctl can gather the evidence into a useful incident ticket.

OpenAI’s hosted Evals platform is closing. Preserve your tests, validate replacement scoring and keep releases covered before the October and November 2026 deadlines.

Decide which AI-assisted pages to keep, improve, combine or remove. Check claims, page overlap and metadata, then put clear review controls into your CMS.

Use October to trial daily AI reorder recommendations before Black Friday. Get your Shopify data, lead times and budget in order before turning recommendations into purchases.

When an OpenAI request stalls, customers need an accurate status. Set sensible retry limits, preserve submissions, and make unresolved work visible.

I learned server operations by breaking my own servers. I want someone who stands next to me while I do it, then does it themselves the week after.

I am good at building and bad at calling. Here is who I want next to me, what is easiest to sell, and how we split it.

An AI assistant can prepare a refund, but a person should approve the exact payment and amount. Here is how to make that approval hold up through execution and retries.