OpenAI Evals Is Closing: Keep Your Workflow Tests Running

Illustrated infographic summarizing: OpenAI Evals Is Closing: Keep Your Workflow Tests Running

By Greg Nowak. Last updated 2026-10-08.

If your agency uses OpenAI’s hosted Evals platform to check AI workflows, its closure creates a practical problem for the next prompt or model change. You still need to know what was tested, what counted as a pass, and which failures should stop a release.

Those decisions need to survive the migration alongside the test inputs. Whoever maintains the workflow afterwards should be able to run the checks and explain why it is ready to deploy.

Plan around both deadlines

OpenAI’s deprecation schedule says existing evaluations become read-only on October 31, 2026. The Evals dashboard and API are scheduled to shut down on November 30, 2026. The graders documented for evaluation workflows are also part of the transition.

Aim to have a working replacement by October 31. That leaves the remaining transition window to investigate differences and finish the handover while you can still consult the hosted dashboard. Waiting until November 30 takes that reference away during the work.

Identify the releases that currently depend on hosted evaluations. Each needs an agreed testing route while you validate the replacement; otherwise, a platform migration can leave a delivery team without the checks it normally relies on.

Keep the context that makes a test useful

Start with an inventory by business workflow. Record who owns each evaluation, which behaviour it protects, and where its results affect a release decision. Give tests used in current delivery priority. Older experiments can have a separate archive status.

For each evaluation, preserve enough information for the next maintainer to understand both the test and the decision it supports.

Evaluation handover checklist: what to keep and how to check the replacement
Asset Keep Ready when
Test cases Inputs, expected outcomes, reference answers and relevant context. Each case has a clear link to the behaviour it tests.
Prompts and configuration Instructions, variables, model identifiers and relevant settings. The replacement runs the intended version.
Grading criteria Rules, rubrics, judge instructions and scoring definitions. Reviewed examples show that grading is acceptable.
Historical results Available outputs, scores, run dates and configuration references. A maintainer can inspect the baseline used for earlier decisions.
Acceptance thresholds Pass conditions, critical failures and review requirements. The release owner can explain what blocks deployment.

Keep historical results in an archive with the context needed to interpret them. If something is missing, label the gap. That lets the next maintainer judge what the archive can reliably tell them.

Thresholds deserve particular attention. A score alone may not explain whether a rule applies to the whole suite, one workflow or an essential behaviour. Write that down, along with who can approve an exception. These are release decisions that need to carry through to the replacement.

Rebuild one evaluation first

The OpenAI Cookbook migration guide describes manually recreating evaluations in Promptfoo. Prompts, providers, test cases and grading behaviour move into a portable configuration, with assertions and metrics replacing hosted testing criteria. The evaluations can then run locally or in continuous integration.

The guide does not rely on an OpenAI Evals export feature. New Promptfoo runs are separate from completed hosted runs, and workflows involving tools or agents may need additional configuration. Allow for that work when setting the scope.

Choose a pilot evaluation that protects a current workflow and has clear success criteria. Use it to settle the file structure, naming, review process and report format. Record any judgement calls made during the rebuild. The pilot then gives the rest of the migration a practical pattern, with unresolved questions visible before you repeat it across the suite.

Check the decisions behind the new scores

The Cookbook warns that similarity scores can differ between systems. Recreated graders, especially model judges, also need validation. An old numerical threshold producing a pass in the new runner does not, by itself, show that the replacement is grading correctly.

Compare individual pilot cases. Keep the inputs and intended behaviour fixed, then investigate where the graders disagree. Ask the workflow owner to review disputed outputs against the written criteria. For each disagreement, record whether it calls for a configuration correction, a clearer rubric or an intentional change to the acceptance rule.

Anthropic’s guidance on agent evaluations recommends repeated trials because outputs vary, and calibration of model judges against human expertise. It also recommends reading transcripts to separate genuine workflow failures from problems with the evaluation itself.

Include a comparison report in the handover. It should explain which release decisions remain consistent, where disagreement remains, and who accepted any revised criteria. That gives the release owner a basis for approving the replacement beyond comparing two overall scores.

For automations, check what actually happened

Anthropic distinguishes the agent’s transcript from the final state it leaves in the environment. That distinction matters when an automation is supposed to take an action: the evaluation needs to check whether the action happened.

For a workflow that updates a record, define a check against the resulting record in the test environment. Response quality can have its own check where it matters, but successful completion needs a separate criterion. Review these outcome checks during migration so the rebuilt suite still tests the business requirement.

Make the checks part of a release

Once you have validated the replacement, automate its execution before releases. Promptfoo’s CI/CD documentation describes evaluation runs, JSON and HTML reports, JUnit XML output, and quality gates that fail a build when thresholds are unmet. These give you ways to connect test results to deployment decisions.

Agree which changes trigger testing, where reviewers find the report, and who investigates failures. Require explicit acceptance of the run before deployment. If a particular behaviour is essential, give it a blocking check of its own; an overall pass rate can leave that requirement unclear.

Before handover, verify the gate with an intentionally failing test case. Keep evidence of both a successful run and a blocked release. Someone other than the implementer should be able to find the report and explain why the failure stopped deployment.

A manageable project with a clear finish

A scoped project with Greg through GrN.dk could begin with one agency workflow. The work would cover its evaluation inventory, preservation of available assets, a rebuild in a portable runner, a comparison of grading decisions, and release checks connected to the validated suite.

Agree the deliverables upfront: an asset inventory, runnable configuration, comparison report, documented acceptance thresholds and a demonstrated release check. Include a rerun and handover session so the agency can operate the suite after the project ends.

The inventory should determine the scope, including any custom workflow setup. Name a business owner for the success criteria and a technical owner for ongoing maintenance. That gives the project a clear finish and the team a dependable starting point for future prompt or model changes.

Contact Greg through GrN.dk to discuss a scoped evaluation migration before the October 31 read-only deadline.

Related on GrN.dk

Need help with this kind of work?

Discuss your evaluation migration with Greg Get in touch with Greg.

Sources

Latest articles

OpenAI’s hosted Evals platform is closing. Preserve your tests, validate replacement scoring and keep releases covered before the October and November 2026 deadlines.

Decide which AI-assisted pages to keep, improve, combine or remove. Check claims, page overlap and metadata, then put clear review controls into your CMS.

Use October to trial daily AI reorder recommendations before Black Friday. Get your Shopify data, lead times and budget in order before turning recommendations into purchases.

When an OpenAI request stalls, customers need an accurate status. Set sensible retry limits, preserve submissions, and make unresolved work visible.

I learned server operations by breaking my own servers. I want someone who stands next to me while I do it, then does it themselves the week after.

I am good at building and bad at calling. Here is who I want next to me, what is easiest to sell, and how we split it.

An AI assistant can prepare a refund, but a person should approve the exact payment and amount. Here is how to make that approval hold up through execution and retries.

AI can pull together onboarding tasks before a new hire’s first day. See how the manager approves specific access and how outstanding tasks are followed through.

An internal AI assistant can cite an obsolete handbook with confidence. Here is how to manage document ownership, updates, deletions, access and answer review.

Cloudflare Free provides useful website protection, but its rate limiting and bot controls have limits. Here is how to assess them for a WordPress site.