By Greg Nowak. Last updated 2026-07-24.
An AI visibility dashboard can look reassuringly precise while measuring something quite unstable. A rising line might mean your source coverage has improved. It might also mean the model changed, the interface changed, one test produced a favourable answer, or somebody edited the prompt set between reports.
If the methodology cannot tell those situations apart, the dashboard is not ready to guide budgets or priorities.
The problem becomes more pressing as generative search enters additional markets. Google has launched AI Overviews and AI Mode in France, adding AI-generated summaries, follow-up questions, multimodal input, and links to the web in another major European search market. Google also says AI Mode can break a question into subtopics and run several searches in parallel. For a multilingual business, visibility may therefore change by query, engine, language, market, interface, and even the route a user takes through follow-up questions.
So the useful commercial question is not simply, āAre we visible in AI?ā It is more specific: for which buyer questions, in which markets, through which engines, and in what form do we appear? And is the pattern stable enough to act on?
AI answers change where attention goes
A 2026 Microsoft Research eye-tracking study examined search interfaces that place generative results above conventional listings. The familiar concentration of attention in the top-left area remained, and participants still engaged with traditional links. But they spent significantly more attention on the generative content.
That makes inclusion in or alongside an AI-generated answer commercially relevant. It does not make every kind of inclusion equal.
A brand name in the prose is not the same as a cited page. A citation is not the same as a recommendation. And none of those signals, on its own, proves that somebody visited the website or completed a commercial action.
A useful dashboard keeps those outcomes separate. Once they are folded into a single visibility score, the number may look cleaner, but it becomes much harder to diagnose or act on.
Portfolio averages can hide the real problem
DeltaV Digitalās 2026 study tracked 21,075 responses across ChatGPT, Perplexity, Gemini, Google AI Overviews, and Google AI Mode. It then analysed 25,337 citations to the most-cited pages across eight industries. The striking part was not one universally successful page type. It was the variation: citation patterns differed markedly by industry, and no page type dominated across the portfolio.
The study also distinguishes retrieval from citation. A system may retrieve a page while assembling an answer without citing that page in the final response. For a team deciding what to fix, this is an important distinction.
If a page appears to be retrieved frequently but is seldom cited, the issue may concern its authority, clarity, or usefulness as a supporting source. If it is absent altogether, the likely investigation shifts towards accessibility, relevance, or discovery. Those are different jobs for different people.
The study is also clear about its limits. Its prompts were commercially oriented rather than a random sample of all AI usage. Classification was automated, the analysis focused on top-cited sources, and the sample covered eight organisations in eight industries. These qualifications do not undo the findings. They show the level of documentation a credible visibility report needs: define the population, unit of analysis, sampling choices, and boundaries before interpreting the charts.
What the dashboard should record
| Measurement layer | Evidence to retain | Question it helps answer |
|---|---|---|
| Test conditions | Prompt, market, language, engine, interface, model or version label, timestamp, and run number | Can we reproduce the test and make a fair comparison? |
| Brand mention | The brand appearance and its surrounding context | Are buyers being exposed to the brand? |
| Citation | Cited domain, URL, answer passage, and citation position where observable | Which sources are supporting the answer? |
| Recommendation | Explicit inclusion, comparative position, qualification, and sentiment | Is the system naming us or actually proposing us? |
| Referral | Verified visits attributable to an AI or search surface where analytics permit | Is that exposure producing website activity? |
| Crawler activity | Bot identity, requested URL, timestamp, response status, and robots treatment | Can the relevant systems access the material? |
All six layers can belong in one reporting system. They should not be compressed into one unexplained score. A crawler request is not a citation; a citation is not a recommendation; and a recommendation is not a referral. Each one describes a different point between technical availability and commercial response.
Set the method before designing the interface
Begin with the decisions the report is supposed to support. A content team may need to identify buyer questions that lack credible supporting pages. Communications may want to see which third-party publications recur in generated answers. A technical team may be looking for access problems. Management may simply need to know whether visibility is improving in priority markets.
Those decisions tell you what belongs in the data. They should come before chart selection.
The next job is to establish a fixed prompt inventory. Prompts should reflect genuine buyer questions and be classified by market, language, product, journey stage, and intent. Freeze a baseline group for trend reporting. New prompts can be introduced as a separately versioned cohort, but silently rewriting the historical set makes any apparent improvement impossible to interpret.
Point Visibleās published methodology offers several useful controls: a fixed set of buyer-intent prompts, engine-level separation, repeated runs, and a log containing model information and timestamps. It also acknowledges non-determinism, model changes, and personalisation. Companies do not need to copy its precise cadence or engine mix. The useful feature is that the protocol is visible enough for another person to inspect.
Repeated testing matters because identical prompts can produce different answers. The basic observation should therefore be a test cell: one prompt, one engine, one market, one language, one timestamp, and one run number. Store the response before calculating summaries. If somebody challenges a quarterly percentage, the team should be able to open the observations behind it and see what moved.
Record retrieval and citation as separate events
The 2026 paper by Kakimov and colleagues proposes a system-agnostic framework that treats retrieval and citation as observable processes across query-document pairs. It also introduces measures conditioned on factors including rank and document provenance. For commercial reporting, the implication is straightforward: generative visibility cannot be treated as a renamed version of conventional search ranking.
The data model needs to retain the relationship between a query, the observed answer, the documents surfaced or cited, and the conditions under which the test ran. Some interfaces will not expose every retrieval step. When that happens, the report should record what was observable and what was not. It should not fill the gap with an assumption.
This discipline also stops unlike outcomes being compared as if they were equivalent. A cited first-party product page, a third-party review that recommends the brand, and an uncited brand mention may all be useful. They are still three different results.
Do not disguise breaks in the series
Model and interface changes belong in the reporting record. If an engine changes substantially, mark the break and consider creating a new comparison baseline. The same applies when the team adds a language, changes the test location, replaces an engine, or modifies prompts.
Governance needs named owners too. Decide who approves prompts, reviews classifications, investigates anomalies, sets the retention period for raw responses, and determines when a methodology change requires a new version. Every chart should lead back to the observations used to produce it.
The finished dashboard may have fewer headline numbers. That is usually a fair trade. It can report citation presence by engine, recommendation frequency by market, consistency across repeated runs, cited-source mix, verified referrals, and crawler access without presenting those signals as one interchangeable measure.
Make it useful for allocating work
The point of an AI visibility dashboard is to help decide where the next investment belongs: technical access, first-party content, third-party authority, localisation, or conversion measurement. That decision is only defensible when the measurement pipeline keeps the context and uncertainty attached to each result.
For GrN.dk, the practical role is to help define buyer prompts by market and language, establish repeatable test conditions, retain timestamped raw evidence, document platform changes, and connect the separate signals in a governed reporting pipeline. Greg can provide the hands-on coordination needed to turn that specification into a reporting process teams can inspect, challenge, and use when assigning work.
Related on GrN.dk
- Googleās AI Search toggle needs a test plan, not a gut decision
- Search Console Can See Social PostsāYour Reports Need a New Map
- When AI writes JSON, one bad field can break the workflow
Need help with this kind of work?
Build a reporting system you can defend Get in touch with Greg.