Control over AI
Blog
Trust and proof 8 min read

A dashboard is not evidence

Numbers show something happened. Evidence requires showing that a control was designed, that it ran, and what it says nothing about.

Auditor reviewing an accountability view
Quick answer

A dashboard shows what was measured. Evidence is different: it answers an auditor's question, and that question is not "how many" but "what is this evidence of, and what does it not cover". Three things are usually missing. No control is named, so nobody knows which obligation the figure supports. No scope is stated, so 97 per cent may describe a tenth of your usage. And nothing is red, which is precisely why the report fails to convince.

01

A number with no control attached substantiates nothing

02

Without scope, nobody knows what a percentage describes

03

A report that is always green is marketing, not evidence

04

Visibility and intervention are two evidence chains and must not be summed

05

What you cannot substantiate belongs inside the view, not outside it

A DPO puts a dashboard in front of the auditor. Three charts: detections per month, distribution across applications, and a figure sitting at 97.

The auditor looks at it for ten seconds and asks one question: what is this evidence of?

That is where the conversation usually stalls, and the numbers are not the problem.

A figure is not evidence

Evidence is a relationship between three things: an obligation, a control, and an observation showing the control worked.

A dashboard typically supplies only the third. There were 312 detections and 97 per cent of them were handled. That is an observation, and it is not yet attached to anything.

The missing step: which article of which framework does this support, and above what line is it enough? Until somebody writes those two down, the reader has to guess. Auditors do not guess.

Three things that are usually absent

The control. A handling percentage might support data minimisation under Article 5(1)(c), or appropriate security under Article 32, or neither. It depends how you set it up. Write it next to the figure, with the article number, or the figure does no work.

The scope. The big one. If detection runs in the supported applications, then 97 per cent describes those applications. Not every AI service your organisation uses. Without that sentence the percentage is not wrong but it is misleading, which is worse, because it does not survive the first question.

The red rows. Every organisation has controls that measurement cannot substantiate. Someone using an AI tool on their phone falls outside the picture. Whether a processing agreement was signed is not something your system knows. Leave those out and the view is incomplete, and that is exactly what gets probed.

Two evidence chains you must not add together

The distinction most reporting misses, and the one that pays off most.

There is a visibility layer: which AI services are in use, how many carry a decision, what risk band applies. It sees a great deal, across hundreds of services, and can intervene nowhere.

And there is an intervention layer: sensitive data getting a highlight before it is sent. That layer can intervene, and only inside the applications it supports.

The temptation is to produce one number, because it reads better. Resist it. The two layers have different scope, and a blended figure hides precisely the distinction a regulator asks about. What you want is an accountability view stating, per control, which layer the evidence comes from. That both layers can be measured without attaching people to them is covered in AI monitoring versus employee monitoring.

Written out, it looks like this:

Dashboard ยท Accountability 3 of 8 substantiated

What we can substantiate over the period 1 May to 29 July (90 days).

47 tools
Visibility
43 of 47 carry their own decision
92%
Seat coverage
118 of 128 seats active, threshold 85%
84%
Intervention volume coverage
share of measured AI usage inside a supported app, threshold 90%

Substantiated from visibility

all 47 observed tools
  • Art. 30 Record of processing activities

    Every observed AI tool is recorded with vendor, hosting region and training policy, and carries a decision from the organisation.

    Substantiated
  • Art. 35 DPIA input per AI processing activity

    No observed tool is left unreviewed. Four are currently sitting without a decision.

    Partial
  • Art. 5(2) Accountability

    Uninterrupted recording across the full 90-day reporting period.

    Substantiated
  • Art. 24 Controller responsibility

    Requires a record of who took which policy decision and when. We do not capture that history yet.

    Not substantiated
  • Art. 28 Processor agreements

    The catalog knows whether a vendor offers one, not whether this organisation signed it. That is register input we do not hold.

    Not substantiated

Substantiated from intervention

84% of measured AI usage
  • Art. 5(1)(c) Data minimisation

    Critical data is handled before it is sent. 97% this period, against a 95% bar.

    Substantiated
  • Art. 32 Appropriate technical measures

    Detection active on the supported applications with handling above the bar. Volume coverage stays under the 90% threshold.

    Partial

Outside what we measure

named, not omitted
  • Art. 33 Breach notification

    BeeSensible provides the signal, not a notification process toward a regulator or data subject.

    Not applicable
  • Substantiated: every condition met
  • Partial: measurement exists, not every threshold met
  • Not substantiated: no measurement available

This view does not establish compliance with any law. It shows which controls can be substantiated by measurement and which cannot. The assessment remains with the controller.

The accountability view from the dashboard. Figures are illustrative. Two things are deliberately visible: which controls cannot be substantiated, and which layer the evidence comes from.

Why there should be red in it

Look at what is not green in that view. Two GDPR controls sit at not substantiated, one at partial, and one is explicitly declared out of scope.

That is not a shortcoming of the report, that is the report. A view that is green everywhere tells an auditor: this party removed the difficult rows. That becomes the first conclusion, and everything that is there loses value with it.

There is a second reason. Writing down what you cannot substantiate tells your own organisation where the next piece of work sits. "We do not record who took which policy decision and when" is a red row and a concrete backlog item at the same time.

What "above the bar" does and does not mean

One more thing that has to be explicit, because it goes wrong in audit conversations otherwise.

Thresholds in an accountability view are self-chosen. No law states that 95 per cent of critical highlights must be handled. What the law asks for is an appropriate measure, and the threshold is your interpretation of that.

Not a weakness, provided you write it that way. "We hold ourselves to 95 per cent and reached 97" is a defensible position. "We meet the standard" is not, because no such standard exists and the auditor knows it.

The same applies to what you deliberately do not norm. Ignoring an amber highlight can be a correct judgement: the employee knows the context and the system does not. Only the red ones carry an instruction, so only those deserve a threshold. The amber figure is context.

The four questions it reduces to

Per control, in this order:

  1. How is it designed? What happens exactly, and why would that help?
  2. Did it run during the period? Not "is it enabled", but did it run uninterrupted across the period you are reporting on.
  3. What was the effect? The numbers, with their scope attached.
  4. Where does it stop? The most important, and the only one nobody volunteers.

Answer those four per control and you have evidence. Have only chart three and you have a dashboard.

Which underlying figures feed this is covered in what are all those AI metrics for.

That gap is also why we built the accountability view as a translation table rather than a report: from what was measured to what an auditor asks, including the rows where the answer is no. A view that shows its own boundary is more useful than one that impresses.

FAQ

Common questions

What is the difference between a dashboard and evidence?

A dashboard shows measurements. Evidence links a measurement to an obligation and states what it does not cover. Without those two additions, a figure is an observation an auditor cannot use.

Why is an all-green report suspicious?

Because no organisation can substantiate everything. If a view shows no red anywhere, it usually means the unsubstantiable controls were left out rather than satisfied. An auditor looks for exactly that omission.

What do you mean by scope on a percentage?

That a figure needs a sentence saying what it describes. If detection runs in the supported applications, then 97 per cent handled is a statement about those applications, not about all AI usage in the organisation. Without that sentence the percentage misleads.

Why can't visibility and intervention be combined?

Because they are two evidence chains with different scope. Seeing hundreds of AI services is not the same as being able to intervene in a handful of applications. One blended figure hides the very distinction a regulator wants to see.

So what does an auditor ask for?

Four things per control: how it is designed, whether it ran during the period, what effect it had, and where it stops. The fourth answer matters most, because it shows you know the limits of your own control.