Consulting innovation · AI quality engineering

Clarification questions

A red, amber, green verdict is a confident answer. Give one to a team that cannot yet defend the data underneath it and you have sold them false confidence, which is worse than none. For a system not yet ready to be scored, the honest artefact is a set of questions: what to clarify first, and the impact if the answer is wrong.

Why

The instinct in a quality readout is to put a colour on the system. It photographs well and it feels like progress. But a verdict is only as good as the data behind it, and most teams starting with AI have not yet measured the things a verdict would be scored against. Paint that system green and you have hidden the risk. Paint it red and you have graded homework the team was never told to do.

Practitioner feedback on our own engine named the failure mode exactly: showing a client a traffic light is selling them the tree before they have learned to crawl. He who wants to fly must first learn to crawl. The team does not need a score they cannot act on. They need to know which questions matter for their system, in what order, and what it costs to get each one wrong. The most damaging outcome in early AI assurance is not a missing answer. It is a confident wrong one.

How

The clarification view is the same coverage gap audit, rendered differently. No new inference, no model in the loop, nothing softened for presentation. Every question is a deterministic reading of a gap the audit already found:

  1. One gap, one question. Each material attribute with no measurement evidence becomes a question phrased for the person who owns the outcome, not the person who owns the pipeline.
  2. Impact band, not a score. The gap's severity, read from the published rubric (regulatory relevance, agency, exposure class), becomes the stakes on the question. It ranks the list. It is not a grade on the system.
  3. The obligations already mapped. Where the ontology ties a gap to a regulatory clause (MAS FEAT, TRM, EU AI Act, NIST AI RMF), the question carries it, so the team sees which regulator cares before they decide how hard to look.
  4. Where accountability lands. When a mapped rule puts the obligation on a named role rather than the model, the question says so.
  5. What to measure. The catalogue's measurement tools for the attribute, so the answer to the question is a concrete next step, not more reading.

A system is shown questions only while it is flagged not yet ready to be assessed, and only when a real audit is available to build them from. The moment the data can stand behind a verdict, the surface reverts to the verdict. The questions are the crawl. The score is the walk. Nothing pretends the team is flying before it is.

What it looks like

Take the customer-facing posture from the coverage gap audit worked example: a supervised, advisory chatbot at a regulated financial institution, 24 of 25 attributes material, 19 gaps. As a verdict that is a wall of red the team cannot yet act on. As questions it is a worked list, ranked worst first. Two of the nineteen:

Critical · impact if unanswered

Can you evidence that advice does not skew across protected groups, and how would you know?

Why it carries weight
Obligations already map here: MAS FEAT fairness criterion, and the Tripartite Guidelines on Fair Employment. A supervised system advising consumers in a regulated setting inherits them directly.
Where accountability lands
Individual accountability sits with the business owner, not the model.
What to measure
Fairlearn, IBM AI Fairness 360.
High · impact if unanswered

Can you show the system refuses unsafe or toxic requests?

Why it carries weight
No obligation is mapped in the catalogue for this attribute. That is a gap in our own coverage, not evidence that none exists. Confirm it with compliance rather than assume.
What to measure
Garak, DeepEval safety and toxicity metrics.

Same audit, same system, same tooling as the traffic-light view. The difference is the artefact handed over: a list the team can start answering on day one, with the cost of each wrong answer already attached.

What it deliberately is not

It is not a softer verdict, and it is not a narrative written by a language model. Every question and every impact band is a lookup over the same versioned ontology the audit uses, so two people can check it line by line. It does not sit beside the traffic light as a second opinion. For a system that is not ready, it replaces the verdict, because a colour on undefended data is the thing the exercise exists to avoid. Once the answers are in, the verdict earns its place back.

Where it fits

Clarification questions are a render mode of Crystal Ball over the coverage gap audit produced by Prism. Questions first, for a team learning to crawl. Measurement next, once they know what to measure. A verdict last, once the data can defend one.

← Coverage gap audit All innovations →