Consulting innovation · AI quality engineering
A red, amber, green verdict is a confident answer. Give one to a team that cannot yet defend the data underneath it and you have sold them false confidence, which is worse than none. For a system not yet ready to be scored, the honest artefact is a set of questions: what to clarify first, and the impact if the answer is wrong.
The instinct in a quality readout is to put a colour on the system. It photographs well and it feels like progress. But a verdict is only as good as the data behind it, and most teams starting with AI have not yet measured the things a verdict would be scored against. Paint that system green and you have hidden the risk. Paint it red and you have graded homework the team was never told to do.
Practitioner feedback on our own engine named the failure mode exactly: showing a client a traffic light is selling them the tree before they have learned to crawl. He who wants to fly must first learn to crawl. The team does not need a score they cannot act on. They need to know which questions matter for their system, in what order, and what it costs to get each one wrong. The most damaging outcome in early AI assurance is not a missing answer. It is a confident wrong one.
The clarification view is the same coverage gap audit, rendered differently. No new inference, no model in the loop, nothing softened for presentation. Every question is a deterministic reading of a gap the audit already found:
A system is shown questions only while it is flagged not yet ready to be assessed, and only when a real audit is available to build them from. The moment the data can stand behind a verdict, the surface reverts to the verdict. The questions are the crawl. The score is the walk. Nothing pretends the team is flying before it is.
Take the customer-facing posture from the coverage gap audit worked example: a supervised, advisory chatbot at a regulated financial institution, 24 of 25 attributes material, 19 gaps. As a verdict that is a wall of red the team cannot yet act on. As questions it is a worked list, ranked worst first. Two of the nineteen:
Can you evidence that advice does not skew across protected groups, and how would you know?
Can you show the system refuses unsafe or toxic requests?
Same audit, same system, same tooling as the traffic-light view. The difference is the artefact handed over: a list the team can start answering on day one, with the cost of each wrong answer already attached.
It is not a softer verdict, and it is not a narrative written by a language model. Every question and every impact band is a lookup over the same versioned ontology the audit uses, so two people can check it line by line. It does not sit beside the traffic light as a second opinion. For a system that is not ready, it replaces the verdict, because a colour on undefended data is the thing the exercise exists to avoid. Once the answers are in, the verdict earns its place back.
Clarification questions are a render mode of Crystal Ball over the coverage gap audit produced by Prism. Questions first, for a team learning to crawl. Measurement next, once they know what to measure. A verdict last, once the data can defend one.