← All articles

004 / AI EXPLAINABILITY

10 min read
Why Explainability Needs Evidence, Not Better Prose

Artificial intelligence has become exceptionally good at explaining itself.

Ask a modern model why it reached a conclusion and it can produce a polished answer in seconds. The language may be clear, structured and persuasive. It can describe assumptions, summarise reasoning and present a result in exactly the tone the reader expects.

That creates a dangerous possibility inside organisations: we may begin to confuse a convincing explanation with an evidential one.

The two are not the same.

If an AI recommends delaying a customer order, changing a supplier, reallocating inventory or approving additional expenditure, the important question is not whether it can tell a good story about the decision.

The important question is whether the decision can be traced back to the operational facts that made it necessary.

Explanation Is a Presentation Layer

Humans naturally value explanations because they compress complexity.

A financial analyst can take thousands of transactions and explain why margin fell. An engineer can describe why a system failed. A doctor can connect symptoms to a diagnosis. A logistics manager can explain why a delivery missed its target.

In each case, the explanation is useful because it sits on top of something more concrete: records, observations, measurements, events and relationships.

AI should be treated the same way.

The language layer can make evidence easier to understand, but the explanation itself should not become the evidence.

Consider two responses to the same question.

"Customer Order 1842 is at risk because the inbound motor-controller shipment has been delayed by six days."

And:

"Customer Order 1842 is at risk because Shipment 771 contains the component batch allocated to Production Order 914. Current usable inventory cannot cover the scheduled requirement, moving projected completion beyond the customer commitment date."

The second answer contains more structure, but even that is still only prose.

A trustworthy system should allow the user to inspect Shipment 771, the batch it contains, the inventory position, the production requirement, the customer order and the commitment date.

The sentence explains the conclusion.

The underlying objects establish it.

A Better Story Can Still Be Wrong

Generative models are unusually good at constructing coherent narratives from incomplete information.

That capability is one of their strengths. It allows them to synthesise large amounts of material, bridge terminology differences and turn fragmented information into something humans can use.

It also means that presentation quality is no guarantee of factual quality.

Suppose an enterprise AI finds that a delivery is running late. It sees a supplier message mentioning manufacturing difficulties, an internal note discussing capacity constraints and a historical report describing previous shortages.

A plausible explanation might attribute the delay to supplier capacity.

The actual cause could be a customs hold that occurred twelve hours later.

The first explanation may be articulate, reasonable and entirely wrong.

In operational environments, plausibility has a limited shelf life.

What matters is whether the conclusion follows from the current evidence.

Evidence Should Have Identity

One of the simplest ways to strengthen machine explanations is to make the evidence itself addressable.

A useful operational claim should be capable of pointing to identifiable things: an order, shipment, contract, event, asset, policy, approval or observation.

That creates a meaningful difference between:

"The supplier is experiencing delays."

and:

"Shipment SHP-204 was scheduled to arrive on 24 August and its latest authenticated tracking event moved the expected arrival to 29 August."

The first statement may have come from anywhere.

The second can be inspected.

Stable identity matters because enterprise reality contains enormous numbers of similar objects. A company may have thousands of shipments, hundreds of suppliers and millions of transactions. Saying that "the shipment" is late becomes meaningless once several systems are trying to reason about the same world simultaneously.

Evidence needs somewhere to point.

Evidence Also Needs Time

A fact without time can be actively misleading.

Imagine an AI answering:

"There are 8,000 units available in the Manchester warehouse."

That statement might have been correct when the inventory system was last updated. Since then, 3,000 units may have been allocated to production and another 1,000 dispatched elsewhere.

The value was true.

It is no longer operationally useful.

Enterprise AI therefore has to care about more than whether a record exists. It needs to understand when the record was valid, whether a newer state supersedes it and how quickly that category of information becomes stale.

This applies across the organisation.

A supplier may have been approved last month and suspended yesterday. A contract may have been active when it was indexed and subsequently replaced. A customer commitment may have been renegotiated. A production schedule may have changed ten minutes ago.

When AI explanations ignore time, old truth can quietly become new fiction.

Provenance Changes the Nature of Trust

There is another question behind every operational fact:

Where did it come from?

A delivery date reported directly by a carrier is not the same thing as an estimate in a planning spreadsheet. A contract clause extracted from the signed agreement is not equivalent to a salesperson's note describing what they remember agreeing. A physical sensor observation is different from a forecast produced by a model.

All of these may be useful.

They should not automatically carry the same evidential weight.

Provenance tells us the origin of a claim and the chain through which it entered the system. That might include the source system, timestamp, transformation, responsible process and any subsequent derivation.

This becomes particularly important when AI combines evidence from many places.

The system should be capable of distinguishing between:

  • directly observed facts;
  • authoritative business records;
  • derived calculations;
  • forecasts and estimates;
  • human assertions;
  • AI-generated interpretations.

Without that distinction, the user receives one seamless answer while the evidence underneath may contain radically different levels of certainty.

Derivation Should Be Visible

Some of the most valuable enterprise conclusions will never exist as a single stored record.

Nobody may have written:

"This disruption exposes £2.4 million of customer commitments."

The number may need to be calculated from multiple operational facts.

A disruption affects a shipment. The shipment contains components. Those components support production. Production supports customer orders. Those orders have commercial values and contractual deadlines.

The conclusion is derived.

Derived conclusions are not inherently less trustworthy than recorded facts. In many cases they are far more useful.

But the derivation should be inspectable.

If a system reports £2.4 million of exposure, a user should be able to ask which orders contributed to the total, which assumptions were applied and which dependency paths connected the original event to those orders.

This gives explainability a stronger foundation.

Instead of asking the model to describe why the number makes sense, we allow the user to inspect how the number came into existence.

Citation Is Not the Same as Evidence

Enterprise AI systems increasingly attach citations to generated answers.

That is a positive development, but citation alone does not solve explainability.

A citation proves that some source material was retrieved. It does not automatically prove that the source supports the claim being made.

A model might cite a fifty-page contract while misunderstanding the clause that applies. It might cite an inventory report containing the right product but the wrong location. It might cite a policy that has already been superseded.

The existence of a citation therefore answers only one question:

Where did the model look?

A stronger evidence model must also answer:

Why does this source support this conclusion?

That requires structure, identity, semantics and often relationships between multiple pieces of evidence.

For simple factual questions, a document reference may be enough.

For operational decisions, the evidential burden becomes much higher.

Confidence Scores Do Not Replace Evidence Either

Another tempting solution is to attach a confidence score to the model's answer.

"92% confident."

The number looks precise. It can also create a false sense of measurement.

Confidence can be useful when it has a defined statistical meaning, a known calibration process and a clear relationship to observed outcomes. A generic model confidence score does not tell the user which operational facts were established correctly.

A system can be highly confident while reasoning from stale information.

It can be uncertain while operating on perfectly valid evidence.

For enterprise systems, confidence should be treated as an additional signal rather than a substitute for provenance.

The user still needs to know what the conclusion rests on.

Explainability Becomes Critical When AI Can Act

The consequences of weak explainability become more serious as AI gains agency.

If a chatbot produces a questionable answer, a human may notice and ignore it. If an agent begins changing production priorities, purchasing inventory, altering schedules or communicating commitments to customers, the cost of an unsupported conclusion rises sharply.

Consider an agent recommending that £80,000 be spent on expedited freight.

The approval interface could display a polished AI summary explaining that the expenditure protects a strategic customer.

That may be useful.

The approver should also be able to inspect:

  • the disruption that triggered the recommendation;
  • the affected shipment and inventory;
  • the production dependency;
  • the customer commitment at risk;
  • the estimated cost of doing nothing;
  • the assumptions behind the alternative;
  • the authority required to approve it.

The decision then becomes reviewable on two levels.

The AI explains the situation in human language.

The system exposes the evidence in operational form.

That separation matters because the approver is ultimately taking responsibility for the action.

The Audit Trail Starts Before the Decision

Explainability is often discussed at the moment an AI produces an answer.

In business systems, the more interesting question may come later.

Three months after an important decision, an auditor, customer, regulator or executive might ask:

Why did we do this?

By then the original model may have changed. The prompt may have changed. The surrounding data may have changed. The employee who approved the action may have moved to another role.

A durable system should still be able to reconstruct what happened.

What was known at the time?

Which version of the evidence was used?

What recommendation was produced?

Which assumptions were applied?

Who reviewed it?

Who had authority?

What action followed?

What happened afterwards?

This is why evidence should survive independently of both the AI model and the human conversation surrounding it.

A transient explanation is useful in the moment.

An evidential decision chain becomes institutional memory.

Better Models Will Make This More Important

There is a curious effect as AI improves.

Weak models make users cautious because their limitations are obvious. Strong models are more persuasive. Their answers are clearer, their reasoning appears more coherent and their mistakes become harder to notice.

The better the model becomes at sounding correct, the more important external verification becomes.

This is especially true inside organisations, where many decisions involve specialised knowledge unavailable to the person reading the answer. An executive may not know the warehouse data. A logistics manager may not know the contract terms. A procurement analyst may not know the production constraints.

The AI can bridge those domains.

The evidence system must allow the organisation to verify that bridge.

Explainability Should Be Architectural

Much of the current discussion around explainable AI focuses on making the model communicate its reasoning more clearly.

That is useful, but enterprise explainability should extend much deeper into the architecture.

Every important operational conclusion should ideally have a path back to authoritative or explicitly qualified evidence.

Every derived value should have an inspectable derivation.

Every recommendation should expose its assumptions.

Every action should record the authority under which it occurred.

Every important state transition should leave history behind.

The language model then has an enormously valuable role: translating that structure into an explanation a human can understand.

That is a better division of responsibility.

Let AI make complexity legible.

Let the system make claims verifiable.

From Persuasion to Proof

Generative AI has given software an unprecedented ability to explain.

The next challenge is making sure those explanations remain anchored to reality.

In enterprise environments, trust will not come from increasingly eloquent answers alone. It will come from systems where claims can be challenged, sources can be inspected, relationships can be followed and decisions can be reconstructed.

That creates a much stronger standard than "the AI provided an explanation."

It asks whether the explanation has somewhere solid to stand.

A trustworthy operational system should never require us to accept a machine conclusion simply because the prose sounds convincing.

The explanation tells us what the system thinks. The evidence shows us why we should believe it.