Evaluation & quality

Explainability

XAI

Explainable AI

The degree to which people can be given understandable reasons for a model's output: which inputs and factors led to this result, in terms that the person affected or the person responsible can act on.

LOAN APPLICATION · DECLINEDillustrative pointsFactors that moved the score mostDebt at 46% of income−38One late payment−21Credit history: 14 months−14Regular income+12With debt below 35% of income: approvedThe model's own accountplausible text, often incompletePost-hoc explanationwhich inputs moved the result, by how much?Reading the mechanisminside the model: still a research fieldThe chart on the left is a post-hoc explanation: it approximates the model's behaviour from outside.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/explainability

In plain terms

A loan officer who says no can be asked why, and the answer can be checked against the file. A model that says no has only produced a number. Explainability is the set of methods that turn that number into reasons a person can use: these three factors weighed most, and this change would have altered the result. With today's large models such reasons are reconstructions made from outside; nobody can yet read the full calculation inside.

Why it matters

Wherever a decision affects a person, somebody will ask why: the customer, an auditor, a court. European data-protection law and the EU AI Act provide, in defined cases, a right to an explanation of decisions made with automated or high-risk systems. Explanations help teams find errors and bias. Two limits apply. For high-stakes decisions a simpler, transparent model can be the better choice despite some loss of accuracy. And a language model's own account of its reasoning is fluent text that can omit what drove the answer, so it cannot serve as an audit trail.

Example

A bank's credit model declines an application. Next to the decision the adviser sees the factors that pushed the score down most: debt at 46% of income, one late payment in the past year and a credit history of 14 months. She can tell the customer what would change it: with debt below 35% the application would have passed. When one factor dominates thousands of rejections, the risk team has a concrete point to review.

Most often confused with

XAI vs. Interpretability

XAIReasons for one output, usually given after the fact
InterpretabilityUnderstanding how the model works inside

The words are often used as synonyms. Where they are separated, interpretability means that the mechanism can be understood: a short decision tree is interpretable by design, and research on neural networks tries to read their internal features. Explainability is the wider, practical goal of giving a person usable reasons, often through an approximation added afterwards. An explanation can be convincing and still be unfaithful to what the model did.

Origin: The abbreviation XAI spread with a DARPA research programme on explainable AI that was announced in 2016.

Under the hood

Three levels. Interpretable models (linear models, small decision trees, scorecards) can be read directly. Post-hoc methods explain a black box from outside: feature attribution with SHAP or LIME, saliency maps, counterfactual explanations (the smallest change that flips the outcome) and surrogate models. They approximate, can disagree with one another and should be tested for fidelity and stability. Mechanistic interpretability works on the internals of neural networks, tracing learned features and circuits; it remains a research field. For language models the practical substitutes are citations, traces of tool calls and structured rationales; a model's visible reasoning is useful evidence, but it has been shown to omit decisive factors. Regulation: the GDPR gives a right to meaningful information about the logic of solely automated decisions; Türkiye's KVKK gives a right to object to an adverse result produced solely by automated analysis; the EU AI Act adds a right to an explanation of decisions taken with high-risk systems. Documents such as model cards describe the model as a whole.

Written by Mehmet Erkek · Last updated: