SHAP and LIME: how to explain artificial intelligence

Share Article

A person is denied a loan. Or a scholarship. Or access to a job offer. Behind that decision is an artificial intelligence system that has processed dozens of variables and returned a score. The person asks the only question anyone would ask: why?

Article 86 of the AI Regulation recognizes their right to an answer. And that's where the problem begins, because explaining an algorithmic decision isn't as simple as opening up the engine and pointing out the faulty part.

To solve this, the technical community has been working for years on two tools that have become the de facto standard: LIME and SHAP. It's important to understand them well, because they will increasingly appear in compliance reports and audits. And it's also important to understand something that almost no one is saying out loud: that we don't know if they are sufficient.

The black box problem

Many of the models that currently make or support important decisions are not equations that a human can read. A deep neural network or a set of thousands of decision trees works reasonably well, but they don't come with a built-in explanation. We know what goes in and what comes out. We don't know, in an intuitive sense, why or how it does it each time.

This is called the black box problem. And it's not solved by asking the model to talk, because the model doesn't have a narrative of itself. What we do, instead, is construct an explanation from the outside, by observing its behavior. This is the origin of what we call "post-hoc" explainability: explanations that come after, and from outside, the decision has already been made.

That nuance—after and from outside—seems technical. It is, in reality, the legal heart of this entire article.

LIME: Looking Up Close Instead of Looking Far Away

LIME stands for "Local Interpretable Model-Agnostic Explanations," and its concept is remarkably elegant.

If a model is too complex to fully understand, let's stop trying to fully understand it. Let's focus on a single case: the specific decision affecting this specific person. Around that case, the model behaves in a much simpler way. In a local way.

The metaphor is perhaps that of a mountain range. No one can describe the complete shape of a mountain range with a simple formula. But if you stand at a precise point on the slope and look only at the square meter you're standing on, that surface looks quite similar to an inclined plane. And an inclined plane can be described. And perhaps that's the only part you really need.

That's what LIME does. It takes the case you want to explain, generates many slightly modified variations of that case, asks the model what it would respond to each one, and with those responses builds a simple model—usually a straight line—that mimics the behavior of the original model at that point. Then it shows you the coefficients of that line: which variables pushed the decision one way and which the other.

It's intuitive and works with any model, without needing to open it. But it's important to remember one thing: what LIME shows you isn't the model. It's a local imitation of the model. A small-scale substitute.

SHAP: Distributing Credit Fairly

SHAP stands for "Shapley Additive Explanations" and its origins lie not in computer science, but in game theory. In 1953, the mathematician Lloyd Shapley posed a seemingly mundane problem: if several people collaborate and obtain a joint benefit, how can that benefit be distributed fairly among them?

His answer was to measure, for each player, their average contribution upon joining the team, considering all possible combinations of teammates. If the team performs the same with you as without you, your contribution is zero. If it collapses without you, your contribution is significant.

SHAP applies this idea to models. Each variable represents a player. The prediction is the benefit to be distributed. And the SHAP value for each variable is its average contribution to moving the decision from a starting point—what the model would predict "on average," without knowing anything about you—to the decision that was actually made with you.

It has a property that lawyers find very attractive: it's additive. The sum of all the contributions, starting from the base value, gives exactly the final prediction. Nothing is left over, nothing is missing. On paper, a complete distribution with no surplus.

The problem lies in calculating all possible combinations of variables, since it becomes unmanageable once there are a few dozen. That's why, in practice, SHAP almost always "approximates." And some of these approximations assume that the variables are independent of each other, something that rarely happens in the real world—where postal code, income, and credit history are intertwined.

What they both share

This is where it's worth pausing, because the two techniques share exactly the same two characteristics. And these are the two that the law has yet to fully grasp.

They are "post-hoc": they don't describe the system's internal reasoning, but rather reconstruct a plausible explanation by observing its behavior from the outside. And they are local: they explain a decision—yours—not the system's general logic.

In other words, SHAP and LIME don't teach you how the model thinks. They offer you a coherent story about how it might have thought, in your case, if it were simpler than it is.

What Article 86 really asks for

It's worth reading carefully, because almost no one does.

Article 86 recognizes the right of a person affected by a decision made by the deployment manager based on the outcome of a high-risk system in Annex III—with the exception of the systems in point 2—that produces legal effects or significantly affects them in a similar and detrimental way to their health, safety, or fundamental rights, to obtain "clear and meaningful explanations about the role of the AI system in the decision-making process and about the main elements of the decision taken.".

There are three things in that sentence that are often overlooked.

First: the obligation falls on the user of the system, not the manufacturer. The financial institution, not the provider of the model.

Second: a technical dump is not required. "Clear and meaningful" explanations are requested, and these are requested regarding two distinct things: the role of the system in the procedure—that is, how much weight the machine carried in the decision, who actually made the decision, what human oversight there was—and the main elements of the decision.

Third: the article is subsidiary. It only applies to the extent that this right is not already provided for in another EU regulation. And that's where the GDPR comes in.

It's worth adding a current point. Following the agreement on the so-called Digital Omnibus, endorsed by the European Parliament in June 2026 and confirmed by the Council shortly thereafter, the obligations for high-risk systems in Annex III have been postponed to December 2027. Since the right under Article 86 applies precisely to these systems, its practical enforceability is linked to this timeline. There is more time. There is also more uncertainty.

So, do SHAP and LIME comply with Article 86?

Well, we don't know yet.

Currently, there is no supervisory pronouncement stating that applying SHAP or LIME satisfies Article 86. There is no guidance from the Commission, no criterion from the AI Office, and no consolidated position from national authorities. Harmonized standards are still lagging behind on this point. No one has certified that a contribution graph constitutes a legally valid explanation.

What can be said, however, is this: Technically, both tools allow us to identify which variables were relevant in a specific case and with what weight. This comes close to what Article 86 calls "the main elements of the decision taken." It is, undoubtedly, much better than silence.

The Four Flaws

The first flaw is that they are approximations. LIME constructs a surrogate model; SHAP estimates values that would be impossible to calculate exactly. What is given to the affected person is not, strictly speaking, the reason for the decision: it is the best available reconstruction of that reason. A lawyer knows that a proven fact is not the same as a highly convincing indication.

The second flaw is that they are local. They explain the outcome, but say absolutely nothing about the role the system played in the procedure. And that is, literally, half of what Article 86 requires. No SHAP value will tell you if there was effective human oversight, if the human could deviate from the outcome, or if they simply signed below.

The third flaw is that they are not causal. Knowing that the variable "seniority" accounted for 12% doesn't answer the question that truly matters to the person: what would have had to be different for the decision to have been different? That's an explanation that neither SHAP nor LIME can offer on their own.

The fourth point is that these techniques can contradict each other. There is established literature showing that SHAP and LIME arrive at different attributions for the same prediction, and that it's possible to construct a discriminatory model that, in the eyes of these explanatory tools, appears harmless.

The only clue we have

Since the AI Regulation hasn't yet addressed this issue, we have to look where it has. And there we find the judgment of the Court of Justice of the European Union in the case of "Dun & Bradstreet Austria" (C-203/22), issued in February 2025, concerning the GDPR's right of access and "meaningful information about the logic applied."

The Court said two things that resonate strongly here. Meaningful information consists of describing, concisely, transparently, intelligibly, and in an easily accessible way, the procedures and principles actually applied, so that the person can understand the reason for the decision. And that delivering the algorithm or a complex technical description does not meet this requirement.

If we apply this standard to Article 86, the provisional conclusion becomes clear: a SHAP value diagram, delivered as is to the affected party, will hardly be "clear and meaningful".

The Gap

We have a regulation that requires clear and meaningful explanations. We have an industry that has decided, by consensus, that these explanations are produced using SHAP and LIME. And we have absolutely no authority that has stated whether this is valid.

This is the gap in which compliance departments, auditors, and lawyers will operate for the next few years. Reports will be signed that endorse a method whose legal sufficiency remains unresolved. And when the first serious litigation arises, the question will not be whether SHAP was used. It will be whether the person understood anything from that report.

What to Do in the Meantime

While this gap is being filled, prudence suggests a few things.

Use SHAP and LIME, yes, but as raw materials, not as a finished product. Translate your results into natural language, understandable to someone who has never seen a scatter plot. Accompany them with counterfactual explanations, which are what truly enable a person to act. Document separately and carefully the human role in the procedure. And above all, resist the temptation to confuse the tool with the obligation. Article 86 doesn't ask you to implement a method. It asks that a person understand a decision that has changed something in their life.

SHAP and LIME are extraordinary advances in engineering, and the law would be wrong to dismiss them. But it would also be wrong to accept them, without further ado, as the fulfillment of a fundamental right that no one has yet interpreted.

For some time, we will have to live with this discomfort: applying an article whose requirement we understand, with tools whose sufficiency we don't know. It's not a comfortable situation for a lawyer.

References

This article contains general legal information and does not constitute legal advice.

You might also like to read this

Legal AI

The Art of Overcoming Fear

The art of overcoming fear. Yet, on this journey of learning, there is one background noise we must silence. It is those three terrible words: 'You can't.'

Móvil mostrando apps de inteligencia artificial junto a un café
Legal AI

AI No Longer Just Answers: Now It Works

Until recently, legal AI did practically only one thing: answer questions and draft documents. To make the most of it, one had to become a specialist in