Back to Blogs
What Is Explainable AI? How Explainability Works When the Model Is an LLM
Explainable AI is the practice of showing a person why an AI system produced a specific output, with evidence that person can understand and check. It is often shortened to XAI, for explainable artificial intelligence. The definition is simple when a model scores a loan application. It gets hard when the model writes a paragraph. So what is explainable AI when the system is a large language model or an agent? That is the question this guide answers.
Key insights
- Write every explanation for a specific reader. If that reader cannot follow it, it does not count.
- NIST separates three ideas: transparency answers what happened, explainability answers how, and interpretability answers why and what it means.
- SHAP and LIME explain a prediction in terms of its input features. An LLM reads free text, so it needs other kinds of evidence as well.
- Evidence for an LLM output comes from the input, the retrieved sources, the training data or the model’s internals, and each shows something different.
- A model’s written reasoning is not proof of why it answered.
What is explainable AI?
Start with the plain version. An AI system produces a score, a recommendation or a piece of text. Someone needs to know why. Explainable AI is the work of producing that why, as evidence the person can inspect, for the output in front of them.
The word explain carries more weight than it looks. An explanation has a reader, and it only counts if that reader can follow it. NIST builds this into its four principles of explainable AI: under the Meaningful principle, an explanation has to be understandable to the people it is meant for.
It also has to be true. NIST’s Explanation Accuracy principle says an explanation must correctly reflect the reason the system produced its output. Large language models make that test hard, and most of this article is about what you can do anyway.
DARPA launched its Explainable AI (XAI) program in 2017 to produce more explainable models while keeping prediction accuracy high, and to help people understand, appropriately trust and manage AI partners. In January 2023 the NIST AI Risk Management Framework named explainable and interpretable as one of the characteristics of trustworthy AI.
Explainability vs interpretability vs transparency
It is easy to treat the three words as synonyms, and mixing them up produces requirements nobody can meet. The clearest split comes from the NIST AI Risk Management Framework.
In NIST’s framing, transparency answers what happened in a system. Explainability answers how a decision was made. Interpretability answers why it was made, and what the output means to the person using it.
| Term | Question it answers (NIST framing) | What it looks like for an LLM output | Who usually needs it |
| Transparency | What happened? | Which model version, instructions, documents and tools were involved | Auditors, regulators |
| Explainability | How was the output produced? | Which inputs, sources or training examples shaped the answer | Developers, validators |
| Interpretability | Why, and what does it mean here? | What the answer means for the decision in front of the user | Business owners, affected people |
One caution before you quote this table. In machine learning, interpretability often means something narrower: a model simple enough to read directly, like a short decision tree. When you write a policy or an RFP, say which meaning you intend.
In practice, most AI explainability work means producing the middle row, then translating it into the third for whoever has to act on it.
How explainable AI works on traditional models
The classic XAI methods explain a prediction in terms of input features, such as income, account age or pixels. Two distinctions organize the toolkit. An intrinsically interpretable model, like a linear model or a small tree, can be read as built, while a post-hoc method explains a model after training. A global explanation describes overall behavior, and a local one covers a single prediction.
The best-known methods are local and post-hoc:
- SHAP (Lundberg and Lee, 2017) assigns each feature an importance value for a particular prediction, based on Shapley values from cooperative game theory. Paper.
- LIME (Ribeiro, Singh and Guestrin, 2016) explains one prediction of any classifier by fitting a simple, interpretable model around it. Paper.
- Integrated Gradients (Sundararajan, Taly and Yan, ICML 2017) attributes a deep network’s prediction to its input features using only standard gradient calls.
- Saliency maps highlight which parts of an image most affect a class score, an approach Simonyan and colleagues described in 2013.
- Counterfactual explanations (Wachter, Mittelstadt and Russell, 2017) describe the smallest change that would lead to a different outcome, without explaining the system’s internal logic.
Picture a credit model that declines an application. SHAP ranks the inputs that pushed the score down, for the analyst. A counterfactual tells the applicant what would have had to be different. Same decision, two readers, two methods.
Why are LLMs harder to explain?
Ask what is explainable AI about a credit score and the answer above works. Ask it about a chatbot answer and it gets much harder, because the methods above describe a prediction in terms of input features, and a chatbot has no fixed feature list.
An LLM’s input is free text, so there is no tidy feature list to rank. Its output is free text too, and one paragraph can blend a retrieved fact, a pattern learned in training and an instruction from the system prompt. What the model knows is spread across its parameters, with no row to point at. Put it inside an agent that calls tools in a loop, and one answer becomes a chain of decisions.
Attention weights are the tempting shortcut. They look like a map of what the model focused on. Jain and Wallace found in 2019 that attention weights largely do not provide meaningful explanations, in experiments on earlier NLP models. Wiegreffe and Pinter replied that it depends on how you define explanation, and proposed tests for when attention can be used. Read attention as a partial, contested signal. On its own, it explains nothing.
Four places an explanation of an LLM output can come from
If you cannot read an LLM’s answer off a feature list, where does the evidence come from? There are four places to look. This is the core of explainable AI for LLMs, and none of the four proves everything.
Figure: the four sources of an LLM explanation
| Source of evidence | What it answers | What it cannot prove |
| Input attribution | Which parts of the question, instructions or context changed the answer? | That the model’s internal process matches the attribution |
| Retrieved sources | Which documents support the answer, and where are they? | That the model actually relied on the cited passage |
| Training data | Which training examples shaped this behavior? | Exact cause and effect, especially for very large models |
| Model internals | Which internal features and circuits were active? | A complete account of how the model reached the answer |
1. Input attribution
Input attribution asks which parts of the input changed the output. For an LLM, the input is more than the user’s question: it includes instructions, retrieved passages, tool results and earlier turns. A practical approach is to remove one piece and run it again. If deleting a passage changes a sentence, that passage mattered to that sentence. Simple enough. The limit is that attribution shows what the output was sensitive to, and it says nothing about the path the model took inside.
2. Retrieved sources
Retrieval-augmented generation, or RAG, gives a model documents to work from, and those documents can be shown as citations. The 2020 RAG paper by Lewis and colleagues named providing provenance for a model’s decisions as an open research problem that retrieval helps address.
For many business users, citations are the explanation they will actually look at, which makes them easy to overtrust. A citation shows where supporting text sits. It does not show that the model used that text faithfully. The stronger check tests influence: take a source away and see whether the statement changes.
3. Training data
Some behavior comes from what the model learned. Training data attribution tries to trace an output back to the training examples that shaped it.
The classic method is influence functions. Koh and Liang used them to trace a prediction back to the training points most responsible for it, in a paper that won the ICML 2017 best paper award. In 2023, Anthropic researchers scaled influence functions to language models with up to 52 billion parameters, using an approximation called EK-FAC. They note the method is hard to scale because it needs an inverse-Hessian-vector product. They also found that influences decay to near zero when the order of key phrases is flipped.
Tracing an output to training examples requires access to those examples. That is a real constraint with foundation models. Stanford’s December 2025 Foundation Model Transparency Index gave major developers a mean score of 41 out of 100 and noted that training data continues to be opaque.
4. Model internals
Mechanistic interpretability tries to read the model itself: which internal features activate, and how they combine into an answer.
In May 2024, Anthropic researchers used sparse autoencoders to pull interpretable features out of Claude 3 Sonnet, a production model, in their Scaling Monosemanticity paper. Some features related to security vulnerabilities, bias and deception. In March 2025, Anthropic published circuit tracing research that uses attribution graphs to partially trace the intermediate steps Claude 3.5 Haiku takes from a prompt to a response.
Anthropic is candid about the limits. The graphs gave satisfying insight for about a quarter of the prompts tried, and they study a replacement model that incompletely and imperfectly captures the original. This is research that is advancing. It is not yet an audit tool you run on every answer.
What about agents? An agent adds a record rather than a fifth source: the trace. Log every model call, tool call, input, output and timestamp, and a reviewer can reconstruct what the agent did, in order. A trace answers the transparency question. Pair it with one of the four sources and it starts to explain.
Is chain-of-thought reasoning an explanation?
Reasoning models write out steps before they answer, and it is natural to read that text as the explanation. The issue is faithfulness: does the written reasoning reflect what actually drove the answer? That is NIST’s Explanation Accuracy principle, applied to one response.
Turpin and colleagues tested this in 2023. They nudged models with biasing features, such as making (A) the answer in every example. Accuracy dropped by up to 36% across 13 BIG-Bench Hard tasks, and the explanations failed to mention the bias. Their conclusion: chain-of-thought explanations can systematically misrepresent the true reason for a model’s prediction.
Anthropic’s April 2025 study inserted hints about the answer into multiple-choice questions and checked whether reasoning models admitted using them. Claude 3.7 Sonnet mentioned a hint it had used 25% of the time on average, and DeepSeek R1 39%. When models were trained to exploit a reward hack, they used it in over 99% of cases but mentioned it in their reasoning less than 2% of the time in most scenarios. Anthropic calls the setup somewhat contrived, with multiple-choice quizzes, hints inserted on purpose and only Anthropic and DeepSeek models. Both studies still point the same way.
So keep the reasoning text, because it helps with debugging. Just do not make it the explanation of record. Check what it claims against evidence outside it: the sources, the tool results, the trace.
Who needs an explanation, and what kind?
NIST recommends tailoring descriptions of how an AI system works to the user’s role, knowledge and skill level.
| Who asks | What they want to know | The explanation that answers it |
| Developer or data scientist | Why did the model do that, and how do I fix it? | Input attribution, training data attribution, run traces |
| Model validator or risk reviewer | Does the explanation reflect what the system really did? | Evidence tested against the actual process, including faithfulness checks |
| Auditor or regulator | Can I reconstruct what happened? | Documentation, logs and data provenance |
| Deployer or business owner | Can I interpret this output and use it appropriately? | Sources, confidence information and known limits |
| Affected person | Why was I declined, and what would change it? | The main reasons in plain language, and what could change the result |
A feature attribution chart helps the developer and tells a declined applicant nothing. A plain-language reason helps the applicant and gives the validator nothing to test. Plan for at least two layers.
NIST’s four principles of explainable AI
NIST IR 8312, published in September 2021, sets out four principles. Use the exact names when you cite them.
- Explanation. A system delivers or contains evidence or reasons for its outputs or processes. For an LLM, the answer arrives with something a person can inspect.
- Meaningful. Explanations are understandable to the people they are meant for, so one answer may need several explanations.
- Explanation Accuracy. The explanation correctly reflects the reason the system produced its output. The chain-of-thought problem lives here.
- Knowledge Limits. The system only operates under the conditions it was designed for and when it reaches sufficient confidence. For an LLM, one way to apply it is to flag low confidence instead of producing a fluent guess.
Explanation Accuracy is the hardest of the four for generative models, and the one to press when a vendor shows you an explanation.
Does the law require explainable AI?
In specific situations, yes, and each law words it differently. This is orientation, not legal advice, as of September 2026. Dates in this area move, so confirm with counsel.
The EU AI Act
The EU AI Act, Regulation (EU) 2024/1689, entered into force on 1 August 2024 and applies in stages. Then the dates moved. The AI Omnibus, Regulation (EU) 2026/1744, entered into force on 27 July 2026 and pushed back the high-risk timeline, so rules for Annex III systems, which include credit scoring and hiring tools, now apply from 2 December 2027. Annex I, for high-risk AI in regulated products such as machinery and lifts, follows on 2 August 2028.
Two articles matter most. Article 13 requires high-risk systems to be transparent enough for deployers to interpret the output and use it appropriately. Under Article 86, a person affected by a decision based on an Annex III high-risk system, with legal or similarly significant adverse effects, can obtain clear and meaningful explanations of the AI system’s role and the main elements of the decision. Both belong to the high-risk rules, which are not yet in force. Article 86 is tied to Annex III systems, whose rules apply from 2 December 2027. The transparency duty in Article 13 follows the high-risk timeline for each system: 2 December 2027 for Annex III and 2 August 2028 for Annex I.
Article 50 works differently. Since 2 August 2026 it has required, among other things, that people are told when they are interacting with an AI system, unless that is obvious. That is disclosure. It does not require an explanation.
The GDPR
Article 22 gives people the right not to be subject to a decision based solely on automated processing, including profiling, with legal or similarly significant effects. The words “to obtain an explanation of the decision reached” sit in Recital 71, outside Article 22’s operative text. Articles 13, 14 and 15 require meaningful information about the logic involved.
In February 2025 the EU Court of Justice ruled in Case C-203/22, often cited as Dun and Bradstreet Austria, that a person can require the controller to explain the procedure and principles actually applied to reach a result such as a credit profile. Handing over a complex formula, or describing every step, does not meet that requirement, because neither is concise and intelligible.
The United States
In credit, Regulation B says that when a lender turns an applicant down, its reasons must be specific and must name the principal ones. Telling an applicant they failed to achieve a qualifying score is not enough. The CFPB withdrew Circulars 2022-03 and 2023-03, its guidance on adverse action notices for complex algorithms, on 12 May 2025, and said the withdrawal is not necessarily final. The Regulation B requirement itself still stands.
For banks, SR 26-2, issued on 17 April 2026 by the Federal Reserve, the OCC and the FDIC, replaced SR 11-7 and SR 21-8 as model risk management guidance. It is guidance, not an enforceable standard, and it states that generative AI and agentic AI models are not within its scope. Sit with that. The models hardest to explain are the ones US model risk guidance leaves out, so a bank’s own governance has to cover them.
Explainable AI examples from the LLM era
The three scenarios below are hypothetical. Each shows the explanation a reviewer would actually receive.
A lender drafts adverse action letters with an LLM. Under Regulation B, the lender has to give the applicant specific reasons and name the principal ones. A careful design keeps those reasons in the underwriting record and uses the LLM only to phrase them. Then a reviewer can match every reason in the letter to that record, instead of trusting reasons that merely sound plausible.
A compliance assistant answers a policy question. An analyst asks whether a vendor contract needs an extra review. The answer cites two policy passages with file and section, plus a confidence indicator. A stronger setup also shows which passage changed the answer when removed, and says so when confidence is low. That last part is the Knowledge Limits principle at work.
An agent drafts an audit finding. The agent reads transaction records, calls a rules lookup tool and writes a finding. The trace lists each model call and tool result in order. The auditor checks each claim in the agent’s reasoning against those tool results before accepting anything.
The limits of explainability
Explanations fail in predictable ways.
- A persuasive explanation can be wrong. The chain-of-thought studies found fluent reasoning that left out what drove the answer.
- The best tools approximate. Circuit tracing studies a replacement model and gave satisfying insight on about a quarter of the prompts tried.
- Attribution can be brittle. Influences decayed to near zero when key phrases swapped order.
- More detail can explain less. The EU Court of Justice held that a formula or a step-by-step description is not an explanation for the person affected.
- A citation shows availability. Whether the model relied on the passage is a separate question.
An explanation is evidence to weigh. It never guarantees the model is right.
Where to go from here
That is the working answer to what is explainable AI once the model is an LLM: evidence from several sources, matched to the person asking. If you are evaluating platforms, our enterprise guide to explainable AI covers requirements and evaluation. Explainability also starts with the data you label, covered in Explainability Starts at the Label.
In SeekrFlow, explainability draws on several of the sources above. For fine-tuned models created in SeekrFlow and trained after 22 September 2025, it surfaces the question and answer pairs from the fine-tuning dataset that most influenced each response. An internal evaluator LLM rates each pair High, Medium or Low impact, and you can click through to the original document chunk. Context attribution tests how an agent’s response changes when sources are removed, to show which sources influenced each statement. Each agent run is captured as a full trace, and every response comes with confidence scores. Seekr describes this explainability as model-agnostic.
See how explainability works on your own models
Book a consultation with an AI expert. Share your challenges and objectives, and our team will connect to explore solutions and walk you through a live demo of SeekrFlow.
Sources
- NIST IR 8312, Four Principles of Explainable Artificial Intelligence, September 2021
- NIST AI 100-1
- DARPA, Explainable Artificial Intelligence (XAI) program page, and Gunning and Aha in AI Magazine (2019)
- Lundberg and Lee (2017)
- Ribeiro, Singh and Guestrin, “Why Should I Trust You?”: Explaining the Predictions of Any Classifier, KDD 2016
- Sundararajan, Taly and Yan (ICML 2017)
- Simonyan et al., Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps (2013)
- Wachter, Mittelstadt and Russell (2017)
- Jain and Wallace (NAACL 2019), with the reply by Wiegreffe and Pinter (EMNLP 2019)
- Lewis et al. (2020)
- Koh and Liang (ICML 2017), and Grosse et al. at Anthropic (2023)
- Stanford CRFM (FMTI, December 2025)
- Anthropic, Scaling Monosemanticity (2024), On the Biology of a Large Language Model (2025) and Reasoning Models Don’t Always Say What They Think (2025)
- Turpin et al. (NeurIPS 2023)
- European Commission, AI Act implementation timeline, AI Omnibus announcement, and AI Act Service Desk pages for Article 13, Article 50 and Article 86
- GDPR
- Court of Justice of the EU, Case C-203/22, judgment of 27 February 2025
- 12 CFR 1002.9
- Federal Register, withdrawal notice of 12 May 2025 (90 FR 20084)
- Federal Reserve, SR 26-2 and its attachment, 17 April 2026
FAQ
What is explainable AI in simple terms?
Explainable AI means an AI system can show why it produced a specific output, in a way the person relying on it can understand and check. For a credit model, that might be the main reasons for a decline. For an LLM, it might be the sources behind an answer and which of them changed it.
What is XAI?
XAI is the usual abbreviation for explainable AI. DARPA used it for its Explainable AI (XAI) program, which aimed to produce more explainable models while keeping prediction accuracy high.
How is explainability different from interpretability and transparency?
In NIST’s framing, transparency answers what happened, explainability answers how a decision was made, and interpretability answers why and what it means to the user.
Can you explain the output of a large language model?
Partly. You can show which inputs changed the answer, which retrieved sources support it, which training examples shaped the behavior and, in research settings, which internal features were active. None gives a complete account, and the model’s own written reasoning is not a reliable substitute.
Why do regulators ask for explainable AI?
Different rules protect different people. Regulation B requires specific principal reasons when credit is denied. The GDPR requires meaningful information about the logic involved in solely automated decisions with significant effects. The EU AI Act requires high-risk systems to be transparent enough for deployers to interpret their output, under rules that apply to Annex III systems from December 2027.
What are examples of explainable AI in enterprise use?
Typical patterns are specific reasons on credit decisions, citations with source locations on answers from internal documents, traces of every step an agent took, and attribution from a fine-tuned model’s output back to the training examples that shaped it.
Accelerate your path to AI impact
Book a consultation with an AI expert. We’re here to help you speed up your time to AI ROI.
Request a demo