Introduction

The first two articles in this series established what interpretability means and how current methods work. This article examines what those methods cannot deliver. Five papers, each from a different angle, document limitations that are not merely technical but conceptual, organisational, and structural.

Ghassemi et al. argue that current explainability methods cannot achieve their stated goals of trust, transparency, and bias mitigation. Watson identifies three conceptual problems embedded in the very design of IML methods. Bhatt et al. show that even when explanations are deployed, they serve engineers rather than affected stakeholders. Casper et al. demonstrate that black-box access, the most common audit mode, is insufficient for rigorous evaluation. Pawlicki et al. document a metrics crisis that makes it impossible to compare methods systematically.

The picture is not that XAI is useless. It is that XAI is much less reliable than its advocates claim, and that the field must confront its limitations honestly to make genuine progress.

This article is not legal advice.

Vocabulary for Limits and Critiques

Severe testing
A framework from the philosophy of science (Mayo, 1996) requiring that a hypothesis must be subjected to tests that it is likely to fail if false. Watson argues that IML methods do not subject their explanations to severe testing.
Product vs. process explanation
Watson distinguishes between explanations treated as a static deliverable computed once (a product) and explanations treated as an iterative exchange between agents engaged in causal inquiry (a process). He argues that current IML methods overwhelmingly treat explanations as products rather than processes.
Deployment gap
The mismatch between the research narrative that explanations serve affected stakeholders and the practice where explanations primarily serve internal ML engineers for debugging.
Outside-the-box access
Access to information about the development process of a model: training methodology, data provenance, documentation, deployment details, and internal evaluation results. Casper et al. argue this access level is necessary for rigorous auditing alongside white-box access.
Metric duplication
The phenomenon where multiple XAI evaluation metrics measure the same construct under different names, creating the illusion of comprehensive evaluation while actually measuring the same thing repeatedly.

Five Dimensions of Limitation

Ghassemi et al. (2021): The Clinical Critique

Ghassemi, Oakden-Rayner, and Beam published a pointed critique from within the medical AI community . Their core argument is that current explainability methods cannot deliver on the promises made on their behalf: engendering trust, providing transparency, and mitigating bias. They highlight a series of ways in which explanations can fail to support clinical practice.

Ambiguity. Different explanation methods applied to the same prediction regularly produce conflicting explanations. Without ground truth, there is no principled way to adjudicate between them. A clinician presented with one explanation from LIME and a different explanation from SHAP has no basis for choosing which to trust.

Inaccuracy. Explanations can be systematically misleading. Saliency-based methods can produce attribution maps that appear reassuring even for untrained networks, so misleading saliency is a genuine risk. Clever Hans phenomena (documented in Methods and Techniques for Explaining Machine Learning Models) show that models can be correct for the wrong reasons, and explanations faithfully reflecting model behaviour will therefore be wrong about reality.

Inscrutability. Even if explanations were accurate, clinicians lack training to critically evaluate them. The cognitive load of assessing the validity of an explanation while making a clinical decision creates a situation where explanations may decrease rather than improve decision quality.

The paper does not argue that explainability is impossible. It argues that the current enthusiasm for explainability as a solution to the trust problems of clinical AI is premature, and that rigorous internal and external validation of AI models is a more reliable path to safe deployment.

Watson (2022): The Conceptual Critique

Watson’s critique targets the conceptual foundations rather than empirical performance . The three challenges are worth restating because they are immune to technical fixes.

Ambiguity of target. When SHAP reports that feature x contributed 0.3 to a prediction, what exactly is being explained? The behaviour of the model (the function SHAP is approximating)? The data-generating process (the correlations in the training distribution)? The domain phenomenon (the real-world relationship the model is meant to capture)? Each is a legitimate target, but SHAP only addresses the first, while being interpreted as addressing the third.

Absent error quantification. LIME fits a local linear model without reporting confidence intervals. SHAP values are point estimates without standard errors. Neither method subjects its explanations to severe testing: constructing tests that the explanation is likely to fail if it is incorrect. In any scientific context, making claims without error quantification is unacceptable.

Product over process. Current methods overwhelmingly treat explanations as static deliverables, computed once and for all. Watson argues that successful explanations are more of a process than a product: an iterative exchange between at least two agents engaged in a certain sort of causal inquiry, requiring dynamic refinements and user interaction. A list of feature attributions, computed once and handed over, does not support this process.

Watson’s prescription is not to abandon IML but to adopt the methodological standards of the sciences it seeks to support: clarifying the target of explanation, quantifying error rates to allow severe testing, and treating explanation as an interactive process rather than a static deliverable.

Bhatt et al. (2020): The Deployment Gap

Bhatt et al. conducted the first empirical study of how organisations actually use explainability in production . The results reveal a fundamental mismatch between research narratives and practice.

The study, spanning roughly thirty people across approximately twenty organisations, found that the majority of explainability deployments serve ML engineers debugging models, not end users affected by model decisions. Explanations were used for: model debugging and feature engineering (most common), regulatory compliance documentation, and building internal stakeholder trust. End-user-facing explanations were rare.

The gap matters because the research community motivates explainability as a tool for transparency and accountability to affected stakeholders. If explanations never reach those stakeholders, the normative goals of XAI remain unfulfilled regardless of technical quality.

The paper proposes a stakeholder-centred framework: establish clear goals for what the explanation should achieve for a specific stakeholder, design for that stakeholder’s cognitive and domain context, and evaluate whether the goals are met. This framework, while sensible, is rarely applied in practice.

Casper et al. (2024): Black-Box Audit Limitations

Casper et al. examine a structural constraint on AI auditing . External auditors typically receive only black-box access (query inputs, observe outputs). The paper demonstrates that this access level is insufficient for rigorous evaluation across multiple dimensions.

Black-box access cannot reliably detect model editing or fine-tuning. Adversarial robustness evaluations are weaker without gradient access. Mechanistic interpretability, understanding how circuits within the model implement behaviours, requires white-box access to weights and activations. Data contamination detection is limited when the training data cannot be inspected.

The access framework of the paper distinguishes black-, grey-, and white-box access, introduces two further categories (de facto white-box and outside-the-box), and maps specific audit tasks to the access levels they require. The central recommendation: auditors must disclose their access level for audit results to be interpretable, and audits of higher-risk systems should be conducted with higher levels of access where this is feasible.

This critique has direct implications for explainability. Many explanation methods require white-box access (gradients, activations). When auditors only have black-box access, the set of available explanation methods is restricted, and the conclusions that can be drawn from explanations are correspondingly limited.

Pawlicki et al. (2024): The Metrics Crisis

Pawlicki et al. document a problem that makes XAI evaluation unreliable even when the methods work correctly . The proliferation of evaluation metrics has created three pathologies:

Metric duplication. Nearly ninety distinct XAI evaluation metrics exist, many measuring the same construct under different names. Faithfulness metrics, for example, appear in multiple variants with different mathematical formulations but overlapping conceptual targets.

Metric inefficacy. Some metrics fail to distinguish between high- and low-quality explanations. When all methods receive similar scores, the metric provides no information for method selection.

Metric confusion. Different metrics applied to the same explanation method produce diverging assessments. One metric may rank SHAP highest while another ranks LIME highest, leaving the practitioner with no basis for choice.

The paper argues for context-specific metric selection, ensuring that practitioners choose the most appropriate metrics for their specific needs, while also calling for a more unified evaluation framework. There is no single metric set that suits every application, and metric results are often inconsistent across frameworks even when the metrics seem theoretically sound.


Synthesis: The Limits Are Not Only Technical

Across these five papers, a pattern emerges that is more troubling than any single critique. The limitations of current XAI are not only technical (better methods will fix them) but structural (the deployment context constrains what can be achieved) and conceptual (the methods are not designed for the questions they are asked to answer).

Ghassemi et al. and Watson together establish that the conceptual foundations are incomplete. Bhatt et al. establish that even good methods do not reach the stakeholders who need them. Casper et al. show that access constraints limit what methods can be applied. Pawlicki et al. demonstrate that the evaluation infrastructure cannot distinguish good methods from poor ones.

The constructive response is not to abandon XAI but to pursue it with greater honesty about its limitations; to match method choice to the specific question and stakeholder; and to invest in the evaluation infrastructure that the field currently lacks.


Conclusion

The critical perspectives examined in this article do not refute the value of XAI. They demand that the field hold itself to higher standards: clear explanatory targets, error quantification, stakeholder-centred deployment, adequate model access, and reliable evaluation metrics. The next article examines how the field is responding to these demands through domain applications, evaluation frameworks, and governance structures.


Frequently Asked Questions

What are the main limitations of current XAI methods identified in the literature?

The literature identifies five interconnected limitations: explanation ambiguity across methods, systematic inaccuracy in explanation methods, a gap between research and deployment practice, structural limits on black-box auditability, and a metrics crisis where evaluation tools fail to distinguish meaningful differences between methods. These limitations are not merely technical but also conceptual and organisational.

Why do Ghassemi et al. argue that explainability requirements may harm clinical AI?

Ghassemi et al. argue that mandating explainability for clinical AI could divert resources from rigorous validation through prospective trials and deployment monitoring. They contend that a thoroughly validated model whose outputs clinicians understand through training may be safer than an explainable model whose behaviour has been less rigorously tested. This argument challenges the assumption that explainability is always a net benefit.

What is the deployment gap that Bhatt et al. identify in XAI practice?

Bhatt et al. interviewed practitioners across organisations that do deploy explainability and found that the majority of deployments primarily serve internal ML engineers debugging models, rather than affected stakeholders. Their conclusion lists specific limitations of current explanations for end-user use, including the need for domain experts to evaluate them, the risk of spurious correlations, and the latency of computing explanations in real time. This gap between research narratives and deployment practice limits the practical impact of XAI.

How does Casper et al. characterise the limits of black-box audits?

Casper et al. demonstrate that black-box access alone is insufficient for reliable auditing because adversarial inputs can be designed to produce misleading explanations, model editing can change behaviour without detection from explanation outputs, and the inherent ambiguity of post-hoc explanations means that multiple contradictory narratives can be supported by the same audit data.

What does Pawlicki et al. identify as the metrics crisis in XAI evaluation?

Pawlicki et al. identify three problems: metric duplication where multiple metrics measure the same underlying property and create false diversity, metric inefficacy where some metrics fail to distinguish between random and meaningful explanations, and metric confusion where contradictory results across metrics make it impossible to determine which method is better for a given task.

Appendix: Source Material

Author and Source Credibility

Source Author profile Venue Impact
Ghassemi et al. (2021) MIT/Harvard Medical The Lancet Digital Health 500+ citations
Watson (2022) University College London Synthese (philosophy journal) Growing citations
Bhatt et al. (2020) CMU/Cambridge/PAI FAT* (now FAccT) 600+ citations
Casper et al. (2024) MIT/Harvard/ multi-institution FAccT 2024 Rapidly growing
Pawlicki et al. (2024) Multi-institution European Neurocomputing New (2024)

Citability Snapshot

Claim category Count Examples
Verified (empirical) 4 Explanation ambiguity, deployment gap, black-box audit limits, metric duplication
Verified (conceptual analysis) 3 Ambiguity of target, absent error quantification, product vs. process
Inferred 3 Trust through validation vs. explanation; end-user inscrutability; structural limits persistent
Speculative 1 Rigorous validation sufficient without explanations (debated)