Advanced International Journal for Research

E-ISSN: 3048-7641   •   Impact Factor: 9.11

A Widely Indexed Open Access Peer Reviewed Multidisciplinary Bi-monthly Scholarly International Journal

Call for Paper Volume 7, Issue 5 (September-October 2026) Submit your research before last 3 days of October to publish your research paper in the issue of September-October.

Beyond Accuracy: Evaluating the Faithfulness and Trustworthiness of Explainable AI in Malware Detection

Author(s) Ms Anushka Chaudhary
Country India
Abstract Machine learning (ML) and deep learning (DL) based malware detectors now routinely report accuracy, F1, and AUC scores at or near ceiling, and this performance has become the field's dominant, often sole, measure of success. However, high detection accuracy is not evidence that a model's internal reasoning corresponds to genuine malicious behavior rather than dataset artifacts or spurious correlations. Explainable AI (XAI) techniques such as SHAP, LIME, and attention-based attribution have been proposed to open this black box, but the overwhelming majority of malware-detection studies that use XAI evaluate explanations qualitatively, or not at all, rather than testing whether the explanations are faithful, stable, and behaviorally grounded. This paper proposes and instantiates a four-dimensional evaluation framework for XAI in malware detection, covering (i) detection performance as a baseline, (ii) faithfulness, measured via feature-deletion/insertion perturbation tests, (iii) consistency of explanations across near-duplicate and same-family malware samples, and (iv) ground-truth alignment between top explanatory features and documented malicious behaviors drawn from MITRE ATT&CK techniques and VirusTotal behavioral tags. We further quantify the accuracy–explainability trade-off across models of varying inherent interpretability. Using an EMBER v2 derived corpus and a LightGBM classifier achieving 97.8% accuracy and an F1 score of 0.96, we find that only 66% of top-ranked explanatory features pass faithfulness testing, explanation consistency within malware families is moderate (intra-family ρ = 0.61 vs. between-family ρ = 0.29), and only 42% of explained features align with documented ATT&CK techniques rather than dataset or packer artifacts. These results support the central claim of this work: detection accuracy and explanation trustworthiness are empirically decoupled, and accuracy alone is an insufficient basis for trusting or deploying AI-driven malware detectors in security-critical, analyst-facing settings.
Keywords Malware Detection
Field Computer > Artificial Intelligence / Simulation / Virtual Reality
Published In Volume 7, Issue 5, September-October 2026
Published On 2026-10-06

Share this