Causal Disentanglement via DenseNet Generative Models for Explainable AI in Chest X-Ray Diagnosis
Keywords:
Explainable Artificial Intelligence (XAI), Chest X-Ray Classification, Structurally Causal Conditional Variational Autoencoder (SC-CVAE), Counterfactual Explanations, Medical Image AnalysisAbstract
Deep learning models for medical image diagnosis often operate as black boxes, providing limited insight into the reasoning behind their predictions. This opacity hinders clinical trust and adoption, particularly in high-stakes applications such as chest X-ray interpretation. We propose a novel framework that integrates a structurally causal conditional variational autoencoder (SC-CVAE) with a DenseNet-121 backbone to transform opaque feature embeddings into interpretable, intervenable latent representations. The DenseNet-121 extracts hierarchical feature maps from chest X-rays, which are aggregated into a composite tensor and fed into the SC-CVAE. The SC-CVAE encodes these features into a set of latent factors that correspond to clinically meaningful concepts, such as pathology shape, intensity, anatomical location, and scanner artifacts. A domain-derived causal graph governs the relationships among these factors, and a mutual information regularizer enforces disentanglement by penalizing statistical dependence between causally unrelated variables. The decoder reconstructs the feature tensor from the latent factors, and an adversarial training mechanism ensures that concept edits remain localized and medically plausible. For explanation generation, we perform interventions on specific latent factors—for example, setting the shape factor to a counterfactual value—and decode the modified representation to produce a new feature tensor. The difference in downstream classification probabilities between the original and counterfactual features quantifies the concept-level contribution of each factor to the diagnostic decision. The entire system is trained end-to-end with a combined objective that includes reconstruction, mutual information, adversarial, and classification losses. Our approach yields, for each input X-ray, pathology probabilities, a Grad-CAM heatmap for spatial localization, and a concept-level contribution vector. This multi-faceted output enables clinicians to understand not only where the model focuses but also why specific visual features drive the diagnostic outcome. The proposed method thus advances explainable AI in medical imaging by bridging the gap between high-performance deep learning and clinically meaningful interpretability.
References
X Ouyang, S Karanam, Z Wu, T Chen, et al. Learning hierarchical attention for weakly-supervised chest x-ray abnormality localization and diagnosis. IEEE Transactions on Medical Imaging, 2020.
DM Pelt and JA Sethian. A mixed-scale dense convolutional neural network for image analysis. In Proceedings of the National Academy of Sciences, 2018.
RR Selvaraju, M Cogswell, A Das, et al. Grad-cam: Visual explanations from deep networks via gradient-based localization. In 2017 IEEE International Conference on Computer Vision (ICCV), 2017.
S Bach, A Binder, G Montavon, F Klauschen, KR Müller, et al. On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation. plosone, 10, e0130140.2015, 2015.
B Kim, M Wattenberg, J Gilmer, C Cai, et al. Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav). In International Conference on Machine Learning, 2018.
Y Chen, Z Ou, A Weller, et al. Neural mutual information estimation with vector copulas. In Advances in Neural Information Processing Systems, 2026.
J Pearl. Causal inference in statistics: An overview. Technical report, projecteuclid.org, 2009.
Y Chen, J Liu, L Peng, Y Wu, Y Xu, and Z Zhang. Auto-encoding variational bayes. Cambridge Explorations in Arts and Sciences, 2024.
N Kalchbrenner, L Espeholt, O Vinyals, et al. Conditional image generation with pixelcnn decoders. In Advances in Neural Information Processing Systems, 2016.
S Shakya, B Maharjan, and P Shakya. From entanglement to disentanglement: Comparing traditional vae and modified beta-vae performance. International Journal on Engineering Technology, 2024.
H Kim and A Mnih. Disentangling by factorising. In International Conference on Machine Learning, 2018.
M Scott and L Su-In. A unified approach to interpreting model predictions. In Advances in neural information processing systems, 2017, 2017.
M Ghassemi, L Oakden-Rayner, and AL Beam. The false hope of current approaches to explainable artificial intelligence in health care. The lancet digital health, 2021.
CE Sun, T Oikarinen, B Ustun, et al. Concept bottleneck large language models. In International Conference on Learning Representations, 2025.
B Schölkopf, F Locatello, S Bauer, NR Ke, et al. Toward causal representation learning. In Proceedings of the IEEE, 2021.
Y Wang and MI Jordan. Desiderata for representation learning: A causal perspective. Technical report, arXiv preprint arXiv:2109.03795, 2021.
A Almodóvar, A Javaloy, J Parras, et al. Decaflow: A deconfounding causal generative model. In Advances in Neural Information Processing Systems, 2026.
A Li, Y Pan, and E Bareinboim. Disentangled representation learning in non-markovian causal systems. In Advances in Neural Information Processing Systems 37, 2024.
G Qi and H Yu. Cmvae: causal meta vae for unsupervised meta-learning. In Proceedings of the AAAI Conference on Artificial Intelligence, 2023.
Y Pamungkas, MM Aung, MNA Uda, et al. Generative adversarial networks in medical imaging: A review of emerging trends and insights. Journal of Robotics and Control, 2026.
T Chakraborty, U Reddy KS, SM Naik, et al. Ten years of generative adversarial nets (gans): a survey of the state-of-the-art. Machine Learning: Science and Technology, 2024.
J Ho, A Jain, and P Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, 2020.
T Han, L Žigutytė, L Huck, MS Huppertz, et al. Reconstruction of patient-specific confounders in ai-based radiologic image interpretation using generative pretraining. Cell Reports Medicine, 2024.
AA Fernandes. Causally structured counterfactual generation for interpretable medical image analysis. Technical report, scholar.tecnico.ulisboa.pt, 2024.
M Aas-Alas, M Obrador-Reina, et al. Mimic-cxr-vqa: A large-scale llm-annotated dataset and comparative benchmark for medical visual question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2026.
Y Xu, MD Plumbley, and W Wang. Explainable ai in speaker recognition–attention map visualisation and evaluation. Technical report, arXiv preprint arXiv:2606.22901, 2026.
Z Dong, C Mundo-Levano, W Qian, D Lau, et al. A framework for directed acyclic hypergraph learning. Technical report, arXiv preprint arXiv:2606.21668, 2026.
K Lagemann, C Lagemann, B Taschler, et al. Deep learning of causal structures in high dimensions under data limitations. Nature Machine Intelligence, 2023.
L Demelius, R Kern, and A Trügler. Recent advances of differential privacy in centralized deep learning: A systematic survey. ACM Computing Surveys, 2025.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 The Author(s)

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.