Explainable Prediction of Open-Source Repository Maintenance Risk from Activity Signals
Keywords:
Explainable Artificial Intelligence, Maintenance-Risk Prediction, Mining Software Repositories, Open-Source Software, Repository Maintenance, Software SustainabilityAbstract
Open-source software sustainability depends on repositories continuing to accept changes, resolve maintenance demand, preserve contributor knowledge, deliver releases, and update technical dependencies. Distinguishing genuine maintenance risk from an intentional pause, a mature low-change state, or completed software is nevertheless difficult. A single observable, such as time since the last commit, cannot represent this distinction because maintenance is socio-technical, temporally uneven, and conditioned by project context. This paper proposes a conceptual framework for forward-looking repository maintenance-risk prediction using longitudinal signals from commits, contributors, issues, pull requests, releases, dependencies, and community interactions. Maintenance risk is defined as a latent probability that a repository will become unable to perform expected maintenance functions during a specified future horizon, rather than as an unconditional abandonment label. The framework constructs windowed levels, trends, volatility, latency, concentration, and cross-signal interactions; adjusts interpretation for repository context; compares interpretable, nonlinear, survival, and temporal models; and places global, repository-specific, temporal, counterfactual, and uncertainty explanations between prediction and stakeholder action. Ten testable propositions and a time-aware validation protocol are developed without claiming empirical performance. The intended contribution is a falsifiable architecture that can support maintainers, adopters, ecosystem managers, and funders while limiting stigmatizing or overconfident use. Longitudinal, cross-ecosystem validation and human evaluation of explanation usefulness remain necessary.
References
K. Manikas and K. M. Hansen, “Software ecosystems—A systematic literature review,” J. Syst. Softw., vol. 86, no. 5, pp. 1294–1306, 2013, doi: 10.1016/j.jss.2012.12.026.
O. Franco-Bedoya, D. Ameller, D. Costal, and X. Franch, “Open source software ecosystems: A systematic mapping,” Inf. Softw. Technol., vol. 91, pp. 160–185, 2017, doi: 10.1016/j.infsof.2017.07.007.
S. Jansen, “Measuring the health of open source software ecosystems: Beyond the scope of project health,” Inf. Softw. Technol., vol. 56, no. 11, pp. 1508–1519, 2014, doi: 10.1016/j.infsof.2014.04.006.
M. Oriol, C. Müller, J. Marco, P. Fernandez, X. Franch, and A. Ruiz-Cortés, “Comprehensive assessment of open source software ecosystem health,” Internet Things, vol. 22, Art. no. 100808, 2023, doi: 10.1016/j.iot.2023.100808.
J. Coelho and M. T. Valente, “Why modern open source projects fail,” in Proc. 11th Joint Meeting Eur. Softw. Eng. Conf. ACM SIGSOFT Symp. Found. Softw. Eng., 2017, pp. 186–196, doi: 10.1145/3106237.3106246.
J. Coelho, M. T. Valente, L. Milen, and L. L. Silva, “Is this GitHub project maintained? Measuring the level of maintenance activity of open-source projects,” Inf. Softw. Technol., vol. 122, Art. no. 106274, 2020, doi: 10.1016/j.infsof.2020.106274.
M. Valiev, B. Vasilescu, and J. Herbsleb, “Ecosystem-level determinants of sustained activity in open-source projects: A case study of the PyPI ecosystem,” in Proc. 26th ACM Joint Meeting Eur. Softw. Eng. Conf. Symp. Found. Softw. Eng., 2018, pp. 644–655, doi: 10.1145/3236024.3236062.
A. Ait, J. L. Cánovas Izquierdo, and J. Cabot, “An empirical study on the survival rate of GitHub projects,” in Proc. 19th Int. Conf. Mining Softw. Repositories, 2022, pp. 365–375, doi: 10.1145/3524842.3527941.
G. Avelino, L. Passos, A. Hora, and M. T. Valente, “A novel approach for estimating truck factors,” in Proc. IEEE 24th Int. Conf. Program Comprehension, 2016, pp. 1–10, doi: 10.1109/ICPC.2016.7503718.
M. Foucault, M. Palyart, X. Blanc, G.C. Murphy, and J.-R. Falleri, “Impact of developer turnover on quality in open-source software,” in Proc. 10th Joint Meeting Found. Softw. Eng., 2015, pp. 829–841, doi: 10.1145/2786805.2786870.
L. Bao, X. Xia, D. Lo, and G. C. Murphy, “A large scale study of long-time contributor prediction for GitHub projects,” IEEE Trans. Softw. Eng., vol. 47, no. 6, pp. 1277–1298, 2021, doi: 10.1109/TSE.2019.2918536.
H. S. Qiu, A. Nolte, A. Brown, A. Serebrenik, and B. Vasilescu, “Going farther together: The impact of social capital on sustained participation in open source,” in Proc. IEEE/ACM 41st Int. Conf. Softw. Eng., 2019, pp. 688–699, doi: 10.1109/ICSE.2019.00078.
I. Steinmacher, M. A. Graciotto Silva, M. A. Gerosa, and D. F. Redmiles, “A systematic literature review on the barriers faced by newcomers to open source software projects,” Inf. Softw. Technol., vol. 59, pp. 67–85, 2015, doi: 10.1016/j.infsof.2014.11.001.
G. Gousios, M. Pinzger, and A. van Deursen, “An exploratory study of the pull-based software development model,” in Proc. 36th Int. Conf. Softw. Eng., 2014, pp. 345–355, doi: 10.1145/2568225.2568260.
J. Tsay, L. Dabbish, and J. Herbsleb, “Influence of social and technical factors for evaluating contribution in GitHub,” in Proc. 36th Int. Conf. Softw. Eng., 2014, pp. 356–366, doi: 10.1145/2568225.2568315.
Y. Yu, H. Wang, V. Filkov, P. Devanbu, and B. Vasilescu, “Wait for it: Determinants of pull request evaluation latency on GitHub,” in Proc. IEEE/ACM 12th Work. Conf. Mining Softw. Repositories, 2015, pp. 367–371, doi: 10.1109/MSR.2015.42.
R. G. Kula, D. M. German, A. Ouni, T. Ishio, and K. Inoue, “Do developers update their library dependencies?” Empir. Softw. Eng., vol. 23, no. 1, pp. 384–417, 2018, doi: 10.1007/s10664-017-9521-5.
A. Decan, T. Mens, and P. Grosjean, “An empirical comparison of dependency network evolution in seven software packaging ecosystems,” Empir. Softw. Eng., vol. 24, no. 1, pp. 381–416, 2019, doi: 10.1007/s10664-017-9589-y.
A. Decan, T. Mens, and E. Constantinou, “On the impact of security vulnerabilities in the npm package dependency network,” in Proc. 15th Int. Conf. Mining Softw. Repositories, 2018, pp. 181–191, doi: 10.1145/3196398.3196401.
H. Borges and M. T. Valente, “What’s in a GitHub star? Understanding repository starring practices in a social coding platform,” J. Syst. Softw., vol. 146, pp. 112–129, 2018, doi: 10.1016/j.jss.2018.09.016.
N. Munaiah, S. Kroh, C. Cabrey, and M. Nagappan, “Curating GitHub for engineered software projects,” Empir. Softw. Eng., vol. 22, no. 6, pp. 3219–3253, 2017, doi: 10.1007/s10664-017-9512-6.
E. Kalliamvakou, G. Gousios, K. Blincoe, L. Singer, D. M. German, and D. Damian, “The promises and perils of mining GitHub,” in Proc. 11th Work. Conf. Mining Softw. Repositories, 2014, pp. 92–101, doi: 10.1145/2597073.2597074.
M. Golzadeh, A. Decan, D. Legay, and T. Mens, “A ground-truth dataset and classification model for detecting bots in GitHub issue and PR comments,” J. Syst. Softw., vol. 175, Art. no. 110911, 2021, doi: 10.1016/j.jss.2021.110911.
A. Barredo Arrieta et al., “Explainable artificial intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI,” Inf. Fusion, vol. 58, pp. 82–115, 2020, doi: 10.1016/j.inffus.2019.12.012.
M. T. Ribeiro, S. Singh, and C. Guestrin, “Why should I trust you?: Explaining the predictions of any classifier,” in Proc. 22nd ACM SIGKDD Int. Conf. Knowl. Discovery Data Mining, 2016, pp. 1135–1144, doi: 10.1145/2939672.2939778.
S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Advances in Neural Information Processing Systems, vol. 30, 2017, pp. 4765–4774, arXiv:1705.07874.
S. Wachter, B. Mittelstadt, and C. Russell, “Counterfactual explanations without opening the black box: Automated decisions and the GDPR,” Harvard J. Law Technol., vol. 31, no. 2, pp. 841–887, 2018, doi: 10.2139/ssrn.3063289.
D. W. Apley and J. Zhu, “Visualizing the effects of predictor variables in black box supervised learning models,” J. Royal Stat. Soc. Series B, vol. 82, no. 4, pp. 1059–1086, 2020, doi: 10.1111/rssb.12377.
J. Jiarpakdee, C. Tantithamthavorn, and A. E. Hassan, “The impact of correlated metrics on the interpretation of defect models,” IEEE Trans. Softw. Eng., vol. 47, no. 2, pp. 320–331, 2021, doi: 10.1109/TSE.2019.2891758.
A. A. Ismail, M. Gunady, H. Corrada Bravo, and S. Feizi, “Benchmarking deep learning interpretability in time series predictions,” in Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 6441–6452, arXiv:2010.13924.
C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On calibration of modern neural networks,” in Proc. 34th Int. Conf. Mach. Learn., ser. PMLR, vol. 70, 2017, pp. 1321–1330, arXiv:1706.04599.
T. Saito and M. Rehmsmeier, “The precision–recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets,” PLoS ONE, vol. 10, no. 3, Art. no. e0118432, 2015, doi: 10.1371/journal.pone.0118432.
C. Tantithamthavorn, S. McIntosh, A. E. Hassan, and K. Matsumoto, “An empirical comparison of model validation techniques for defect prediction models,” IEEE Trans. Softw. Eng., vol. 43, no. 1, pp. 1–18, 2017, doi: 10.1109/TSE.2016.2584050.
H. Ishwaran, U. B. Kogalur, E. H. Blackstone, and M. S. Lauer, “Random survival forests,” Ann. Appl. Stat., vol. 2, no. 3, pp. 841–860, 2008, doi: 10.1214/08-AOAS169.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 The Author(s)

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.