Early Prediction of GitHub Issue Escalation Using Text and Collaboration Features

Authors

  • I Wayan Budi Sentana School of Computing, Faculty of Science and Engineering, Macquarie University, Sydney, NSW, Australia; Department of Information Technology, Politeknik Negeri Bali, Jimbaran, Badung, Bali, Indonesia Author
  • Gurdip Kaur Saminder Singh Open University Malaysia (OUM), Kuala Lumpur Author

Keywords:

GitHub Issues, Issue Escalation, Mining Software Repositories, Natural Language Processing, Collaboration Networks, Early Prediction, Software Maintenance

Abstract

Problem: GitHub issues can develop into prolonged, contentious, or coordination-intensive episodes, but escalation is usually recognized only after substantial maintainer effort has been spent.

Motivation: An early warning could help maintainers prioritize attention while preserving human judgment.

Gap: Existing work largely predicts assignment, severity, priority, or resolution time; these targets do not jointly represent escalation, and many approaches underuse the interaction process visible at the beginning of an issue.

Framework: This paper proposes a leakage-aware conceptual framework that predicts escalation at fixed early cutoffs, including issue creation, the first maintainer response, and the first few comments. Observable future events—such as reopening, priority promotion, cross-repository transfer or reference, senior-contributor assignment, abnormal unresolved duration, discussion growth, and validated conflict—define the outcome but are excluded from predictors. Early predictors combine title, body, and comment representations with participant growth, response timing, interaction networks, maintainer involvement, and repository controls. Interpretable linear and tree models and multimodal neural fusion are specified, together with calibrated, alert-budget-aware evaluation.

Contributions: The proposal provides an operational escalation construct, a temporal feature taxonomy, a reproducible multi-repository methodology, and an explanation layer for maintainer-facing alerts. No empirical performance claim is made; the expected contribution is a testable design for future evaluation and responsible deployment.

References

T. F. Bissyandé, D. Lo, L. Jiang, L. Réveillère, J. Klein, and Y. L. Traon, “Got issues? who cares about it? a large scale investigation of issue trackers from GitHub,” in 2013 IEEE 24th International Symposium on Software Reliability Engineering (ISSRE), 2013, pp. 188–197.

L. Dabbish, C. Stuart, J. Tsay, and J. Herbsleb, “Social coding in GitHub: Transparency and collaboration in an open software repository,” in Proceedings of the ACM 2012 Conference on Computer Supported Cooperative Work, 2012, pp. 1277–1286.

E. Kalliamvakou, G. Gousios, K. Blincoe, L. Singer, D. M. German, and D. Damian, “The promises and perils of mining GitHub,” in Proceedings of the 11th Working Conference on Mining Software Repositories, 2014, pp. 92–101.

J. Tsay, L. Dabbish, and J. Herbsleb, “Let’s talk about it: Evaluating contributions through discussion in GitHub,” in Proceedings of the 22nd ACM SIGSOFT International Symposium on Foundations of Software Engineering, 2014, pp. 144–154.

J. Anvik, L. Hiew, and G. C. Murphy, “Who should fix this bug?” in Proceedings of the 28th International Conference on Software Engineering, 2006, pp. 361–370.

G. Jeong, S. Kim, and T. Zimmermann, “Improving bug triage with bug tossing graphs,” in Proceedings of the 7th Joint Meeting of the European Software Engineering Conference and the ACM SIGSOFT Symposium on the Foundations of Software Engineering, 2009, pp. 111–120.

A. Lamkanfi, S. Demeyer, E. Giger, and B. Goethals, “Predicting the severity of a reported bug,” in 2010 7th IEEE Working Conference on Mining Software Repositories, 2010, pp. 1–10.

E. Giger, M. Pinzger, and H. C. Gall, “Predicting the fix time of bugs,” in Proceedings of the 2nd International Workshop on Recommendation Systems for Software Engineering, 2010, pp. 52–56.

R. Kikas, M. Dumas, and D. Pfahl, “Using dynamic and contextual features to predict issue lifetime in GitHub projects,” in Proceedings of the 13th International Conference on Mining Software Repositories, 2016, pp. 291–302.

M. Izadi, S. Akbari, and A. Heydarnoori, “Predicting the objective and priority of issue reports in software repositories,” Empirical Software Engineering, vol. 27, no. 2, p. 50, 2022.

E. Shihab, A. Ihara, Y. Kamei, W. M. Ibrahim, M. Ohira, B. Adams, A. E. Hassan, and K. ichi Matsumoto, “Studying re-opened bugs in open source software,” Empirical Software Engineering, vol. 18, no. 5, pp. 1005–1042, 2013.

N. Novielli, F. Calefato, and F. Lanubile, “The challenges of sentiment detection in the social programmer ecosystem,” in Proceedings of the 7th International Workshop on Social Software Engineering, 2015, pp. 33–40.

F. Calefato, F. Lanubile, F. Maiorano, and N. Novielli, “Sentiment polarity detection for software development,” Empirical Software Engineering, vol. 23, no. 3, pp. 1352–1382, 2018.

G. Destefanis, M. Ortu, S. Counsell, S. Swift, M. Marchesi, and R. Tonelli, “Software development: Do good manners matter?” PeerJ Computer Science, vol. 2, p. e73, 2016.

C. Miller, S. Cohen, D. Klug, B. Vasilescu, and C. Kästner, “Did you miss my comment or what? understanding toxicity in open source discussions,” in Proceedings of the 44th International Conference on Software Engineering, 2022, pp. 710–722.

I. Ferreira, B. Adams, and J. Cheng, “How heated is it? understanding GitHub locked issues,” in Proceedings of the 19th International Conference on Mining Software Repositories, 2022, pp. 309–320.

H. S. Qiu, B. Vasilescu, C. Kästner, S. Egelman, C. Jaspan, and E. Murphy-Hill, “Detecting interpersonal conflict in issues and code review: Cross-pollinating open- and closed-source approaches,” in Proceedings of the 44th International Conference on Software Engineering: Software Engineering in Society, 2022, pp. 41–55.

C. Bird, D. Pattison, R. D’Souza, V. Filkov, and P. Devanbu, “Latent social structure in open source projects,” in Proceedings of the 16th ACM SIGSOFT International Symposium on Foundations of Software Engineering, 2008, pp. 24–35.

A. Meneely and L. Williams, “Socio-technical developer networks: Should we trust our measurements?” in Proceedings of the 33rd International Conference on Software Engineering, 2011, pp. 281–290.

M. Joblin, W. Mauerer, S. Apel, J. Siegmund, and D. Riehle, “From developer networks to verified communities: A fine-grained approach,” in 2015 IEEE/ACM 37th IEEE International Conference on Software Engineering, vol. 1, 2015, pp. 563–573.

J. Tsay, L. Dabbish, and J. Herbsleb, “Influence of social and technical factors for evaluating contribution in GitHub,” in Proceedings of the 36th International Conference on Software Engineering, 2014, pp. 356–366.

G. Gousios and D. Spinellis, “GHTorrent: GitHub’s data from a firehose,” in 2012 9th IEEE Working Conference on Mining Software Repositories, 2012, pp. 12–21.

M. Golzadeh, A. Decan, D. Legay, and T. Mens, “A ground-truth dataset and classification model for detecting bots in GitHub issue and PR comments,” Journal of Systems and Software, vol. 175, p. 110911, 2021.

J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, vol. 1, 2019, pp. 4171–4186.

N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using siamese BERT-networks,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, 2019, pp. 3982–3992.

Z. Feng, D. Guo, D. Tang, N. Duan, X. Feng, M. Gong, L. Shou, B. Qin, T. Liu, D. Jiang, and M. Zhou, “CodeBERT: A pre-trained model for programming and natural languages,” in Findings of the Association for Computational Linguistics: EMNLP 2020, 2020, pp. 1536–1547.

S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Advances in Neural Information Processing Systems 30, 2017, pp. 4765–4774.

C. Rudin, “Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead,” Nature Machine Intelligence, vol. 1, no. 5, pp. 206–215, 2019.

C. Tantithamthavorn, S. McIntosh, A. E. Hassan, and K. ichi Matsumoto, “An empirical comparison of model validation techniques for defect prediction models,” IEEE Transactions on Software Engineering, vol. 43, no. 1, pp. 1–18, 2017.

D. Falessi, J. Huang, L. Narayana, J. F. Thai, and B. Turhan, “On the need of preserving order of data when validating within-project defect classifiers,” Empirical Software Engineering, vol. 25, no. 6, pp. 4805–4830, 2020.

Y. Kamei, E. Shihab, B. Adams, A. E. Hassan, A. Mockus, A. Sinha, and N. Ubayashi, “A large-scale empirical study of just-in-time quality assurance,” IEEE Transactions on Software Engineering, vol. 39, no. 6, pp. 757–773, 2013.

B. Krawczyk, “Learning from imbalanced data: Open challenges and future directions,” Progress in Artificial Intelligence, vol. 5, no. 4, pp. 221–232, 2016.

T. Saito and M. Rehmsmeier, “The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets,” PLOS ONE, vol. 10, no. 3, p. e0118432, 2015.

N. E. Gold and J. Krinke, “Ethics in the mining of software repositories,” Empirical Software Engineering, vol. 27, no. 1, p. 17, 2022.

Downloads

Published

2026-09-30

Issue

Section

Articles