Automating Severity Prediction of Bug Reports Using Large Language Models (LLM)
Keywords:
Large Language Model, Machine Learning, Natural Language Processing, Deep Learning, Meta AIAbstract
Accurate prediction of bug report severity is essential for prioritizing software-maintenance activities, but manual severity assignment is costly, time-consuming, and potentially inconsistent. This study investigates prompt-based open-weight Large Language Models for automatically classifying bug reports as severe or non-severe without task-specific fine-tuning. The experiments used 3,535 bug reports from the CDT, JDT, Bugzilla, and Thunderbird projects. The original severity labels were mapped into two categories: blocker, critical, and major as severe, and normal, minor, and trivial as non-severe. LLaMA 3.1:8B and Qwen3:8B were evaluated under zero-shot and few-shot prompting and compared with Support Vector Machine, Multinomial Naive Bayes, Random Forest, and Long Short-Term Memory models using TF-IDF and Word2Vec representations. On the held-out test set, Qwen3:8B with few-shot prompting achieved 98.6% accuracy, 100% precision, 97.5% recall, and a 98.7% F1-score. Compared with the strongest conventional baseline, Random Forest with TF-IDF, it improved accuracy and F1-score by 26.6 and 24.2 percentage points, respectively. On an additional dataset of 200 synthetically generated bug reports, LLaMA 3.1:8B with zero-shot prompting achieved 97.99% accuracy and a 98.51% F1-score, while Qwen3:8B with few-shot prompting achieved 97.9% accuracy and a 97.5% F1-score. These results indicate that prompt-based LLMs can substantially reduce task-specific preprocessing and training requirements while providing high predictive performance. However, further evaluation on independent, human-authored, cross-project datasets is required before making broad claims regarding real-world generalization.
References
“(PDF) Software Engineering: A Practitioner’s Approach 9 th Edition.” Accessed: Jul. 06, 2026. [Online]. Available: https://www.researchgate.net/publication/365946272_Software_Engineering_A_Practitioner’s_Approach_9_th_Edition
Muhammad Ali Arshad, Adnan Riaz, “SevPredict: Exploring the Potential of Large Language Models in Software Maintenance,” AI, vol. 5, no. 4, pp. 2739–2760, 2024, doi: https://doi.org/10.3390/ai5040132.
P. Mahajan, K. Choudhary, N. Rana, R. Kumar, and S. Deshmukh, “Software Bugs Classification Using SVM, RF, DT Algorithms,” Lect. Notes Networks Syst., vol. 1237, pp. 51–62, 2025, doi: 10.1007/978-981-96-1185-0_5/SAVE-RESEARCH.
G. Sharma and P. Singh, “Comparative Study: Word2Vec Versus TF-IDF in Software Defect Predictions,” Lect. Notes Networks Syst., vol. 1165, pp. 95–107, 2025, doi: 10.1007/978-981-97-8336-6_8/SAVE-RESEARCH.
Ömer Köksal, Bedir Tekinerdogan, “Automated Classification of Unstructured Bilingual Software Bug Reports: An Industrial Case Study Research,” Appl. Sci., vol. 12, no. 1, p. 338, 2022, doi: https://doi.org/10.3390/app12010338.
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, “LLaMA: Open and Efficient Foundation Language Models,” arXiv:2302.13971, 2023, [Online]. Available: https://arxiv.org/abs/2302.13971
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, Graham Neubig, “Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing,” arXiv:2107.13586, 2021, [Online]. Available: https://arxiv.org/abs/2107.13586
“Application of Large Language Models to Software Engineering Tasks: Opportunities, Risks, and Implications | IEEE Journals & Magazine | IEEE Xplore.” Accessed: Jul. 06, 2026. [Online]. Available: https://ieeexplore.ieee.org/document/10109345
Ehsan Mashhadi, Shaiful Chowdhury, “An empirical study on bug severity estimation using source code metrics and static analysis,” J. Syst. Softw., vol. 217, p. 112179, 2024, doi: https://doi.org/10.1016/j.jss.2024.112179.
Quanjun Zhang, Chunrong Fang, Yuxiang Ma, Weisong Sun, Zhenyu Chen, “A Survey of Learning-based Automated Program Repair,” arXiv:2301.03270, 2023, [Online]. Available: https://arxiv.org/abs/2301.03270
Kevin Moran, “Enhancing Android Application Bug Reporting,” arXiv:1706.01118, 2017, [Online]. Available: https://arxiv.org/abs/1706.01118
Ashima Kukkar, Rajni Mohana, “Does bug report summarization help in enhancing the accuracy of bug severity classification?,” Procedia Comput. Sci., vol. 167, pp. 1345–1353, 2020, doi: https://doi.org/10.1016/j.procs.2020.03.345.
“Deep Neural Network-Based Severity Prediction of Bug Reports | IEEE Journals & Magazine | IEEE Xplore.” Accessed: Jul. 06, 2026. [Online]. Available: https://ieeexplore.ieee.org/document/8685106
J. N. Singh, K. V. Gupta, S. Kumar, R. Priya, and K. Saini, “Bugzilla Bug Tracking System,” Proc. - IEEE 2023 5th Int. Conf. Adv. Comput. Commun. Control Networking, ICAC3N 2023, pp. 1522–1526, 2023, doi: 10.1109/ICAC3N60023.2023.10541310.
“Revolutionize IT Support with Jira Service Management | Atlassian.” Accessed: Jul. 06, 2026. [Online]. Available: https://www.atlassian.com/software/jira/service-management?campaign=20339374771&adgroup=152515828964&targetid=kwd-1683663083742&matchtype=p&network=g&device=c&device_model=&creative=672220291174&keyword=jira service management ticketing tool&placement=&target=&ds_eid=700000001721198&ds_e1=GOOGLE&gad_source=1&gad_campaignid=20339374771&gbraid=0AAAAADAVAKcuW8U5-PmnTAmLnQcfu_ALU&gclid=CjwKCAjwpK3SBhASEiwAtV1SPFyF4746okaEe5U82fQqux4wYXLEjq3UmoO_e5zJPiaSBOmN0QB-5hoCGYsQAvD_BwE
M. A. Arshad and H. Zhiqiu, “Using CNN to Predict the Resolution Status of Bug Reports,” J. Phys. Conf. Ser., vol. 1828, no. 1, p. 012106, Feb. 2021, doi: 10.1088/1742-6596/1828/1/012106.
Antonios Saravanos, Matthew X. Curinga, “Simulating the Software Development Lifecycle: The Waterfall Model,” Appl. Syst. Innov., vol. 6, no. 6, p. 108, 2023, doi: https://doi.org/10.3390/asi6060108.
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, Jeffrey Dean, “Distributed Representations of Words and Phrases and their Compositionality,” arXiv:1310.4546, 2013, [Online]. Available: https://arxiv.org/abs/1310.4546
Yang, J. Kim and G., “Bug Severity Prediction Algorithm Using Topic-Based Feature Selection and CNN-LSTM Algorithm,” IEEE Access, vol. 10, pp. 94643–94651, 2022, doi: 10.1109/ACCESS.2022.3204689.
“A Light Bug Triage Framework for Applying Large Pre-trained Language Model | Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering.” Accessed: Jul. 11, 2026. [Online]. Available: https://dl.acm.org/doi/10.1145/3551349.3556898
A. K. Dipongkor and K. Moran, “A Comparative Study of Transformer-Based Neural Text Representation Techniques on Bug Triaging,” Proc. - 2023 38th IEEE/ACM Int. Conf. Autom. Softw. Eng. ASE 2023, pp. 1012–1023, 2023, doi: 10.1109/ASE56229.2023.00217.
Asif Ali, Yuanqing Xia, “BERT based severity prediction of bug reports for the maintenance of mobile applications,” J. Syst. Softw., vol. 208, p. 111898, 2024, doi: https://doi.org/10.1016/j.jss.2023.111898.
“Graph Neural Network vs. Large Language Model: A Comparative Analysis for Bug Report Priority and Severity Prediction | Proceedings of the 20th International Conference on Predictive Models and Data Analytics in Software Engineering.” Accessed: Jul. 11, 2026. [Online]. Available: https://dl.acm.org/doi/10.1145/3663533.3664042
M. Theingi and N. L. Wah, “Automated Prediction of Bug Reports’ Severity using Attention-based Convolutional Neural Network,” Proc. 22nd IEEE Int. Conf. Comput. Appl. ICCA 2025, 2025, doi: 10.1109/ICCA65395.2025.11011201.
F. Thung, D. Lo, and L. Jiang, “Automatic defect categorization,” Proc. - Work. Conf. Reverse Eng. WCRE, pp. 205–214, 2012, doi: 10.1109/WCRE.2012.30.
Piotr Bojanowski, Edouard Grave, Armand Joulin, Tomas Mikolov, “Enriching Word Vectors with Subword Information,” arXiv:1607.04606, 2017, [Online]. Available: https://arxiv.org/abs/1607.04606
S. Kumar, S. Muttoo, and V. B. Singh, “Classification of Software Defects Using Orthogonal Defect Classification,” Int. J. Open Source Softw. Process., vol. 13, no. 1, pp. 1–16, Jan. 2022, doi: 10.4018/IJOSSP.300749.
Y. Zhou, Y. Tong, R. Gu, and H. Gall, “Combining text mining and data mining for bug report classification,” Proc. - 30th Int. Conf. Softw. Maint. Evol. ICSME 2014, pp. 311–320, Dec. 2014, doi: 10.1109/ICSME.2014.53.
S. Kumar, M. Sharma, V. B. Singh, and S. K. Muttoo, “Bug Report Classification into Orthogonal Defect Classification Defect Type using Long Short Term Memory,” Proc. - 2021 3rd Int. Conf. Adv. Comput. Commun. Control Networking, ICAC3N 2021, pp. 285–287, 2021, doi: 10.1109/ICAC3N53548.2021.9725398.
“CaPBug-A Framework for Automatic Bug Categorization and Prioritization Using NLP and Machine Learning Algorithms | IEEE Journals & Magazine | IEEE Xplore.” Accessed: Jul. 06, 2026. [Online]. Available: https://ieeexplore.ieee.org/document/9388660
Ahmed Fawzi Otoom, Sara Al-Jdaeh, “Automated Classification of Software Bug Reports,” ACM Int. Conf. Proceeding Ser., 2019, [Online]. Available: https://dl.acm.org/doi/10.1145/3357419.3357424
Ashima Kukkar, Rajni Mohana, “A Supervised Bug Report Classification with Incorporate and Textual field Knowledge,” Procedia Comput. Sci., vol. 132, pp. 352–361, 2018, doi: https://doi.org/10.1016/j.procs.2018.05.194.
A. Lamkanfi, J. Pérez, and S. Demeyer, “The eclipse and mozilla defect tracking dataset: A genuine dataset for mining bug information,” IEEE Int. Work. Conf. Min. Softw. Repos., pp. 203–206, 2013, doi: 10.1109/MSR.2013.6624028.
S. Suthaharan, “Machine Learning Models and Algorithms for Big Data Classification,” vol. 36, 2016, doi: 10.1007/978-1-4899-7641-3.
H. M. Tran, S. Van Nguyen, S. V. U. Ha, and T. Q. Le, “An analysis of software bug reports using random forest,” Lect. Notes Comput. Sci. (including Subser. Lect. Notes Artif. Intell. Lect. Notes Bioinformatics), vol. 11251 LNCS, pp. 273–285, 2018, doi: 10.1007/978-3-030-03192-3_21/SAVE-RESEARCH.
Ahmed Fawzi Otoom, Sara Al-Jdaeh, “An implementation of naive bayes classifier,” ACM Int. Conf. Proceeding Ser., 2019, [Online]. Available: https://dl.acm.org/doi/10.1145/3357419.3357424
X. Ye, F. Fang, J. Wu, R. Bunescu, and C. Liu, “Bug Report Classification Using LSTM Architecture for More Accurate Software Defect Locating,” Proc. - 17th IEEE Int. Conf. Mach. Learn. Appl. ICMLA 2018, pp. 1438–1445, Jul. 2018, doi: 10.1109/ICMLA.2018.00234.
“Classifying Bug Reports into Bugs and Non-bugs Using LSTM | Proceedings of the 10th Asia-Pacific Symposium on Internetware.” Accessed: Jul. 06, 2026. [Online]. Available: https://dl.acm.org/doi/10.1145/3275219.3275239
Edward Loper, Steven Bird, “NLTK: The Natural Language Toolkit,” arXiv:cs/0205028, 2002, [Online]. Available: https://arxiv.org/abs/cs/0205028
“TextBlob: Simplified Text Processing — TextBlob 0.19.0 documentation.” Accessed: Jul. 06, 2026. [Online]. Available: https://textblob.readthedocs.io/en/dev/
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 50sea

This work is licensed under a Creative Commons Attribution 4.0 International License.


















