A Novel Approach Based on Deep Learning for Violence Detection in Public Places

Authors

  • Shah Faisal Khan Department of Computer Science, University of Science and Technology Bannu, KP, Pakistan
  • Said Khalid Shah Department of Computer Science, University of Science and Technology Bannu, KP, Pakistan
  • Faheem Ullah Khan Department of Software Engineering, University of Science and Technology Bannu, KP, Pakistan
  • Wasiat Khan Department of Software Engineering, University of Science and Technology Bannu, KP, Pakistan
  • Fouzia Idrees Department of Computer Science, Shaheed Benazir Bhutto Women University Peshawar, KP, Pakistan

DOI:

https://doi.org/10.33411/IJIST/1773

Keywords:

Violence Detection, Public-Place Surveillance, Key-Frame Selection, Activity Recognition, Transfer Learning, ResNet-50

Abstract

This study presents a lightweight deep learning pipeline for violence detection in public place surveillance imagery. The central idea is to perform offline key frame selection before model training so that redundant, transitional, and visually ambiguous frames are discarded, and only informative violent or non-violent scenes are retained. Unlike clip-level methods that rely on dense temporal stacks or computationally heavy 3D networks, the proposed approach uses curated key frames to construct a balanced 7,000-image dataset and then trains a transfer learning classifier on those frames. Images are resized to 256 × 256, normalized, and augmented before being processed by a ResNet-50 backbone, followed by two fully connected layers of 512 and 256 units and dropout regularization. On Hockey Fight dataset test images, the model achieves 97.25% accuracy, 97.62% precision, 96.45% recall, and 97.03% F1-score, outperforming EfficientNet-B0 and VGG-19 under the same experimental pipeline. The results indicate that careful offline frame curation can provide a practical compromise between computational efficiency and recognition performance for early violence screening in surveillance settings. The paper also discusses current limitations, including manual curation bias, the absence of explicit temporal modeling, and the lack of actor-level localization.

References

L. Minh Dang, Kyungbok Min, “Sensor-based and vision-based human activity recognition: A comprehensive survey,” Pattern Recognit., vol. 108, p. 107561, 2020, doi: https://doi.org/10.1016/j.patcog.2020.107561.

Yanjinlkham Myagmar-Ochir, Wooseong Kim, “A Survey of Video Surveillance Systems in Smart City,” Electronics, vol. 12, no. 17, p. 3567, 2023, doi: https://doi.org/10.3390/electronics12173567.

N. Mumtaz et al., “An overview of violence detection techniques: current challenges and future directions,” Artif. Intell. Rev. 2022 565, vol. 56, no. 5, pp. 4641–4666, Oct. 2022, doi: 10.1007/S10462-022-10285-3.

Fath U.Min Ullah, Amin Ullah, “Violence Detection Using Spatiotemporal Features with 3D Convolutional Neural Network,” Sensors, vol. 19, no. 11, p. 2472, 2019, doi: https://doi.org/10.3390/s19112472.

Simone Accattoli, Paolo Sernani, “Violence Detection in Videos by Combining 3D Convolutional Neural Networks and Support Vector Machines,” Appl. Artif. Intell., vol. 34, no. 4, 2020, [Online]. Available: https://www.tandfonline.com/doi/full/10.1080/08839514.2020.1723876

Romas Vijeikis, Vidas Raudonis, “Efficient Violence Detection in Surveillance,” Sensors, vol. 22, no. 6, p. 2216, 2022, doi: https://doi.org/10.3390/s22062216.

Guillermo Garcia-Cobo, Juan C. SanMiguel, “Human skeletons and change detection for efficient violence detection in surveillance videos,” Comput. Vis. Image Underst., vol. 233, p. 103739, 2023, doi: https://doi.org/10.1016/j.cviu.2023.103739.

Fernando J. Rendón-Segador, Juan A. Álvarez-García, “CrimeNet: Neural Structured Learning using Vision Transformer for violence detection,” Neural Networks, vol. 161, pp. 318–329, 2023, doi: https://doi.org/10.1016/j.neunet.2023.01.048.

A. Datta, M. Shah, and N. Da Vitoria Lobo, “Person-on-person violence detection in video data,” Proc. - Int. Conf. Pattern Recognit., vol. 16, no. 1, pp. 433–438, 2002, doi: 10.1109/ICPR.2002.1044748.

E. Bermejo Nievas, O. Deniz Suarez, G. Bueno García, and R. Sukthankar, “Violence Detection in Video Using Computer Vision Techniques,” Lect. Notes Comput. Sci. (including Subser. Lect. Notes Artif. Intell. Lect. Notes Bioinformatics), vol. 6855 LNCS, no. PART 2, pp. 332–339, 2011, doi: 10.1007/978-3-642-23678-5_39.

P. Bilinski and F. Bremond, “Human violence recognition and detection in surveillance videos,” 2016 13th IEEE Int. Conf. Adv. Video Signal Based Surveillance, AVSS 2016, pp. 30–36, Nov. 2016, doi: 10.1109/AVSS.2016.7738019.

Muhammad Shoaib, Nasir Sayed, “A Deep Learning Based System for the Detection of Human Violence in Video Data,” Trait. du signal, vol. 38, no. 6, p. 12, 2021, [Online]. Available: https://www.iieta.org/journals/ts/paper/10.18280/ts.380606

V. Akula and I. Kavati, “Human Violence Detection in Videos Using Key Frame Identification and 3D CNN with Convolutional Block Attention Module,” Circuits, Syst. Signal Process. 2024 4312, vol. 43, no. 12, pp. 7924–7950, Aug. 2024, doi: 10.1007/S00034-024-02824-W.

Naz Dündar, Ali Seydi Keçeli, “A shallow 3D convolutional neural network for violence detection in videos,” Egypt. Informatics J., vol. 26, p. 100455, 2024, doi: https://doi.org/10.1016/j.eij.2024.100455.

Marwa Qaraqe, Yin David Yang, Elizabeth B Varghese, Emrah Basaran & Almiqdad Elzein, “Crowd behavior detection: leveraging video swin transformer for crowd size and violence level analysis,” Appl. Intell., vol. 54, pp. 10709–10730, 2024, [Online]. Available: https://link.springer.com/article/10.1007/s10489-024-05775-6

Vesal Khean, Chomyong Kim, “Human Interaction Recognition in Surveillance Videos Using Hybrid Deep Learning and Machine Learning Models,” Comput. Mater. Contin., vol. 81, no. 1, pp. 773–787, 2024, doi: https://doi.org/10.32604/cmc.2024.056767.

J. Mahmoodi and H. Nezamabadi-pour, “A spatio-temporal model for violence detection based on spatial and temporal attention modules and 2D CNNs,” Pattern Anal. Appl. 2024 272, vol. 27, no. 2, pp. 46-, Apr. 2024, doi: 10.1007/S10044-024-01265-0.

P. Zhang, L. Dong, X. Zhao, W. Lei, and W. Zhang, “An end-to-end framework for real-time violent behavior detection based on 2D CNNs,” J. Real-Time Image Process. 2024 212, vol. 21, no. 2, pp. 57-, Mar. 2024, doi: 10.1007/S11554-024-01443-7.

V. A. M. Chidambaram and K. P. Chandrasekaran, “F3DNN-Net: behaviours violence detection via fine-tuned fused feature based deep neural network from surveillance video,” Signal, Image Video Process. 2024 1811, vol. 18, no. 11, pp. 7655–7669, Aug. 2024, doi: 10.1007/S11760-024-03418-4.

Gurmeet Kaur, Sarbjeet Singh, “Revisiting vision-based violence detection in videos: A critical analysis,” Neurocomputing, vol. 597, p. 128113, 2024, doi: https://doi.org/10.1016/j.neucom.2024.128113.

Pablo Negre, Ricardo S. Alonso, “Literature Review of Deep-Learning-Based Detection of Violence in Video,” Sensors, vol. 24, no. 12, p. 4016, 2024, doi: https://doi.org/10.3390/s24124016.

Fernando J. Rendón-Segador, Juan A. Álvarez-García, “Transformer and Adaptive Threshold Sliding Window for Improving Violence Detection in Videos,” Sensors, vol. 24, no. 16, p. 5429, 2024, doi: https://doi.org/10.3390/s24165429.

Abu Bakar Siddique Mahi, Farhana Sultana Eshita, “VID: A comprehensive dataset for violence detection in various contexts,” Data Br., vol. 57, p. 110875, 2024.

Damith Chamalke Senadeera, Xiaoyun Yang, Dimitrios Kollias, Gregory Slabaugh, “CUE-Net: Violence Detection Video Analytics with Spatial Cropping, Enhanced UniformerV2 and Modified Efficient Additive Attention,” arXiv:2404.18952, 2024, [Online]. Available: https://arxiv.org/abs/2404.18952

F. Cao, Y. Miao, and W. Zhang, “Implementation and Application of Violence Detection System Based on Multi-head Attention and LSTM,” Lect. Notes Comput. Sci. (including Subser. Lect. Notes Artif. Intell. Lect. Notes Bioinformatics), vol. 14868 LNCS, pp. 77–88, 2024, doi: 10.1007/978-981-97-5600-1_7.

N. Mumtaz, N. Ejaz, I. Rida, M. A. Khan, and M. Y. Lee, “Towards Real-world Violence Recognition via Efficient Deep Features and Sequential Patterns Analysis,” Mob. Networks Appl., vol. 29, no. 4, pp. 1326–1335, Aug. 2024, doi: 10.1007/S11036-024-02319-7/METRICS.

L. Hsairi, S. M. Alosaimi, and G. A. Alharaz, “Violence Detection Using Deep Learning,” Arab. J. Sci. Eng. 2024 5015, vol. 50, no. 15, pp. 11669–11679, Sep. 2024, doi: 10.1007/S13369-024-09536-Y.

M. Evany Anne, M. Brindha, and N. Sivakumaran, “Advancing intelligent surveillance: A comprehensive hybrid deep learning framework for anomaly detection, violence recognition, and person re-identification,” Multimed. Tools Appl. 2025 8438, vol. 84, no. 38, pp. 46863–46909, Jul. 2025, doi: 10.1007/S11042-025-21005-8.

N. Tran, H. Nguyen, D. Ly, K. Ngo, and H. D. Nguyen, “Advancing Violence Detection with Graph-Based Skeleton Motion Analysis,” SN Comput. Sci. 2025 66, vol. 6, no. 6, pp. 595-, Jun. 2025, doi: 10.1007/S42979-025-04118-7.

Hong Huang, Qingping Jiang, “IDG-ViolenceNet: A Video Violence Detection Model Integrating Identity-Aware Graphs and 3D-CNN,” Sensors, vol. 25, no. 20, 2025, doi: https://doi.org/10.3390/s25206272.

Chenghao Li, Gang Liang, “MEN-VVDF: Multipath excitation network-based video violence detection framework focusing on human activity in keyframes,” J. Vis. Commun. Image Represent., vol. 112, p. 104573, 2025, doi: https://doi.org/10.1016/j.jvcir.2025.104573.

Downloads

Published

2026-02-14
CITATION
Published: 2026-02-14
Crossref Citation Count: Loading...

How to Cite

Shah Faisal Khan, Said Khalid Shah, Faheem Ullah Khan, Wasiat Khan, & Fouzia Idrees. (2026). A Novel Approach Based on Deep Learning for Violence Detection in Public Places. International Journal of Innovations in Science & Technology, 8(1), 406–415. https://doi.org/10.33411/IJIST/1773