An End-to-End Deep Learning Pipeline for Child Detection and Activity Recognition in Surveillance Systems

Authors

  • Samad Riaz Department of Electrical Engineering, UET Peshawar
  • Shayan Riaz Center for Intelligent Systems and Network Research, UET Peshawar, Pakistan.
  • Abdul Razzaq Center for Intelligent Systems and Network Research, UET Peshawar, Pakistan.
  • Umar Sadique Department: Computer Systems Engineering UEAS Swat, Pakistan.
  • Shahid Bashir Department of Electrical Engineering, UET Peshawar, Pakistan.
  • Jehad Ur Rahman Department of Electrical Engineering, UET Peshawar, Pakistan.

DOI:

https://doi.org/10.33411/IJIST/1873

Keywords:

Child Safety, Activity Recognition, Deep Learning, Computer Vision, Video Surveillance

Abstract

Children in developing countries often face significant safety risks such as domestic accidents, unsafe environments, and abusive social exposure. The current surveillance systems are able to identify children but, in most cases, they fail to comprehend their activities over time. This creates a critical gap in real-time child safety monitoring. This paper addresses that gap by developing an end-to-end video understanding pipeline for child detection in video frames and recognizes their activities accurately over time. A custom child detection dataset containing 19,890 annotated images was developed and split into 16,107 training and 3,783 validation samples. Several state-of-the-art object detection models, including YOLOv8s, YOLO26s, and RT-DETR-L, were trained and compared based on their performance metrics. DeepSORT is used to preserve the identities of children detected across frames. For activity recognition, a custom video dataset comprising approximately 1,960 clips across 47 activity classes was created. Four deep learning architectures (I3D, ResNet3D, ViViT, and VideoMAE) were trained using standardized configurations with 16-frame clips, 224×224 input resolution, Adam optimization, and cross-entropy loss. YOLO26s achieved the best detection performance with Precision = 97.17%, Recall = 96.18%, mAP@0.50 = 98.51%, and mAP@0.50–0.95 = 87.80%. For activity recognition, VideoMAE outperformed competing models on 448 clips with test accuracy of 86.38% and Macro F1-score of 90.49%, significantly surpassing ViViT (79.24%), ResNet3D (51.34%), and I3D (9.60%). The proposed method provides a scalable and practical solution by combining high-quality spatial localization with temporal activity understanding, enabling real-time monitoring of child activities.

References

“FAST FACTS: Violence against children widespread, affecting millions globally.” Accessed: May 09, 2026. [Online]. Available: https://www.unicef.org/press-releases/fast-facts-violence-against-children-widespread-affecting-millions-globally

K. R. Tanveer, M. S. Luqman, and A. Qureshi, “Ensuring Child Safety: An IoT-Based Surveillance System for Remote Monitoring and Detection of Anomalous Behavior,” pp. 1–5, Nov. 2024, doi: 10.1109/ICETST62952.2024.10737980.

J. H. Tan and C. P. Goh, “Enhancing Child Safety: Computer Vision-Based Accident Detection for Infants and Toddlers,” 2024 3rd Int. Conf. Digit. Transform. Appl., pp. 179–183, 2024, doi: 10.1109/ICDXA61007.2024.10470712.

G. Singh, A. R. Shekhar, X. Yu, and J. Saniie, “Smart Infant Monitoring System Using Computer Vision and AI,” IEEE Int. Conf. Electro Inf. Technol., vol. 2023-May, pp. 347–352, 2023, doi: 10.1109/EIT57321.2023.10187295.

A. Prathyanga, P. Shyaminda, P. Chamikara, S. Lakshan, S. Thelijjagoda, and D. Kasthurirathna, “Intelligent Daycare: Enhancing Child Safety with IoT and Machine Learning Innovations,” Proc. 9th Int. Conf. Commun. Electron. Syst. ICCES 2024, pp. 530–538, 2024, doi: 10.1109/ICCES63552.2024.10859472.

Shehzad Ali, Md Tanvir Islam, “CABAD: A video dataset for benchmarking child aggression recognition,” Alexandria Eng. J., vol. 127, pp. 1143–1157, 2025, doi: https://doi.org/10.1016/j.aej.2025.06.035.

Lucia Migliorelli, Sara Moccia, “The babyPose dataset,” Data Br., 2020, [Online]. Available: https://pubmed.ncbi.nlm.nih.gov/33083503/

Pengbo Wei, David Ahmedt-Aristizabal, “Vision-based activity recognition in children with autism-related behaviors,” Heliyon, vol. 9, no. 6, 2023, doi: https://doi.org/10.1016/j.heliyon.2023.e16763.

Somaieh Amraee, Bishoy Galoaa, Matthew Goodwin, Elaheh Hatamimajoumerd, Sarah Ostadabbas, “Multiple Toddler Tracking in Indoor Videos,” arXiv:2311.17656, 2023, [Online]. Available: https://arxiv.org/abs/2311.17656

Chiradeep Roy, Mahsan Nourani, “Explainable Activity Recognition in Videos using Deep Learning and Tractable Probabilistic Models,” ACM Trans. Interact. Intell. Syst., vol. 13, no. 14, 2023, [Online]. Available: https://dl.acm.org/doi/10.1145/3626961

“Integrating YOLOv8 and IoT in a Computer Vision System for Child Detection in Smart Cities.” Accessed: May 09, 2026. [Online]. Available: https://thesai.org/Publications/ViewPaper?Volume=16&Issue=9&Code=IJACSA&SerialNo=60

K. L. Tran, M. N. Dang, T. N. Trong, H. N. Quoc, and L. N. Kieu, “Enhancing YOLOv11n for Reliable Child Detection in Noisy Surveillance Footage,” Feb. 2026, Accessed: May 09, 2026. [Online]. Available: http://arxiv.org/abs/2602.10592

“Opportunities, Applications, and Challenges of Edge-AI Enabled Video Analytics in Smart Cities: A Systematic Review | IEEE Journals & Magazine | IEEE Xplore.” Accessed: May 09, 2026. [Online]. Available: https://ieeexplore.ieee.org/document/10198424

“(PDF) Smart Home Monitoring System for Early Childhood Using Computer Vision Technology.” Accessed: May 09, 2026. [Online]. Available: https://www.researchgate.net/publication/395824244_Smart_Home_Monitoring_System_for_Early_Childhood_Using_Computer_Vision_Technology

“(PDF) Child Activity Recognition using Deep Learning.” Accessed: May 09, 2026. [Online]. Available: https://www.researchgate.net/publication/354778046_Child_Activity_Recognition_using_Deep_Learning

A. Sandygulova, A. Yershov, A. Zhanatkyzy, and Z. Telisheva, “ChildACT: Child Action Recognition Dataset in RGB Data,” ACM/IEEE Int. Conf. Human-Robot Interact., pp. 1088–1092, 2025, doi: 10.1109/HRI61500.2025.10974037.

“Falling Detection of Toddlers Based on Improved YOLOv8 Models - PubMed.” Accessed: May 09, 2026. [Online]. Available: https://pubmed.ncbi.nlm.nih.gov/39409491/

S. S. Hidayat, D. Aprilia, S. Hadwi, I. Mujahidin, M. C. A. Prabowo, and F. A. Rakhman, “Child Presence Detection for Child Safety with Deep Neural Networks,” J. Inform. J. Pengemb. IT, vol. 10, no. 2, pp. 370–381, Apr. 2025, doi: 10.30591/JPIT.V10I2.6540.

Y. Zhang et al., “ByteTrack: Multi-Object Tracking by Associating Every Detection Box,” Apr. 2022, Accessed: May 09, 2026. [Online]. Available: http://arxiv.org/abs/2110.06864

S. Zhu et al., “A Video-based End-to-end Pipeline for Non-nutritive Sucking Action Recognition and Segmentation in Young Infants,” Mar. 2023, Accessed: May 09, 2026. [Online]. Available: http://arxiv.org/abs/2303.16867

“EduNet: A New Video Dataset for Understanding Human Activity in the Classroom Environment.” Accessed: May 09, 2026. [Online]. Available: https://www.mdpi.com/1424-8220/21/17/5699

Klim Kireev, Ana-Maria Creţu, Raphael Meier, Sarah Adel Bargal, Elissa Redmiles, Carmela Troncoso, “A Manually Annotated Image-Caption Dataset for Detecting Children in the Wil,” arXiv:2506.10117, 2025, [Online]. Available: https://arxiv.org/abs/2506.10117

T. S. Yashwanth, Y. S. Royal, M. Kashyap, V. R. Shreya, and D. K. N, “Real Time Child Abduction and Detection System,” pp. 1–6, Dec. 2025, doi: 10.1109/SITA67914.2025.11273371.

V. Godase, “Edge AI for Smart Surveillance: Real-time Human Activity Recognition on Low-power Devices,” SSRN Electron. J., 2025, doi: 10.2139/SSRN.5383804.

Zihan Wang, Yang Yang, Zhi Liu, Yifan Zheng, “Deep Neural Networks in Video Human Action Recognition: A Review,” arXiv:2305.15692, 2023, [Online]. Available: https://arxiv.org/abs/2305.15692

Zhan Tong, Yibing Song, Jue Wang, Limin Wang, “VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training,” arXiv:2203.12602, 2022, [Online]. Available: https://arxiv.org/abs/2203.12602

Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, Cordelia Schmid, “ViViT: A Video Vision Transformer,” arXiv:2103.15691, 2021, [Online]. Available: https://arxiv.org/abs/2103.15691

Downloads

Published

2026-06-03
CITATION
Published: 2026-06-03
Crossref Citation Count: Loading...

How to Cite

Riaz, S., Shayan Riaz, Abdul Razzaq, Umar Sadique, Shahid Bashir, & Jehad Ur Rahman. (2026). An End-to-End Deep Learning Pipeline for Child Detection and Activity Recognition in Surveillance Systems. International Journal of Innovations in Science & Technology, 8(3), 1029–1048. https://doi.org/10.33411/IJIST/1873